Files
mcp-gateway-nexus/docs/modules/16-infrastructure-discovery.md
drjones 8d3ffef920 docs: initial architecture and design for MCP Nexus
MCP Nexus is a self-hosted control plane for MCP servers — "Kubernetes for
MCP." Agents connect to one endpoint; Nexus discovers services, installs the
right MCP servers, aggregates them behind a namespaced router, secures access
with auth/RBAC, and heals/updates them via a reconcile loop.

This first commit is design-phase only (no runnable code yet):
- README.md            project front door + module map
- docs/ARCHITECTURE.md target design: tenets, two-plane split, reconcile
                       loop, domain types, storage, security, deployment
- docs/ROADMAP.md      phased delivery (foundations -> walking skeleton ->
                       discover+install -> secure -> operate -> extend/scale)
- docs/modules/01-20   one design doc per module, all cross-linked

Backend stack decision: Go (single static binary, embedded SQLite + embedded
React dashboard). Repo initialized in /root with a whitelist .gitignore so
only project files are tracked.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-07 04:38:27 +00:00

99 lines
7.8 KiB
Markdown

# Infrastructure Discovery
> Module 16 · Plane: Control · Roadmap phase: 2
> Part of [MCP Nexus architecture](../ARCHITECTURE.md).
## Purpose
Discover and connect to the **compute substrates** Nexus can deploy MCP containers onto — and enumerate the workloads already running on them. Where credentials allow, auto-connect and register each substrate as a **Runtime** target the [Auto Installer](03-auto-installer.md) can use. This is the "where can things run?" half of discovery, complementing [Discovery Engine](01-discovery-engine.md)'s "what services exist?".
## Responsibilities
- Detect and probe substrates: **Docker**, **Docker Compose** projects, **Kubernetes**, **Nomad**, **Proxmox**, **LXC**, **VirtualBox**, **VMware**, **Hyper-V**, and **bare metal** (SSH-reachable hosts).
- **Auto-connect** where credentials/sockets are available (local Docker socket, in-cluster K8s service account, `KUBECONFIG`, Proxmox API token, SSH keys), with graceful degradation to "detected but not connected" otherwise.
- Register each connected substrate as a **Runtime target** — capabilities (can it run OCI containers? host-network? privileged?), capacity hints, and a health status the reconciler/Installer can place against.
- Enumerate existing workloads on each substrate (containers, pods, VMs, jails) and emit them so Discovery Engine's fingerprinters can classify them (a running container is both a *host workload* here and a candidate *service* there).
- Track substrate lifecycle: appear, become reachable/unreachable, drain, disappear — emitted on the shared event bus.
## Non-goals
- **Does not fingerprint application services** into product types — it emits raw workloads for [Discovery Engine](01-discovery-engine.md) to classify.
- **Does not create/start MCP containers** — it provides the Runtime *targets*; the [Auto Installer](03-auto-installer.md) does placement and lifecycle.
- **Does not implement the Runtime abstraction** (create/exec/logs/destroy). It *populates* Runtime targets; the Runtime interface itself is owned alongside the Installer.
- **Does not manage substrate credentials at rest** — those live in [Secrets](07-secrets-manager.md).
- **Does not schedule/bin-pack** beyond exposing capability + capacity hints (advanced placement is future work).
## How the two discovery modules relate
Both are Control-plane, Phase 2, and **share the event bus and Inventory store**. The split is by *target of discovery*:
| | Module 1 — Discovery Engine | Module 16 — Infrastructure Discovery |
|---|---|---|
| Finds | Application **services** (Postgres, Home Assistant…) | Compute **substrates**/hosts (Docker, K8s, Proxmox…) |
| Emits | `DiscoveredResource` (a thing to *manage via* MCP) | `RuntimeTarget` (a place to *run* MCP containers) + raw workloads |
| Feeds | [Recipes](15-smart-recipes.md) → what MCP to install | [Installer](03-auto-installer.md) → where to install it |
They cooperate: Module 16 connects a Docker host, enumerates its containers, and forwards each as an Observation; Module 1's Docker fingerprinter then classifies those containers into typed `DiscoveredResource`s. Conversely, when the Installer needs to place an `MCPInstance`, it picks from the `RuntimeTarget`s Module 16 registered.
## Interfaces
```go
// A prober detects and (optionally) connects to one class of substrate.
type SubstrateProber interface {
Name() string // "docker", "kubernetes", "proxmox", ...
Detect(ctx context.Context, scope Scope) ([]Candidate, error)
Connect(ctx context.Context, c Candidate, creds SecretRef) (RuntimeTarget, error)
}
// RuntimeTarget is a connected substrate the Installer can place onto.
type RuntimeTarget struct {
ID string
Kind string // "docker" | "kubernetes" | "proxmox" | ...
Endpoint string // socket / api url / ssh host
Capabilities RuntimeCaps // OCI, host-network, privileged, gpu...
Capacity CapacityHint // cpu/mem/limits, best-effort
Health Health
Workloads []WorkloadRef // existing containers/pods/vms/jails
}
type Registry interface {
Register(t RuntimeTarget) error
Targets(ctx context.Context, filter TargetFilter) ([]RuntimeTarget, error)
Watch(ctx context.Context) (<-chan TargetEvent, error)
}
```
Internal HTTP (RBAC-guarded, not agent-facing):
- `GET /api/v1/infra/targets` — list runtime targets + health/capabilities.
- `POST /api/v1/infra/connect` — attempt connection to a detected candidate with a secret ref.
- `GET /api/v1/infra/targets/{id}/workloads` — enumerate existing workloads.
## Data
- **Writes** `RuntimeTarget` records and their enumerated `WorkloadRef`s into the **Inventory** store (§8), alongside `DiscoveredResource`s.
- **Reads** substrate connection settings from **Config**; substrate credentials from **Secrets** (by ref only).
- Emits workloads as Observations consumed by [Discovery Engine](01-discovery-engine.md).
## Dependencies
- Feeds the Runtime targeting used by the [Auto Installer](03-auto-installer.md).
- Shares event bus + Inventory with [Discovery Engine](01-discovery-engine.md).
- Prober interfaces provided by the [Plugin System](11-plugin-system.md); new runtime adapters (containerd, K8s operator, LXC) widen this in Phase 5.
- Credentials via [Secrets](07-secrets-manager.md); connect actions gated by [RBAC](08-rbac.md) and recorded in Audit.
- Substrates + workloads render as nodes in the [Service Graph](17-service-graph.md).
## Failure modes & handling
| Failure | Behavior |
|---|---|
| Substrate detected but no credentials | Recorded as `detected/unconnected`; surfaced for operator to attach a secret. No auto-connect. |
| Credentials invalid / connection refused | Target marked `unreachable` with reason; backoff retry; Installer excludes it from placement. |
| Substrate becomes unreachable | Bound `MCPInstance`s flagged for the reconciler; target drained, not deleted, until a grace period elapses. |
| Docker socket present but daemon dead | Distinguished from "no Docker" — reported as `error`, not absent, to avoid flapping. |
| Ambiguous host (both K8s node and Docker host) | Multiple targets registered, deduped by endpoint natural key; capabilities merged. |
| Privileged/host-network capability requested but substrate can't offer it | Capability advertised as false; Installer refuses placement rather than degrading isolation. |
## Security notes
Honors Architecture §9: substrate credentials live only in [Secrets](07-secrets-manager.md), are injected at connect time, and are never logged or written into Inventory. Auto-connect is **least-privilege** — Nexus requests the minimum scope (e.g. a read+deploy K8s role, not cluster-admin) and prefers local sockets over broad network API tokens. Connecting to and enumerating a substrate is a mutating, audited action. Capability flags (privileged, host-network) are explicit so the Installer can honor the sandboxing defaults; a substrate that cannot sandbox is not silently used.
## Open questions
- Is `RuntimeTarget` a distinct first-class domain object, or a specialization of `DiscoveredResource` with `source=infra`? (Leaning: distinct, sharing the store.)
- How much capacity/placement intelligence belongs here vs the Installer/reconciler?
- Hyper-V/VMware/VirtualBox are explicitly *not early* (Roadmap "not early") — ship as Phase 5 plugins?
- Credential discovery UX: how far do we go auto-detecting `KUBECONFIG`/SSH configs before prompting?
## Milestone
Delivered in **Phase 2** (Discover → install). Thin slice that lands first: **local Docker socket** detection + auto-connect, registered as a `RuntimeTarget`, with existing container enumeration forwarded to [Discovery Engine](01-discovery-engine.md). This is the substrate the Phase 2 exit proof installs onto. Kubernetes, Proxmox, SSH/bare-metal, and the VM hypervisors follow as additional probers.