MCP Nexus is a self-hosted control plane for MCP servers — "Kubernetes for
MCP." Agents connect to one endpoint; Nexus discovers services, installs the
right MCP servers, aggregates them behind a namespaced router, secures access
with auth/RBAC, and heals/updates them via a reconcile loop.
This first commit is design-phase only (no runnable code yet):
- README.md project front door + module map
- docs/ARCHITECTURE.md target design: tenets, two-plane split, reconcile
loop, domain types, storage, security, deployment
- docs/ROADMAP.md phased delivery (foundations -> walking skeleton ->
discover+install -> secure -> operate -> extend/scale)
- docs/modules/01-20 one design doc per module, all cross-linked
Backend stack decision: Go (single static binary, embedded SQLite + embedded
React dashboard). Repo initialized in /root with a whitelist .gitignore so
only project files are tracked.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
7.8 KiB
Infrastructure Discovery
Module 16 · Plane: Control · Roadmap phase: 2 Part of MCP Nexus architecture.
Purpose
Discover and connect to the compute substrates Nexus can deploy MCP containers onto — and enumerate the workloads already running on them. Where credentials allow, auto-connect and register each substrate as a Runtime target the Auto Installer can use. This is the "where can things run?" half of discovery, complementing Discovery Engine's "what services exist?".
Responsibilities
- Detect and probe substrates: Docker, Docker Compose projects, Kubernetes, Nomad, Proxmox, LXC, VirtualBox, VMware, Hyper-V, and bare metal (SSH-reachable hosts).
- Auto-connect where credentials/sockets are available (local Docker socket, in-cluster K8s service account,
KUBECONFIG, Proxmox API token, SSH keys), with graceful degradation to "detected but not connected" otherwise. - Register each connected substrate as a Runtime target — capabilities (can it run OCI containers? host-network? privileged?), capacity hints, and a health status the reconciler/Installer can place against.
- Enumerate existing workloads on each substrate (containers, pods, VMs, jails) and emit them so Discovery Engine's fingerprinters can classify them (a running container is both a host workload here and a candidate service there).
- Track substrate lifecycle: appear, become reachable/unreachable, drain, disappear — emitted on the shared event bus.
Non-goals
- Does not fingerprint application services into product types — it emits raw workloads for Discovery Engine to classify.
- Does not create/start MCP containers — it provides the Runtime targets; the Auto Installer does placement and lifecycle.
- Does not implement the Runtime abstraction (create/exec/logs/destroy). It populates Runtime targets; the Runtime interface itself is owned alongside the Installer.
- Does not manage substrate credentials at rest — those live in Secrets.
- Does not schedule/bin-pack beyond exposing capability + capacity hints (advanced placement is future work).
How the two discovery modules relate
Both are Control-plane, Phase 2, and share the event bus and Inventory store. The split is by target of discovery:
| Module 1 — Discovery Engine | Module 16 — Infrastructure Discovery | |
|---|---|---|
| Finds | Application services (Postgres, Home Assistant…) | Compute substrates/hosts (Docker, K8s, Proxmox…) |
| Emits | DiscoveredResource (a thing to manage via MCP) |
RuntimeTarget (a place to run MCP containers) + raw workloads |
| Feeds | Recipes → what MCP to install | Installer → where to install it |
They cooperate: Module 16 connects a Docker host, enumerates its containers, and forwards each as an Observation; Module 1's Docker fingerprinter then classifies those containers into typed DiscoveredResources. Conversely, when the Installer needs to place an MCPInstance, it picks from the RuntimeTargets Module 16 registered.
Interfaces
// A prober detects and (optionally) connects to one class of substrate.
type SubstrateProber interface {
Name() string // "docker", "kubernetes", "proxmox", ...
Detect(ctx context.Context, scope Scope) ([]Candidate, error)
Connect(ctx context.Context, c Candidate, creds SecretRef) (RuntimeTarget, error)
}
// RuntimeTarget is a connected substrate the Installer can place onto.
type RuntimeTarget struct {
ID string
Kind string // "docker" | "kubernetes" | "proxmox" | ...
Endpoint string // socket / api url / ssh host
Capabilities RuntimeCaps // OCI, host-network, privileged, gpu...
Capacity CapacityHint // cpu/mem/limits, best-effort
Health Health
Workloads []WorkloadRef // existing containers/pods/vms/jails
}
type Registry interface {
Register(t RuntimeTarget) error
Targets(ctx context.Context, filter TargetFilter) ([]RuntimeTarget, error)
Watch(ctx context.Context) (<-chan TargetEvent, error)
}
Internal HTTP (RBAC-guarded, not agent-facing):
GET /api/v1/infra/targets— list runtime targets + health/capabilities.POST /api/v1/infra/connect— attempt connection to a detected candidate with a secret ref.GET /api/v1/infra/targets/{id}/workloads— enumerate existing workloads.
Data
- Writes
RuntimeTargetrecords and their enumeratedWorkloadRefs into the Inventory store (§8), alongsideDiscoveredResources. - Reads substrate connection settings from Config; substrate credentials from Secrets (by ref only).
- Emits workloads as Observations consumed by Discovery Engine.
Dependencies
- Feeds the Runtime targeting used by the Auto Installer.
- Shares event bus + Inventory with Discovery Engine.
- Prober interfaces provided by the Plugin System; new runtime adapters (containerd, K8s operator, LXC) widen this in Phase 5.
- Credentials via Secrets; connect actions gated by RBAC and recorded in Audit.
- Substrates + workloads render as nodes in the Service Graph.
Failure modes & handling
| Failure | Behavior |
|---|---|
| Substrate detected but no credentials | Recorded as detected/unconnected; surfaced for operator to attach a secret. No auto-connect. |
| Credentials invalid / connection refused | Target marked unreachable with reason; backoff retry; Installer excludes it from placement. |
| Substrate becomes unreachable | Bound MCPInstances flagged for the reconciler; target drained, not deleted, until a grace period elapses. |
| Docker socket present but daemon dead | Distinguished from "no Docker" — reported as error, not absent, to avoid flapping. |
| Ambiguous host (both K8s node and Docker host) | Multiple targets registered, deduped by endpoint natural key; capabilities merged. |
| Privileged/host-network capability requested but substrate can't offer it | Capability advertised as false; Installer refuses placement rather than degrading isolation. |
Security notes
Honors Architecture §9: substrate credentials live only in Secrets, are injected at connect time, and are never logged or written into Inventory. Auto-connect is least-privilege — Nexus requests the minimum scope (e.g. a read+deploy K8s role, not cluster-admin) and prefers local sockets over broad network API tokens. Connecting to and enumerating a substrate is a mutating, audited action. Capability flags (privileged, host-network) are explicit so the Installer can honor the sandboxing defaults; a substrate that cannot sandbox is not silently used.
Open questions
- Is
RuntimeTargeta distinct first-class domain object, or a specialization ofDiscoveredResourcewithsource=infra? (Leaning: distinct, sharing the store.) - How much capacity/placement intelligence belongs here vs the Installer/reconciler?
- Hyper-V/VMware/VirtualBox are explicitly not early (Roadmap "not early") — ship as Phase 5 plugins?
- Credential discovery UX: how far do we go auto-detecting
KUBECONFIG/SSH configs before prompting?
Milestone
Delivered in Phase 2 (Discover → install). Thin slice that lands first: local Docker socket detection + auto-connect, registered as a RuntimeTarget, with existing container enumeration forwarded to Discovery Engine. This is the substrate the Phase 2 exit proof installs onto. Kubernetes, Proxmox, SSH/bare-metal, and the VM hypervisors follow as additional probers.