Files
mcp-gateway-nexus/docs/modules/16-infrastructure-discovery.md
drjones 8d3ffef920 docs: initial architecture and design for MCP Nexus
MCP Nexus is a self-hosted control plane for MCP servers — "Kubernetes for
MCP." Agents connect to one endpoint; Nexus discovers services, installs the
right MCP servers, aggregates them behind a namespaced router, secures access
with auth/RBAC, and heals/updates them via a reconcile loop.

This first commit is design-phase only (no runnable code yet):
- README.md            project front door + module map
- docs/ARCHITECTURE.md target design: tenets, two-plane split, reconcile
                       loop, domain types, storage, security, deployment
- docs/ROADMAP.md      phased delivery (foundations -> walking skeleton ->
                       discover+install -> secure -> operate -> extend/scale)
- docs/modules/01-20   one design doc per module, all cross-linked

Backend stack decision: Go (single static binary, embedded SQLite + embedded
React dashboard). Repo initialized in /root with a whitelist .gitignore so
only project files are tracked.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-07 04:38:27 +00:00

7.8 KiB

Infrastructure Discovery

Module 16 · Plane: Control · Roadmap phase: 2 Part of MCP Nexus architecture.

Purpose

Discover and connect to the compute substrates Nexus can deploy MCP containers onto — and enumerate the workloads already running on them. Where credentials allow, auto-connect and register each substrate as a Runtime target the Auto Installer can use. This is the "where can things run?" half of discovery, complementing Discovery Engine's "what services exist?".

Responsibilities

  • Detect and probe substrates: Docker, Docker Compose projects, Kubernetes, Nomad, Proxmox, LXC, VirtualBox, VMware, Hyper-V, and bare metal (SSH-reachable hosts).
  • Auto-connect where credentials/sockets are available (local Docker socket, in-cluster K8s service account, KUBECONFIG, Proxmox API token, SSH keys), with graceful degradation to "detected but not connected" otherwise.
  • Register each connected substrate as a Runtime target — capabilities (can it run OCI containers? host-network? privileged?), capacity hints, and a health status the reconciler/Installer can place against.
  • Enumerate existing workloads on each substrate (containers, pods, VMs, jails) and emit them so Discovery Engine's fingerprinters can classify them (a running container is both a host workload here and a candidate service there).
  • Track substrate lifecycle: appear, become reachable/unreachable, drain, disappear — emitted on the shared event bus.

Non-goals

  • Does not fingerprint application services into product types — it emits raw workloads for Discovery Engine to classify.
  • Does not create/start MCP containers — it provides the Runtime targets; the Auto Installer does placement and lifecycle.
  • Does not implement the Runtime abstraction (create/exec/logs/destroy). It populates Runtime targets; the Runtime interface itself is owned alongside the Installer.
  • Does not manage substrate credentials at rest — those live in Secrets.
  • Does not schedule/bin-pack beyond exposing capability + capacity hints (advanced placement is future work).

How the two discovery modules relate

Both are Control-plane, Phase 2, and share the event bus and Inventory store. The split is by target of discovery:

Module 1 — Discovery Engine Module 16 — Infrastructure Discovery
Finds Application services (Postgres, Home Assistant…) Compute substrates/hosts (Docker, K8s, Proxmox…)
Emits DiscoveredResource (a thing to manage via MCP) RuntimeTarget (a place to run MCP containers) + raw workloads
Feeds Recipes → what MCP to install Installer → where to install it

They cooperate: Module 16 connects a Docker host, enumerates its containers, and forwards each as an Observation; Module 1's Docker fingerprinter then classifies those containers into typed DiscoveredResources. Conversely, when the Installer needs to place an MCPInstance, it picks from the RuntimeTargets Module 16 registered.

Interfaces

// A prober detects and (optionally) connects to one class of substrate.
type SubstrateProber interface {
    Name() string                          // "docker", "kubernetes", "proxmox", ...
    Detect(ctx context.Context, scope Scope) ([]Candidate, error)
    Connect(ctx context.Context, c Candidate, creds SecretRef) (RuntimeTarget, error)
}

// RuntimeTarget is a connected substrate the Installer can place onto.
type RuntimeTarget struct {
    ID           string
    Kind         string            // "docker" | "kubernetes" | "proxmox" | ...
    Endpoint     string            // socket / api url / ssh host
    Capabilities RuntimeCaps       // OCI, host-network, privileged, gpu...
    Capacity     CapacityHint      // cpu/mem/limits, best-effort
    Health       Health
    Workloads    []WorkloadRef     // existing containers/pods/vms/jails
}

type Registry interface {
    Register(t RuntimeTarget) error
    Targets(ctx context.Context, filter TargetFilter) ([]RuntimeTarget, error)
    Watch(ctx context.Context) (<-chan TargetEvent, error)
}

Internal HTTP (RBAC-guarded, not agent-facing):

  • GET /api/v1/infra/targets — list runtime targets + health/capabilities.
  • POST /api/v1/infra/connect — attempt connection to a detected candidate with a secret ref.
  • GET /api/v1/infra/targets/{id}/workloads — enumerate existing workloads.

Data

  • Writes RuntimeTarget records and their enumerated WorkloadRefs into the Inventory store (§8), alongside DiscoveredResources.
  • Reads substrate connection settings from Config; substrate credentials from Secrets (by ref only).
  • Emits workloads as Observations consumed by Discovery Engine.

Dependencies

  • Feeds the Runtime targeting used by the Auto Installer.
  • Shares event bus + Inventory with Discovery Engine.
  • Prober interfaces provided by the Plugin System; new runtime adapters (containerd, K8s operator, LXC) widen this in Phase 5.
  • Credentials via Secrets; connect actions gated by RBAC and recorded in Audit.
  • Substrates + workloads render as nodes in the Service Graph.

Failure modes & handling

Failure Behavior
Substrate detected but no credentials Recorded as detected/unconnected; surfaced for operator to attach a secret. No auto-connect.
Credentials invalid / connection refused Target marked unreachable with reason; backoff retry; Installer excludes it from placement.
Substrate becomes unreachable Bound MCPInstances flagged for the reconciler; target drained, not deleted, until a grace period elapses.
Docker socket present but daemon dead Distinguished from "no Docker" — reported as error, not absent, to avoid flapping.
Ambiguous host (both K8s node and Docker host) Multiple targets registered, deduped by endpoint natural key; capabilities merged.
Privileged/host-network capability requested but substrate can't offer it Capability advertised as false; Installer refuses placement rather than degrading isolation.

Security notes

Honors Architecture §9: substrate credentials live only in Secrets, are injected at connect time, and are never logged or written into Inventory. Auto-connect is least-privilege — Nexus requests the minimum scope (e.g. a read+deploy K8s role, not cluster-admin) and prefers local sockets over broad network API tokens. Connecting to and enumerating a substrate is a mutating, audited action. Capability flags (privileged, host-network) are explicit so the Installer can honor the sandboxing defaults; a substrate that cannot sandbox is not silently used.

Open questions

  • Is RuntimeTarget a distinct first-class domain object, or a specialization of DiscoveredResource with source=infra? (Leaning: distinct, sharing the store.)
  • How much capacity/placement intelligence belongs here vs the Installer/reconciler?
  • Hyper-V/VMware/VirtualBox are explicitly not early (Roadmap "not early") — ship as Phase 5 plugins?
  • Credential discovery UX: how far do we go auto-detecting KUBECONFIG/SSH configs before prompting?

Milestone

Delivered in Phase 2 (Discover → install). Thin slice that lands first: local Docker socket detection + auto-connect, registered as a RuntimeTarget, with existing container enumeration forwarded to Discovery Engine. This is the substrate the Phase 2 exit proof installs onto. Kubernetes, Proxmox, SSH/bare-metal, and the VM hypervisors follow as additional probers.