Files
mcp-gateway-nexus/docs/modules/01-discovery-engine.md
drjones 8d3ffef920 docs: initial architecture and design for MCP Nexus
MCP Nexus is a self-hosted control plane for MCP servers — "Kubernetes for
MCP." Agents connect to one endpoint; Nexus discovers services, installs the
right MCP servers, aggregates them behind a namespaced router, secures access
with auth/RBAC, and heals/updates them via a reconcile loop.

This first commit is design-phase only (no runnable code yet):
- README.md            project front door + module map
- docs/ARCHITECTURE.md target design: tenets, two-plane split, reconcile
                       loop, domain types, storage, security, deployment
- docs/ROADMAP.md      phased delivery (foundations -> walking skeleton ->
                       discover+install -> secure -> operate -> extend/scale)
- docs/modules/01-20   one design doc per module, all cross-linked

Backend stack decision: Go (single static binary, embedded SQLite + embedded
React dashboard). Repo initialized in /root with a whitelist .gitignore so
only project files are tracked.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-07 04:38:27 +00:00

7.5 KiB
Raw Permalink Blame History

Discovery Engine

Module 01 · Plane: Control · Roadmap phase: 2 Part of MCP Nexus architecture.

Purpose

Continuously scan the network and its runtimes for application services (Home Assistant, Postgres, Grafana, Ollama, …), fingerprint each into a typed DiscoveredResource with a confidence score, and emit it onto the event bus so the reconciler can match it against Recipes. This is the observe stage of the reconciliation loop (Architecture §6).

Responsibilities

  • Own a set of pluggable discovery methods that enumerate endpoints: mDNS/Bonjour/DNS-SD/Zeroconf, SSDP/UPnP, ICMP/ARP sweeps, HTTP/HTTPS probing, TLS certificate inspection, Docker API (containers + labels + Compose project metadata), Kubernetes API, Proxmox/VMware/LXC/Hyper-V inventories, SSH host-key fingerprints, SNMP, MQTT topic discovery, and Home Assistant discovery.
  • Own a set of pluggable fingerprinters that turn raw signals (open ports, HTTP titles/headers, TLS SANs, banners, container image names, labels) into a concrete type + version + capabilities. Starter type set: Home Assistant, UniFi, Immich, Paperless, Grafana, Prometheus, Jellyfin, Ollama, Postgres, MySQL/MariaDB, Redis, MinIO, Frigate, Synology, TrueNAS, Pi-hole, AdGuard, Traefik, Portainer, Nginx Proxy Manager, OpenWebUI, OpenAI-compatible APIs, LM Studio, vLLM, llama.cpp, ComfyUI, Automatic1111.
  • Assign every discovered resource a stable uuid (derived from a natural key so re-discovery is idempotent, not a new row) and a confidence score aggregated across the methods/fingerprinters that observed it.
  • Deduplicate and merge multi-source sightings of the same service (e.g. mDNS + Docker label + HTTP probe) into one DiscoveredResource, updating first_seen/last_seen.
  • Emit DiscoveredResource create/update/lost events onto the event bus and upsert them into Inventory.
  • Run on a schedule (periodic resync) and edge-triggered (e.g. a Docker event stream) per tenet #2.

Non-goals

  • Does not discover compute substrates/hosts to deploy onto — that is Infrastructure Discovery (Module 16). Module 1 finds what services exist; Module 16 finds where MCP containers can run.
  • Does not decide what to install. Matching a fingerprint to a package/config is Smart Recipes + the reconciler.
  • Does not install, configure, or run anything — that is the Auto Installer and Runtime.
  • Does not define plugin transport. It consumes the interfaces owned by the Plugin System; out-of-process discovery plugins land in Phase 5.
  • Does not store credentials. Method configs that need creds reference Secrets by ref.

Interfaces

// A discovery method enumerates candidate endpoints on the network/runtime.
type Method interface {
    Name() string                 // "mdns", "docker", "http-probe", ...
    Scan(ctx context.Context, scope Scope) (<-chan Observation, error)
}

// A fingerprinter inspects observations and proposes a typed identity + confidence.
type Fingerprinter interface {
    Name() string
    Match(ctx context.Context, o Observation) (Fingerprint, bool)
}

// Fingerprint is a scored, typed claim about an observation.
type Fingerprint struct {
    Type         string            // "home-assistant", "postgres", ...
    Version      string            // best-effort
    Capabilities []string          // e.g. ["rest-api","websocket"]
    Confidence   float64           // 0.0–1.0 for THIS signal
    Evidence     map[string]string // http_title, tls_san, image, banner...
}

// The engine fans methods → fingerprinters → merged resources.
type Engine interface {
    Register(m Method) error
    RegisterFingerprinter(f Fingerprinter) error
    Run(ctx context.Context) error            // periodic + edge-triggered
    Rescan(ctx context.Context, scope Scope) error
}

Internal HTTP (control API, not agent-facing; guarded by RBAC):

  • GET /api/v1/discovery/resources — list DiscoveredResources (filter by type/confidence/source).
  • POST /api/v1/discovery/rescan — trigger an on-demand scan (optionally scoped).
  • GET /api/v1/discovery/methods — list registered methods + last-run status.

Data

  • Writes DiscoveredResource (§7) to the Inventory store (§8) — {uuid, type, version, ip, hostname, ports, capabilities, health, confidence, source, first_seen, last_seen}.
  • Reads method configuration from the Config store and any method credentials from Secrets (by ref only).
  • Module-local persistence: a small per-method scan cursor/cache (last-seen sets, backoff timers) to make rescans cheap and dedup stable across restarts.

Dependencies

Failure modes & handling

Failure Behavior
A method errors or times out Isolated per-method; other methods continue. Error logged + surfaced in method status; exponential backoff before retry.
Aggressive scan (ARP/ICMP sweep) trips IDS / rate limits Sweeps are opt-in, rate-limited, and scoped to configured CIDRs; passive methods (mDNS/Docker) are default-on.
Two methods disagree on type Highest-confidence fingerprint wins; conflicting evidence lowers aggregate confidence and is retained in Evidence.
Confidence below threshold Resource is recorded but not auto-acted-on; flagged for human-in-the-loop review (tenet #6).
Service disappears Marked lost after a grace period (not deleted); emits a resource lost event so the reconciler can decide on the bound MCPInstance.
Duplicate sightings Merged by natural key into one uuid; never spawns duplicate instances.

Security notes

Honors Architecture §9: active scanning is least-privilege and opt-in — passive/local methods default on, network sweeps require explicit CIDR scope. Method credentials come from Secrets and are never logged or written into DiscoveredResource. Every scan and every discovery event is audited with source + principal. Discovery observes only; it holds no write access to runtimes. TLS inspection reads certificates without trusting them for auth.

Open questions

  • Confidence math: fixed weighted sum per signal, or a learned/Bayesian combiner tunable per environment?
  • How aggressive may default network sweeps be before they need explicit opt-in in a homelab?
  • Should lost grace periods be per-type (a laptop MCP vs a rack Postgres differ)?
  • Do we expose fingerprint Evidence to the dashboard for debugging, and does it risk leaking internal topology?

Milestone

Delivered in Phase 2 (Discover → install). Thin slice that lands first: Docker API + mDNS/DNS-SD + HTTP(S) probing methods with fingerprinters for the starter set (Home Assistant, Postgres, Ollama, Grafana, Portainer), emitting DiscoveredResource with confidence onto the event bus. Exit proof (with Recipes + Installer): start a Postgres container on the LAN → it is discovered, fingerprinted, and postgres.query appears at the Gateway within one reconcile cycle.