docs: initial architecture and design for MCP Nexus

MCP Nexus is a self-hosted control plane for MCP servers — "Kubernetes for
MCP." Agents connect to one endpoint; Nexus discovers services, installs the
right MCP servers, aggregates them behind a namespaced router, secures access
with auth/RBAC, and heals/updates them via a reconcile loop.

This first commit is design-phase only (no runnable code yet):
- README.md            project front door + module map
- docs/ARCHITECTURE.md target design: tenets, two-plane split, reconcile
                       loop, domain types, storage, security, deployment
- docs/ROADMAP.md      phased delivery (foundations -> walking skeleton ->
                       discover+install -> secure -> operate -> extend/scale)
- docs/modules/01-20   one design doc per module, all cross-linked

Backend stack decision: Go (single static binary, embedded SQLite + embedded
React dashboard). Repo initialized in /root with a whitelist .gitignore so
only project files are tracked.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
drjones
2026-07-07 04:38:27 +00:00
commit 8d3ffef920
24 changed files with 3000 additions and 0 deletions

View File

@@ -0,0 +1,48 @@
# Future Vision
> Module 20 · Plane: — · Roadmap phase: ongoing
> Part of [MCP Nexus architecture](../ARCHITECTURE.md).
## Purpose
This is the north star the other nineteen modules serve. Where the engineering docs describe *how* Nexus reconciles discovery, install, health, and routing, this one describes *why it matters* and what the finished thing feels like — so every design decision can be checked against a single question: **does this move us closer to AI infrastructure that configures itself?**
MCP Nexus aims to be the **operating system for AI infrastructure**. Not another MCP server, and not a directory of servers — the layer that makes an entire environment's worth of services instantly, safely usable by any AI agent, with zero manual MCP wiring.
## The North Star
You install Nexus on one host. It scans the environment — the LAN, the Docker daemon, the Kubernetes cluster, the homelab rack, the enterprise subnet. It finds every compatible service: Home Assistant, Postgres, Grafana, Ollama, UniFi, GitHub, the NAS, the reverse proxy. For each, it downloads the correct MCP package, configures it against the discovered service, secures it, keeps it healthy, and keeps it updated. Then it exposes **one endpoint**.
Every AI agent you own — Claude, Cursor, an internal copilot, a fleet of autonomous workers — points at that one endpoint and, in that instant, gains governed access to the *entire* infrastructure. No per-agent MCP config. No hand-maintained server lists. No copy-pasted credentials. The agent asks for what it's allowed to do; Nexus decides, routes, and audits.
## The guiding metaphor: plug and play for AI tools
Think about plugging a USB device into a computer. You don't hunt for a driver, edit a config file, or restart. The OS detects the device, identifies it, loads the right driver, and it just works.
**Nexus is that experience for AI tools.** Stand up a new service on your network and — the way a USB device is detected, matched to a driver, and made available — Nexus discovers it, matches it to a recipe, installs the right MCP "driver," and the corresponding tools appear behind the single endpoint. The agent that connected yesterday can use the new service today without anyone touching its config. Unplug the service and the tools drain away just as cleanly.
That is the whole promise in one sentence: **infrastructure that AI agents can use should configure itself, the way peripherals already do.**
## What "done" feels like
- **For the homelabber:** install one binary, open the dashboard, watch your whole rack light up as discovered services become live tools. Point Claude Desktop at Nexus once; never edit an MCP config again.
- **For the enterprise operator:** a single, audited, least-privilege gateway between every AI agent and every internal system. Roles decide who sees what; secrets never touch an agent; every tool call is logged. Turning on a new team of agents is a role assignment, not an integration project.
- **For the agent:** one endpoint, a tool set that matches exactly what it's permitted to do, and tools that appear and disappear as the real world changes — no stale integrations, no missing capabilities.
- **For the ecosystem:** anyone can publish a discovery method, a recipe, or an MCP package, and every Nexus instance can pick it up — because everything is a plugin (tenet #3).
## Progression
The vision is not a leap; it is the roadmap's phases compounding:
1. **One endpoint proven** (Phases 1–2) — the core loop: discover a service and its MCP just appears behind one address. This is the differentiating claim, proven early.
2. **Safe for many agents** (Phase 3) — auth, RBAC, secrets, audit turn "it works" into "it's trustworthy in a shared environment."
3. **Runs unattended** (Phase 4) — health self-healing, updates, metrics, notifications, and the dashboard make it a system you can leave alone and still understand.
4. **Open and at scale** (Phase 5) — the plugin system and HA topologies turn Nexus from a product into a platform the community extends and enterprises depend on.
5. **The OS layer** (beyond) — richer autonomy: recipes that self-tune, discovery that reaches more substrates, a service graph that reasons about dependencies, and confidence-driven automation that safely does more on its own over time.
Each phase already ends in something demonstrable; the north star is simply what they add up to.
## Risks & unknowns
- **Discovery is probabilistic, not certain** (tenet #6). "Plug and play" must stay honest: high-confidence detections auto-act; ambiguous ones ask a human. Over-automating erodes trust faster than under-automating.
- **Recipe & package trust.** A self-installing system is only as safe as its supply chain. Community recipes and packages need signing, governance, and a real trust model (Architecture §12) before "downloads the correct package" can be fully hands-off.
- **The USB metaphor has limits.** Real services have credentials, network policy, and blast radius a USB stick never does. Security-by-default and least privilege are what keep the convenience from becoming a liability.
- **Breadth vs. depth.** The value scales with how many services are recognized — but every new fingerprinter/recipe is surface area to maintain. The plugin system exists so breadth can grow without the core growing.
- **Autonomy pacing.** How much Nexus should *do on its own* versus *propose* is a dial, not a constant; it should move toward more autonomy only as confidence, auditability, and rollback prove themselves.
## Milestone
Ongoing — this doc has no ship date; it is the standard the shippable phases are measured against. The nearest concrete embodiment is the Phase 2 exit: **start a service on the LAN and watch its tools appear behind the single endpoint with zero manual MCP config.** That is the north star in miniature. Everything after makes it secure, unattended, extensible, and universal.