docs: initial architecture and design for MCP Nexus

MCP Nexus is a self-hosted control plane for MCP servers — "Kubernetes for
MCP." Agents connect to one endpoint; Nexus discovers services, installs the
right MCP servers, aggregates them behind a namespaced router, secures access
with auth/RBAC, and heals/updates them via a reconcile loop.

This first commit is design-phase only (no runnable code yet):
- README.md            project front door + module map
- docs/ARCHITECTURE.md target design: tenets, two-plane split, reconcile
                       loop, domain types, storage, security, deployment
- docs/ROADMAP.md      phased delivery (foundations -> walking skeleton ->
                       discover+install -> secure -> operate -> extend/scale)
- docs/modules/01-20   one design doc per module, all cross-linked

Backend stack decision: Go (single static binary, embedded SQLite + embedded
React dashboard). Repo initialized in /root with a whitelist .gitignore so
only project files are tracked.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
drjones
2026-07-07 04:38:27 +00:00
commit 8d3ffef920
24 changed files with 3000 additions and 0 deletions

10
.gitignore vendored Normal file
View File

@@ -0,0 +1,10 @@
# The repository lives in /root (a home directory), so ignore everything by
# default and explicitly whitelist only project files. This prevents shell
# dotfiles, credentials, and caches from ever being tracked.
/*
# Whitelisted project paths
!/.gitignore
!/README.md
!/LICENSE
!/docs

100
README.md Normal file
View File

@@ -0,0 +1,100 @@
<h1 align="center">MCP Nexus</h1>
<p align="center"><b>The control plane for MCP. One endpoint. Every tool. Zero manual config.</b></p>
<p align="center">
<i>Kubernetes for MCP — discovery, install, routing, auth, and healing for every MCP server on your network.</i>
</p>
---
> **Status: design phase.** This repository currently contains the architecture and design docs.
> No runnable code yet. See the [Roadmap](./docs/ROADMAP.md) for the delivery plan and
> [Architecture](./docs/ARCHITECTURE.md) for the full design.
## The problem
Every AI agent (Claude, Cursor, VSCode, Goose, OpenWebUI, Aider, …) has to be told about every MCP server individually — install it, configure it, hold its secrets, keep it updated, wire it into each agent's config. Multiply that by every agent × every service on a homelab or in an enterprise and it collapses under its own weight.
## The idea
Run **one** thing. Agents connect to a single endpoint:
```
https://mcp.company.lan
```
Behind that endpoint, **MCP Nexus**:
- **Discovers** the services on your network (Home Assistant, Proxmox, Postgres, Ollama, GitHub, UniFi, …).
- **Looks up** the right MCP server for each in a package registry.
- **Installs** it as a sandboxed container and **configures** it automatically.
- **Aggregates** every MCP server behind one endpoint with **namespaced** tools (`homeassistant.turn_on`, `postgres.query`, `github.create_issue`).
- **Secures** access with auth + RBAC, so each agent sees only the tools it's allowed to.
- **Heals** unhealthy servers and **updates** them, with rollback.
> Think of it like plugging a USB device into a modern OS: it's detected, the right driver is installed, and it just works — but for every AI tool and service across your infrastructure.
## How it works (30-second version)
MCP Nexus is a **reconciler** — like Kubernetes, but for MCP servers. It continuously drives *what's actually running* toward *what should be running* given the services it discovers and the recipes/policies you set.
```
Discovery ─▶ Registry/Recipes ─▶ Installer ─▶ Runtime ─▶ Inventory
│
Agents ─▶ Gateway ─▶ Auth/RBAC ─▶ Router ─▶ upstream MCP servers ─▶ your services
```
Full picture: **[docs/ARCHITECTURE.md](./docs/ARCHITECTURE.md)**.
## Design tenets
1. One endpoint, unlimited agents.
2. Reconcile, don't script.
3. Everything is a plugin.
4. Single static binary first (Go); Postgres/K8s optional for scale-out.
5. Secure by default — sandboxed, least-privilege, audited, secrets never leak to agents.
6. Discovery is probabilistic — every result carries a confidence score.
## Tech stack
- **Backend / control plane:** Go (single static binary, embedded SQLite, embedded dashboard).
- **Protocol:** MCP over JSON-RPC 2.0 (Streamable HTTP + stdio bridge).
- **Runtime:** Docker first; containerd / Kubernetes adapters planned.
- **Dashboard:** React, embedded in the binary.
- **Observability:** Prometheus metrics + shipped Grafana dashboards.
## Modules
The system is 20 modules, each with its own design doc in [`docs/modules/`](./docs/modules/):
| Control plane | Data plane | Cross-cutting |
|---|---|---|
| [Discovery Engine](./docs/modules/01-discovery-engine.md) | [Gateway](./docs/modules/04-gateway.md) | [Secrets Manager](./docs/modules/07-secrets-manager.md) |
| [Package Registry](./docs/modules/02-package-registry.md) | [Dynamic Tool Registry](./docs/modules/05-dynamic-tool-registry.md) | [Plugin System](./docs/modules/11-plugin-system.md) |
| [Auto Installer](./docs/modules/03-auto-installer.md) | [Authentication](./docs/modules/06-authentication.md) | [Notifications](./docs/modules/13-notifications.md) |
| [Health Monitoring](./docs/modules/09-health-monitoring.md) | [RBAC](./docs/modules/08-rbac.md) | [Service Graph](./docs/modules/17-service-graph.md) |
| [Update Manager](./docs/modules/10-update-manager.md) | [Web Dashboard](./docs/modules/12-web-dashboard.md) | [Metrics](./docs/modules/18-metrics.md) |
| [Smart Recipes](./docs/modules/15-smart-recipes.md) | [AI Agent Profiles](./docs/modules/14-agent-profiles.md) | [Security](./docs/modules/19-security.md) |
| [Infrastructure Discovery](./docs/modules/16-infrastructure-discovery.md) | | [Future Vision](./docs/modules/20-future-vision.md) |
## Roadmap at a glance
| Phase | Goal |
|---|---|
| **0 — Foundations** | Repo, domain types, storage, config, event bus. |
| **1 — Walking skeleton** | Gateway aggregates a manually-registered MCP server behind one endpoint. |
| **2 — Discover → install** | Discovery + registry + recipes + installer close the reconcile loop. |
| **3 — Secure** | Auth, RBAC, secrets, per-agent tool scoping. |
| **4 — Operate** | Health, updates, metrics, notifications, dashboard. |
| **5 — Extend & scale** | Plugin system, service graph, K8s runtime, HA. |
Details: **[docs/ROADMAP.md](./docs/ROADMAP.md)**.
## Contributing
Design-phase feedback is the most valuable contribution right now — open an issue on the [design docs](./docs/). Coding conventions and a `CONTRIBUTING.md` will land with Phase 0.
## License
TBD (intended to be a permissive open-source license — see roadmap).

185
docs/ARCHITECTURE.md Normal file
View File

@@ -0,0 +1,185 @@
# MCP Nexus — Architecture
> Status: **Design draft (v0)**. This document defines the target architecture.
> No production code exists yet; the [Roadmap](./ROADMAP.md) sequences delivery.
## 1. What MCP Nexus is
MCP Nexus is a **self-hosted control plane for [Model Context Protocol (MCP)](https://modelcontextprotocol.io) servers**. It is the single source of truth for every MCP server on a network.
AI agents connect to **one endpoint**. Behind it, Nexus discovers infrastructure, installs the right MCP servers, configures them, secures them, keeps them healthy and up to date, and routes agent tool calls to the correct backend.
The mental model: **"Kubernetes for MCP."** Nexus is a reconciler that continuously drives *observed state* (what MCP servers are actually running) toward *desired state* (what should be running, given discovered services + recipes + policy).
## 2. Design tenets
These are load-bearing. Every module doc must be consistent with them.
1. **One endpoint, many agents.** Agents never learn about individual MCP servers. They speak MCP to the Gateway; the Gateway is an MCP server to them and an MCP client to everything upstream.
2. **Reconcile, don't script.** State changes flow through a control loop (desired vs observed), not imperative one-shot commands. This is what makes discovery→install→heal→update robust.
3. **Everything is a plugin.** Discovery methods, fingerprinters, package providers, installers, notifiers, and auth backends are all interfaces with swappable implementations. The core knows only the interfaces.
4. **Single static binary first.** Nexus ships as one Go binary with embedded storage and embedded dashboard assets. External Postgres/Redis are *optional* scale-out choices, never requirements.
5. **Secure by default.** Least privilege, sandboxed MCP containers, secrets never exposed to agents unless policy allows, every action audited, TLS everywhere, signed packages.
6. **Confidence, not certainty.** Discovery/fingerprinting is probabilistic. Every discovered resource carries a confidence score; automated actions above a threshold, human-in-the-loop below it.
## 3. Technology choices
| Concern | Choice | Rationale |
|---|---|---|
| Control-plane language | **Go** | Single static binary, first-class Docker/K8s client libs, strong concurrency for discovery + routing, trivial to self-host. |
| MCP transport (agent ↔ Nexus) | **Streamable HTTP** (+ stdio shim) | The current MCP HTTP transport; stdio shim lets local agents (Claude Desktop, Cursor) attach via a thin stdio→HTTP bridge. |
| MCP transport (Nexus ↔ upstream) | **stdio and Streamable HTTP** | Upstream MCP servers vary; the client layer abstracts transport. |
| Wire protocol | **JSON-RPC 2.0** | MCP is JSON-RPC 2.0. |
| Embedded state | **SQLite (via `modernc.org/sqlite`, cgo-free)** | Single-binary friendly; keeps registry, inventory, audit, config. |
| Optional scale-out state | **Postgres** | For HA / multi-node deployments. |
| Container runtime | **Docker API** (containerd/K8s adapters later) | Installer targets Docker first; runtime is an interface. |
| Dashboard | **React (embedded via `embed.FS`)** | Served by the same binary; no separate deploy. |
| Metrics | **Prometheus exposition** | `/metrics`; Grafana dashboards shipped as JSON. |
| Config | **Declarative YAML + API** | Desired state is data, matching tenet #2. |
## 4. High-level component map
```
AI Agents (Claude, Cursor, VSCode, Goose, OpenWebUI, …)
│ MCP (Streamable HTTP / stdio bridge)
▼
┌───────────────────────────────────────────────────────────────────────┐
│ MCP NEXUS CONTROL PLANE │
│ │
│ ┌─────────────┐ ┌──────────────┐ ┌───────────────────────────┐ │
│ │ API Gateway │──▶│ Auth / RBAC │──▶│ MCP Router + Aggregator │ │
│ │ (agent edge)│ │ (policy) │ │ (namespaced tool routing) │ │
│ └─────────────┘ └──────────────┘ └────────────┬──────────────┘ │
│ │ │ │
│ ▼ ▼ │
│ ┌─────────────┐ ┌───────────────────┐ │
│ │ Dashboard │ │ Upstream MCP client│ │
│ │ (React) │ │ pool (per server) │ │
│ └─────────────┘ └─────────┬─────────┘ │
│ │ │
│ ══════════════════ CONTROL LOOP (reconciler) ══════╪═════════════════ │
│ │ │
│ ┌───────────┐ ┌──────────┐ ┌───────────┐ ┌──────┴─────┐ │
│ │ Discovery │─▶│ Registry │─▶│ Installer │─▶│ Runtime │ │
│ │ Engine │ │ + Recipes│ │ │ │ (Docker…) │ │
│ └───────────┘ └──────────┘ └───────────┘ └────────────┘ │
│ │ │ │
│ ▼ ▼ │
│ ┌───────────┐ ┌──────────┐ ┌─────────┐ ┌────────┐ ┌──────────┐ │
│ │ Inventory │ │ Health │ │ Updater │ │Secrets │ │ Audit / │ │
│ │ (state) │ │ Monitor │ │ │ │ Vault │ │ Metrics │ │
│ └───────────┘ └──────────┘ └─────────┘ └────────┘ └──────────┘ │
└───────────────────────────────────────────────────────────────────────┘
│
Docker · K8s · LXC · VMs · Bare metal
│
Home Assistant · UniFi · Proxmox · Postgres · Ollama · GitHub · Grafana · …
```
## 5. The two planes
Nexus has a clean split that the module docs inherit:
### Data plane (hot path, per request)
`Agent → API Gateway → Auth/RBAC → MCP Router → upstream MCP client → MCP server → real service`
Latency-sensitive. Must stay up even while the control plane reconciles. Concerns: connection pooling, tool namespacing, request scoping to the agent's RBAC-visible tool set, streaming responses, rate limiting, audit.
### Control plane (cold path, continuous)
`Discovery → Registry/Recipes → Installer → Runtime → Inventory`, with `Health`, `Updater`, `Secrets`, and `Notifications` reacting to inventory state.
Eventually consistent. Runs as background reconcilers. A failure here degrades *management* (no new installs) but must **not** take down the data plane.
## 6. The reconciliation loop
The heart of the system (tenet #2). One loop, many controllers, all edge-triggered off an event bus plus periodic resync:
```
observe Discovery emits DiscoveredResource events → Inventory
decide Reconciler diffs desired (Recipes ⨯ Inventory ⨯ Policy) vs observed
→ produces a plan of actions
act Installer / Updater / Health execute actions via Runtime
record Inventory + Audit updated; Router refreshes upstream set
notify Notifications fire on meaningful transitions
```
Desired state = for each discovered service with confidence ≥ threshold, the recipe-matched MCP server should be installed, configured, running, healthy, and current. The reconciler is idempotent: re-running with the same inputs is a no-op.
## 7. Core domain objects
These types are shared vocabulary across every module.
- **DiscoveredResource** — `{uuid, type, version, ip, hostname, ports, capabilities, health, confidence, source, first_seen, last_seen}`. Output of Discovery.
- **Recipe** — declarative match+install+config rule (see [Smart Recipes](./modules/15-smart-recipes.md)). Maps a fingerprint to an MCP package + config template.
- **Package** — a resolvable MCP server artifact `{name, source(github|oci|dockerhub|local), image/ref, versions, config_schema, signature}` (see [Registry](./modules/02-package-registry.md)).
- **MCPInstance** — a managed running MCP server `{id, package, version, config, container_ref, state, health, bound_resource_uuid}`. The unit the reconciler manages.
- **Tool** — an MCP tool exposed by an instance, namespaced as `{namespace}.{tool}` at the Gateway (see [Dynamic Tool Registry](./modules/05-dynamic-tool-registry.md)).
- **AgentProfile** — `{id, identity, allowed_roles}` → resolves to a visible tool set (see [AI Agent Profiles](./modules/14-agent-profiles.md)).
- **Role** — RBAC grant mapping principals to allowed namespaces/tools (see [RBAC](./modules/08-rbac.md)).
- **Secret** — an encrypted credential referenced by config templates, never returned to agents (see [Secrets](./modules/07-secrets-manager.md)).
## 8. Storage model
Single embedded SQLite database (Postgres-compatible schema for scale-out). Logical stores:
| Store | Holds | Notes |
|---|---|---|
| Inventory | DiscoveredResources, MCPInstances | Source of observed + desired state |
| Registry | Packages, Recipes | Syncable from remote indexes |
| Identity | Users, AgentProfiles, Roles, API keys | RBAC + auth |
| Secrets | Encrypted secrets, envelope keys | Encrypted at rest; see Secrets module |
| Audit | Append-only action log | Immutable; every mutating action |
| Config | Desired-state overrides, settings | Declarative YAML mirrored here |
## 9. Security architecture (summary)
Full detail in [Security](./modules/19-security.md). Cross-cutting rules every module honors:
- **Isolation:** each MCP instance runs in its own sandboxed container; least-privilege, read-only rootfs where possible, no host network unless the recipe demands it.
- **Secret handling:** secrets are injected into MCP containers at runtime (env/file mount), never persisted in config sent to agents, never logged.
- **Request scoping:** the Router only ever exposes an agent the tools its resolved RBAC allows — an agent cannot call, or even see, a tool outside its role.
- **Supply chain:** packages are signature-verified before install; pinned by digest.
- **Audit:** every mutating control-plane action and every agent tool call is recorded with principal, target, and outcome.
- **Transport:** TLS terminated at the API Gateway; internal component calls over localhost/socket.
## 10. Deployment topologies
1. **Single binary** (default) — `nexus` runs on a host with Docker; embedded SQLite; embedded dashboard. Target: homelab.
2. **Container** — the same binary in a container with the Docker socket mounted (or a remote Docker/K8s endpoint configured).
3. **HA / multi-node** — multiple `nexus` replicas behind a load balancer, shared Postgres, leader-elected reconciler. Target: enterprise.
## 11. Module index
Each module has a dedicated design doc under [`docs/modules/`](./modules/). They share the template: *Purpose · Responsibilities · Interfaces · Data · Dependencies · Failure modes · Open questions · Milestone.*
| # | Module | Plane | Doc |
|---|---|---|---|
| 1 | Discovery Engine | Control | [01](./modules/01-discovery-engine.md) |
| 2 | MCP Package Registry | Control | [02](./modules/02-package-registry.md) |
| 3 | Auto Installer | Control | [03](./modules/03-auto-installer.md) |
| 4 | Gateway | Data | [04](./modules/04-gateway.md) |
| 5 | Dynamic Tool Registry | Data | [05](./modules/05-dynamic-tool-registry.md) |
| 6 | Authentication | Data | [06](./modules/06-authentication.md) |
| 7 | Secrets Manager | Cross-cutting | [07](./modules/07-secrets-manager.md) |
| 8 | RBAC | Data | [08](./modules/08-rbac.md) |
| 9 | Health Monitoring | Control | [09](./modules/09-health-monitoring.md) |
| 10 | Update Manager | Control | [10](./modules/10-update-manager.md) |
| 11 | Plugin System | Cross-cutting | [11](./modules/11-plugin-system.md) |
| 12 | Web Dashboard | Data | [12](./modules/12-web-dashboard.md) |
| 13 | Notifications | Cross-cutting | [13](./modules/13-notifications.md) |
| 14 | AI Agent Profiles | Data | [14](./modules/14-agent-profiles.md) |
| 15 | Smart Recipes | Control | [15](./modules/15-smart-recipes.md) |
| 16 | Infrastructure Discovery | Control | [16](./modules/16-infrastructure-discovery.md) |
| 17 | Service Graph | Cross-cutting | [17](./modules/17-service-graph.md) |
| 18 | Metrics | Cross-cutting | [18](./modules/18-metrics.md) |
| 19 | Security | Cross-cutting | [19](./modules/19-security.md) |
| 20 | Future Vision | — | [20](./modules/20-future-vision.md) |
## 12. Open architectural questions
Tracked here until resolved; each module may add its own.
- **Stdio-only agents:** ship a `nexus connect` stdio↔HTTP bridge binary, or document per-agent proxy config? (Leaning: ship the bridge.)
- **Multi-tenancy depth:** is a "project"/"tenant" a first-class object above roles, or is RBAC enough for v1? (Leaning: RBAC only for v1.)
- **Recipe distribution:** community recipe repo governance and trust model.
- **K8s runtime parity:** how much of the Docker installer semantics map cleanly to a K8s operator vs a separate controller.

122
docs/ROADMAP.md Normal file
View File

@@ -0,0 +1,122 @@
# MCP Nexus — Roadmap
> This roadmap sequences the [architecture](./ARCHITECTURE.md) into shippable phases.
> Ordering principle: **prove the core loop end-to-end early, then widen and harden.**
> Each phase ends in something demonstrable, not just more scaffolding.
## Guiding constraints
- Every phase must leave `main` runnable (after Phase 1) — no long-lived broken states.
- Data plane and control plane evolve on separate tracks; the data plane must never depend on an in-progress control-plane feature to serve traffic.
- Security is not a phase you bolt on. Auth/secrets/audit hooks are stubbed from Phase 0 and filled in Phase 3, but the *seams* exist from the start.
---
## Phase 0 — Foundations
**Goal:** a Go project you can build, test, and configure. No behavior yet.
- Repo layout (`cmd/nexus`, `internal/…`, `pkg/…`), Makefile, CI (build + test + lint), `CONTRIBUTING.md`, license.
- Core domain types (§7 of Architecture): `DiscoveredResource`, `Package`, `Recipe`, `MCPInstance`, `Tool`, `AgentProfile`, `Role`, `Secret`.
- Storage layer: embedded SQLite with migrations; the six logical stores as interfaces.
- Config loader: declarative YAML → desired state; hot reload.
- Event bus + reconciler skeleton (register controllers, no controllers yet).
- Structured logging + audit-log sink interface.
**Exit criteria:** `nexus serve` boots, loads config, exposes `/healthz` and `/metrics`, persists to SQLite. Nothing MCP yet.
---
## Phase 1 — Walking skeleton (data plane)
**Goal:** an agent connects to Nexus and calls a tool on a **manually registered** upstream MCP server.
- MCP server implementation (agent edge): Streamable HTTP transport, JSON-RPC 2.0, `initialize`/`tools/list`/`tools/call`.
- Upstream MCP **client** (stdio + Streamable HTTP) with a connection pool per instance.
- **Router + Aggregator:** multiplex N upstream servers, namespace tools (`ns.tool`), route `tools/call` to the owning instance, stream results back.
- Static instance registration via config (skip discovery/installer for now).
- `nexus connect` stdio↔HTTP bridge so local agents (Claude Desktop, Cursor) can attach.
**Exit criteria:** register a filesystem MCP + one HTTP MCP in YAML; connect Claude; see `filesystem.*` and `other.*` tools; call one successfully. **This is the core value proposition proven.**
Modules: [Gateway](./modules/04-gateway.md), [Dynamic Tool Registry](./modules/05-dynamic-tool-registry.md).
---
## Phase 2 — Discover → install (control plane)
**Goal:** close the reconcile loop. Discover a service, auto-install its MCP, and it shows up at the Gateway with no manual step.
- **Discovery Engine** v1: Docker API + mDNS/DNS-SD + HTTP(S) probing; emits `DiscoveredResource` with confidence.
- **Fingerprinting** for a starter set (Home Assistant, Postgres, Ollama, Grafana, Portainer…).
- **Package Registry** v1: local + GitHub/OCI package index; config schemas.
- **Smart Recipes** v1: match→install→config rules; a starter recipe pack.
- **Auto Installer** + **Runtime (Docker)**: pull, create sandboxed container, template config, register instance.
- **Reconciler** wires it together: discovered→matched→installed→routed, idempotently.
- **Inventory** persistence + first dashboard-less status API.
**Exit criteria:** start a Postgres container on the LAN → within one reconcile cycle, `postgres.query` is live at the Gateway, no human action.
Modules: [Discovery](./modules/01-discovery-engine.md), [Registry](./modules/02-package-registry.md), [Recipes](./modules/15-smart-recipes.md), [Installer](./modules/03-auto-installer.md), [Infra Discovery](./modules/16-infrastructure-discovery.md).
---
## Phase 3 — Secure (data plane hardening)
**Goal:** multi-agent, least-privilege access. An agent sees only what its role allows.
- **Authentication:** API keys first; then OIDC/OAuth; JWT sessions. Pluggable backends (LDAP/AD/SAML later).
- **RBAC:** roles → allowed namespaces/tools; enforced in the Router at both `tools/list` and `tools/call`.
- **AI Agent Profiles:** identity → roles → resolved visible tool set.
- **Secrets Manager:** encrypted-at-rest vault; envelope encryption; runtime injection into MCP containers; never returned to agents or logged.
- **Audit:** every tool call + mutating action recorded with principal/target/outcome.
- **Security baseline:** container sandboxing defaults, TLS termination, rate limiting, signed-package verification.
**Exit criteria:** two agents with different roles connect; each sees a different tool set; a secret-backed MCP works without the secret ever appearing in any agent-visible payload or log.
Modules: [Auth](./modules/06-authentication.md), [RBAC](./modules/08-rbac.md), [Profiles](./modules/14-agent-profiles.md), [Secrets](./modules/07-secrets-manager.md), [Security](./modules/19-security.md).
---
## Phase 4 — Operate (day-2)
**Goal:** it runs unattended and you can see what it's doing.
- **Health Monitoring:** liveness/readiness per instance; auto-restart; backoff; quarantine.
- **Update Manager:** watch image/release/registry updates; auto or approval-gated; version pinning; rollback.
- **Metrics:** Prometheus exposition (requests, latency, errors, token usage, discovery counts, instance states); Grafana dashboards.
- **Notifications:** Discord/Slack/Email/ntfy/Telegram/Pushover/Signal on discovery, updates, failures, offline, security alerts.
- **Web Dashboard:** React UI embedded in the binary — services, instances, containers, logs, health, updates, secrets, users, roles, discovery, agents.
**Exit criteria:** kill a managed MCP container → it's auto-restarted, a Discord alert fires, and the dashboard reflects the transition in real time.
Modules: [Health](./modules/09-health-monitoring.md), [Updates](./modules/10-update-manager.md), [Metrics](./modules/18-metrics.md), [Notifications](./modules/13-notifications.md), [Dashboard](./modules/12-web-dashboard.md).
---
## Phase 5 — Extend & scale
**Goal:** third parties extend Nexus; it runs in production topologies.
- **Plugin System:** stable interfaces + out-of-process plugin transport for discovery, fingerprinters, installers, package providers, notifiers, auth, dashboards, tool transformers.
- **Service Graph:** live dependency graph (agent→gateway→MCP→service→device) with interactive visualization.
- **Runtime adapters:** containerd, Kubernetes operator, LXC.
- **HA:** multi-replica Nexus, shared Postgres, leader-elected reconciler.
- **Community registry & recipe governance:** trust/signing model for shared recipes and packages.
**Exit criteria:** a community-authored discovery plugin + recipe installs a new MCP with no core changes; Nexus runs 3-replica HA against Postgres.
Modules: [Plugins](./modules/11-plugin-system.md), [Service Graph](./modules/17-service-graph.md), [Future Vision](./modules/20-future-vision.md).
---
## What is intentionally *not* early
- Full auth backend matrix (SAML/AD) — API keys + OIDC cover the 90% first.
- K8s runtime — Docker proves the model; K8s is an adapter, not a rewrite.
- Multi-tenancy above RBAC — deferred until there's demand (see Architecture §12).
- Windows/Hyper-V discovery — Linux/Docker homelab is the beachhead.
## Sequencing rationale
The riskiest, most differentiating claim is **"discover a service and its MCP just appears behind one endpoint."** Phases 1–2 prove exactly that as fast as possible. Everything after widens (more discovery methods, more recipes), hardens (security, ops), or scales (plugins, HA) — none of which is worth building before the core loop is real.

View File

@@ -0,0 +1,93 @@
# Discovery Engine
> Module 01 · Plane: Control · Roadmap phase: 2
> Part of [MCP Nexus architecture](../ARCHITECTURE.md).
## Purpose
Continuously scan the network and its runtimes for **application services** (Home Assistant, Postgres, Grafana, Ollama, …), fingerprint each into a typed `DiscoveredResource` with a confidence score, and emit it onto the event bus so the reconciler can match it against [Recipes](15-smart-recipes.md). This is the `observe` stage of the reconciliation loop (Architecture §6).
## Responsibilities
- Own a set of pluggable **discovery methods** that enumerate endpoints: mDNS/Bonjour/DNS-SD/Zeroconf, SSDP/UPnP, ICMP/ARP sweeps, HTTP/HTTPS probing, TLS certificate inspection, Docker API (containers + labels + Compose project metadata), Kubernetes API, Proxmox/VMware/LXC/Hyper-V inventories, SSH host-key fingerprints, SNMP, MQTT topic discovery, and Home Assistant discovery.
- Own a set of pluggable **fingerprinters** that turn raw signals (open ports, HTTP titles/headers, TLS SANs, banners, container image names, labels) into a concrete `type` + `version` + `capabilities`. Starter type set: Home Assistant, UniFi, Immich, Paperless, Grafana, Prometheus, Jellyfin, Ollama, Postgres, MySQL/MariaDB, Redis, MinIO, Frigate, Synology, TrueNAS, Pi-hole, AdGuard, Traefik, Portainer, Nginx Proxy Manager, OpenWebUI, OpenAI-compatible APIs, LM Studio, vLLM, llama.cpp, ComfyUI, Automatic1111.
- Assign every discovered resource a stable `uuid` (derived from a natural key so re-discovery is idempotent, not a new row) and a **confidence score** aggregated across the methods/fingerprinters that observed it.
- Deduplicate and merge multi-source sightings of the same service (e.g. mDNS + Docker label + HTTP probe) into one `DiscoveredResource`, updating `first_seen`/`last_seen`.
- Emit `DiscoveredResource` create/update/lost events onto the event bus and upsert them into Inventory.
- Run on a schedule (periodic resync) and edge-triggered (e.g. a Docker event stream) per tenet #2.
## Non-goals
- **Does not discover compute substrates/hosts** to deploy onto — that is [Infrastructure Discovery](16-infrastructure-discovery.md) (Module 16). Module 1 finds *what services exist*; Module 16 finds *where MCP containers can run*.
- **Does not decide what to install.** Matching a fingerprint to a package/config is [Smart Recipes](15-smart-recipes.md) + the reconciler.
- **Does not install, configure, or run** anything — that is the [Auto Installer](03-auto-installer.md) and Runtime.
- **Does not define plugin transport.** It consumes the interfaces owned by the [Plugin System](11-plugin-system.md); out-of-process discovery plugins land in Phase 5.
- **Does not store credentials.** Method configs that need creds reference [Secrets](07-secrets-manager.md) by ref.
## Interfaces
```go
// A discovery method enumerates candidate endpoints on the network/runtime.
type Method interface {
Name() string // "mdns", "docker", "http-probe", ...
Scan(ctx context.Context, scope Scope) (<-chan Observation, error)
}
// A fingerprinter inspects observations and proposes a typed identity + confidence.
type Fingerprinter interface {
Name() string
Match(ctx context.Context, o Observation) (Fingerprint, bool)
}
// Fingerprint is a scored, typed claim about an observation.
type Fingerprint struct {
Type string // "home-assistant", "postgres", ...
Version string // best-effort
Capabilities []string // e.g. ["rest-api","websocket"]
Confidence float64 // 0.0–1.0 for THIS signal
Evidence map[string]string // http_title, tls_san, image, banner...
}
// The engine fans methods → fingerprinters → merged resources.
type Engine interface {
Register(m Method) error
RegisterFingerprinter(f Fingerprinter) error
Run(ctx context.Context) error // periodic + edge-triggered
Rescan(ctx context.Context, scope Scope) error
}
```
Internal HTTP (control API, not agent-facing; guarded by RBAC):
- `GET /api/v1/discovery/resources` — list `DiscoveredResource`s (filter by type/confidence/source).
- `POST /api/v1/discovery/rescan` — trigger an on-demand scan (optionally scoped).
- `GET /api/v1/discovery/methods` — list registered methods + last-run status.
## Data
- **Writes** `DiscoveredResource` (§7) to the **Inventory** store (§8) — `{uuid, type, version, ip, hostname, ports, capabilities, health, confidence, source, first_seen, last_seen}`.
- **Reads** method configuration from the **Config** store and any method credentials from **Secrets** (by ref only).
- Module-local persistence: a small per-method scan cursor/cache (last-seen sets, backoff timers) to make rescans cheap and dedup stable across restarts.
## Dependencies
- Emits to the reconciler via the event bus; consumed by [Smart Recipes](15-smart-recipes.md) and the [Auto Installer](03-auto-installer.md).
- Method/fingerprinter interfaces are provided by the [Plugin System](11-plugin-system.md).
- Shares inventory + event bus with [Infrastructure Discovery](16-infrastructure-discovery.md).
- Discovered types feed the dependency view in [Service Graph](17-service-graph.md).
- Credentials via [Secrets](07-secrets-manager.md); scan scope may be constrained by [RBAC](08-rbac.md).
## Failure modes & handling
| Failure | Behavior |
|---|---|
| A method errors or times out | Isolated per-method; other methods continue. Error logged + surfaced in method status; exponential backoff before retry. |
| Aggressive scan (ARP/ICMP sweep) trips IDS / rate limits | Sweeps are opt-in, rate-limited, and scoped to configured CIDRs; passive methods (mDNS/Docker) are default-on. |
| Two methods disagree on `type` | Highest-confidence fingerprint wins; conflicting evidence lowers aggregate confidence and is retained in `Evidence`. |
| Confidence below threshold | Resource is recorded but **not** auto-acted-on; flagged for human-in-the-loop review (tenet #6). |
| Service disappears | Marked `lost` after a grace period (not deleted); emits a `resource lost` event so the reconciler can decide on the bound `MCPInstance`. |
| Duplicate sightings | Merged by natural key into one `uuid`; never spawns duplicate instances. |
## Security notes
Honors Architecture §9: active scanning is **least-privilege and opt-in** — passive/local methods default on, network sweeps require explicit CIDR scope. Method credentials come from [Secrets](07-secrets-manager.md) and are never logged or written into `DiscoveredResource`. Every scan and every discovery event is audited with source + principal. Discovery observes only; it holds no write access to runtimes. TLS inspection reads certificates without trusting them for auth.
## Open questions
- Confidence math: fixed weighted sum per signal, or a learned/Bayesian combiner tunable per environment?
- How aggressive may default network sweeps be before they need explicit opt-in in a homelab?
- Should `lost` grace periods be per-type (a laptop MCP vs a rack Postgres differ)?
- Do we expose fingerprint `Evidence` to the dashboard for debugging, and does it risk leaking internal topology?
## Milestone
Delivered in **Phase 2** (Discover → install). Thin slice that lands first: **Docker API + mDNS/DNS-SD + HTTP(S) probing** methods with fingerprinters for the starter set (Home Assistant, Postgres, Ollama, Grafana, Portainer), emitting `DiscoveredResource` with confidence onto the event bus. Exit proof (with Recipes + Installer): start a Postgres container on the LAN → it is discovered, fingerprinted, and `postgres.query` appears at the Gateway within one reconcile cycle.

View File

@@ -0,0 +1,186 @@
# MCP Package Registry
> Module 02 · Plane: Control · Roadmap phase: 2
> Part of [MCP Nexus architecture](../ARCHITECTURE.md).
## Purpose
The Package Registry is the catalog that answers one question for the rest of the
control plane: *"Given a service type, which MCP server should I install, from
where, at what version, and configured how?"* It maintains a searchable, locally
cached index of `Package` objects (Architecture §7) synced from one or more
remote indexes plus operator-authored local packages. It is the source of truth
for artifact provenance (source, digest, signature) and for the `config_schema`
that [Recipes](15-smart-recipes.md) bind to at install time.
The Registry does not run anything. It resolves and describes. The
[Auto Installer](03-auto-installer.md) consumes its resolutions; a `Package`
never becomes an `MCPInstance` inside this module.
## Responsibilities
- Maintain the local `Package` catalog in the **Registry** store (Architecture §8),
co-resident with Recipes.
- **Sync** from multiple remote indexes (a Nexus index is a signed JSON/HTTP
manifest listing packages) and merge them with local packages, with a
deterministic precedence order on collision.
- Provide **resolution**: `service type → recommended Package` (e.g. Home
Assistant → `homeassistant-mcp`).
- Provide **version selection**: given a `Package` and a constraint, return a
concrete version **pinned by digest**.
- Verify **signatures** on packages and index manifests before a package is
eligible for install (tenet #5, Architecture §9).
- Expose the `config_schema` and config **template reference** each package
carries, so Recipes can render config against it.
- Full-text and faceted **search** over name, service type, source, and tags for
the Dashboard and API.
## Non-goals
- Pulling images or creating containers — that is the [Installer](03-auto-installer.md)
via the Runtime.
- Deciding *whether* to install (matching a `DiscoveredResource` to a package) —
that is [Smart Recipes](15-smart-recipes.md) and the reconciler.
- Storing secrets or rendered config — templates are declarative; values come
from [Secrets](07-secrets-manager.md) at install time.
- Being an OCI registry itself. Nexus reads from GitHub / OCI / Docker Hub /
self-hosted repos; it does not host artifacts.
## Interfaces
```go
// Source of a package artifact (mirrors Package.source in Architecture §7).
type Source string
const (
SourceGitHub Source = "github"
SourceOCI Source = "oci"
SourceDockerHub Source = "dockerhub"
SourceLocal Source = "local"
)
// A concrete, install-ready version pinned by digest.
type PackageVersion struct {
Version string // semver-ish tag, e.g. "1.4.2"
Ref string // "ghcr.io/project/homeassistant-mcp:1.4.2"
Digest string // "sha256:..." — the install pin
Signature Signature
}
type Package struct {
Name string // "homeassistant-mcp"
ServiceType string // "home-assistant" — resolution key
Source Source
Versions []PackageVersion
ConfigSchema json.RawMessage // JSON Schema for config
ConfigTemplateRef string // e.g. "homeassistant.yaml"
Signature Signature
}
type Registry interface {
// Resolution: service type -> recommended package.
Resolve(ctx context.Context, serviceType string) (Package, error)
// Version selection against a constraint ("", "latest", "~1.4", "1.4.2").
SelectVersion(ctx context.Context, pkg string, constraint string) (PackageVersion, error)
Get(ctx context.Context, name string) (Package, error)
Search(ctx context.Context, q Query) ([]Package, error)
}
// Everything-is-a-plugin (tenet #3): remote index providers are swappable.
type IndexProvider interface {
Name() string
Fetch(ctx context.Context) ([]Package, IndexMeta, error) // signed manifest
}
type SyncManager interface {
Sync(ctx context.Context) (SyncReport, error) // merge all providers + local
AddIndex(url string, opts IndexOpts) error
}
```
HTTP/API surface (control-plane API, behind Auth):
- `GET /api/v1/packages?service_type=&source=&q=` — search/list.
- `GET /api/v1/packages/{name}` — full package incl. versions + schema.
- `POST /api/v1/packages` — register/upsert a **local** package.
- `POST /api/v1/registry/sync` — trigger a sync of all configured indexes.
- `GET /api/v1/registry/indexes` — configured remote indexes + last sync state.
## Data
- **Reads/writes:** `Package` objects in the **Registry** store (Architecture §8),
alongside `Recipe`s (which reference packages by name + `ConfigTemplateRef`).
- **Module-local persistence:** `registry_index` rows (index URL, public key,
last-synced digest, ETag/cursor), `package_version` rows (digest, ref,
signature status), and a cached copy of each fetched index manifest for
offline/deterministic resolution.
- Config templates referenced by `ConfigTemplateRef` are stored as opaque blobs
in the Registry store and rendered later by the Installer/Recipe engine.
Resolution flow: a `DiscoveredResource.type` (Architecture §7) is mapped by a
Recipe to a `serviceType`; `Resolve` returns the recommended `Package`;
`SelectVersion` pins it by digest. Example entry:
```
service_type: home-assistant
name: homeassistant-mcp
source: oci
ref: ghcr.io/project/homeassistant-mcp
versions: [1.4.2 @ sha256:…, 1.4.1 @ sha256:…]
config_template_ref: homeassistant.yaml
```
## Dependencies
- [Smart Recipes](15-smart-recipes.md) — consumes packages + config templates.
- [Auto Installer](03-auto-installer.md) — consumes resolved, digest-pinned versions.
- [Secrets Manager](07-secrets-manager.md) — config templates declare secret refs
filled at install time, not here.
- [Security](19-security.md) — signature/trust model for indexes and packages.
- [Plugin System](11-plugin-system.md) — `IndexProvider` is a plugin point.
- [Update Manager](10-update-manager.md) — watches index updates as one update source.
## Failure modes & handling
- **Remote index unreachable:** sync is best-effort; last good cached manifest
continues to serve `Resolve`/`SelectVersion`. Sync failures are surfaced via
[Notifications](13-notifications.md), never block the control loop.
- **Signature verification fails:** the offending package/version is marked
`untrusted` and excluded from resolution; it is never returned as installable.
- **Index collision (same package from two indexes):** deterministic precedence
(local > pinned-trusted index > community index); the shadowed entry is logged.
- **No package for a service type:** `Resolve` returns a typed `ErrNoPackage`;
the reconciler leaves the resource un-installed and low-confidence handling
applies (tenet #6).
- **Missing/invalid `config_schema`:** package is listed but flagged
`unschematized`; Recipes referencing it fail validation early, not at runtime.
## Security notes
- Every remote index manifest and every package is **signature-verified** before
eligibility (tenet #5, Architecture §9); trust anchors (public keys) are
configured per index.
- Versions are always **pinned by digest**; tags are advisory only. The digest is
what the Installer pulls.
- The Registry stores no secrets and renders no secret values; templates carry
only *references*.
- Mutating API calls (register local package, add index, sync) are audited
(Architecture §9) and gated by [RBAC](08-rbac.md).
## Open questions
- Community index governance and the trust/signing model for shared packages
(mirrors Architecture §12 "Recipe distribution").
- Do we support multiple recommended packages per service type with a ranking,
or strictly one canonical recommendation?
- Should `SelectVersion` honor per-package update policy hints, or is that owned
entirely by the [Update Manager](10-update-manager.md)?
- Config-template format: single YAML template vs. a small templating dialect
shared with Recipes.
## Milestone
Phase 2 — *Discover → install*. Registry v1 ships with local + GitHub/OCI package
indexes, digest-pinned versions, config schemas, and a starter package set so the
[Installer](03-auto-installer.md) can resolve (e.g.) Postgres → `postgres-mcp`
during the Phase 2 exit-criteria loop.

View File

@@ -0,0 +1,175 @@
# Auto Installer
> Module 03 · Plane: Control · Roadmap phase: 2
> Part of [MCP Nexus architecture](../ARCHITECTURE.md).
## Purpose
The Auto Installer is the actor that turns a matched triple —
`(DiscoveredResource + Package + Recipe)` — into a running, registered
`MCPInstance` (Architecture §7) with **no manual work**. It is the `act` step of
the reconciliation loop (Architecture §6) for new installs: given a plan from the
reconciler, it pulls the pinned image, creates a **sandboxed** container through
the [Runtime](16-infrastructure-discovery.md) abstraction, renders config,
injects secrets at runtime, starts, health-gates, and registers the instance so
the [Router/Gateway](04-gateway.md) begins exposing its tools.
Per tenet #2 it is *reconcile, not script*: the Installer is invoked by the
reconciler with a desired spec and must be **idempotent** — re-running with the
same inputs yields no duplicate containers and no duplicate instances.
## Responsibilities
- Accept an install `Plan` (bound resource + resolved, digest-pinned package
version + recipe) and drive it to a healthy, registered `MCPInstance`.
- **Registry lookup / resolution:** obtain the concrete `PackageVersion` (digest)
from the [Registry](02-package-registry.md) (the reconciler usually pre-resolves;
the Installer re-validates the pin).
- **Pull** the image by digest through the target Runtime.
- **Create** a sandboxed container: least privilege, read-only rootfs where
possible, no host network unless the recipe demands it (Architecture §9).
- **Render** the recipe's config template against the package `config_schema`,
injecting [Secrets](07-secrets-manager.md) at runtime (env/file mount) — never
baking secret values into the image, the stored spec, or logs.
- **Start**, then **health-gate**: block registration until the instance passes
an initial readiness check (handed off to [Health](09-health-monitoring.md)).
- **Register** the `MCPInstance` in the **Inventory** store and emit an event so
the Router refreshes its upstream set.
- Be **idempotent** and support **rollback** on any failed step.
- Target multiple **Runtimes** (Docker first; K8s/containerd later) discovered by
[Infrastructure Discovery](16-infrastructure-discovery.md).
## Non-goals
- Deciding *what* to install or *whether* to (matching + policy) — that is the
reconciler + [Recipes](15-smart-recipes.md).
- Choosing the version or verifying signatures — that is the
[Registry](02-package-registry.md) (the Installer trusts the digest-pinned,
verified version handed to it, and re-checks the pin).
- Ongoing liveness/restart — owned by [Health Monitoring](09-health-monitoring.md).
- Applying version upgrades — a re-install driven by the
[Update Manager](10-update-manager.md) reuses this module.
## Interfaces
```go
type Plan struct {
Resource DiscoveredResource // bound_resource_uuid target
Version PackageVersion // digest-pinned (Registry)
Recipe Recipe // config template + runtime constraints
}
type Installer interface {
// Idempotent: same Plan -> same MCPInstance, no duplicate container.
Install(ctx context.Context, p Plan) (MCPInstance, error)
Rollback(ctx context.Context, instanceID string) error
}
// Runtime abstraction (tenet #3). Docker first; K8s/containerd adapters later,
// selected per-target from Infrastructure Discovery (Module 16).
type Runtime interface {
Name() string // "docker", "k8s", ...
Pull(ctx context.Context, ref, digest string) error
Create(ctx context.Context, spec ContainerSpec) (ContainerRef, error)
Start(ctx context.Context, ref ContainerRef) error
Stop(ctx context.Context, ref ContainerRef) error
Remove(ctx context.Context, ref ContainerRef) error
Inspect(ctx context.Context, ref ContainerRef) (ContainerStatus, error)
}
// Sandbox defaults come from Security (Module 19); recipes may widen with reason.
type ContainerSpec struct {
Ref string
Digest string
Labels map[string]string // includes nexus.instance-id for idempotency
Env []EnvVar // secret refs resolved at start
Mounts []Mount // incl. secret file mounts (tmpfs)
ReadOnlyRoot bool
NoNewPrivs bool
Network NetworkMode
}
```
Idempotency key: the Installer derives a stable `instance-id` from
`(package, version-digest, bound_resource_uuid, recipe-hash)` and stamps it as a
container **label** and the `MCPInstance.id`. Before creating anything it queries
the Runtime for a container carrying that label and adopts it instead of
re-creating.
No public HTTP surface of its own; it is invoked in-process by the reconciler.
Install progress is observable via Inventory state transitions and
[Metrics](18-metrics.md).
## Data
- **Writes:** `MCPInstance` `{id, package, version, config, container_ref, state,
health, bound_resource_uuid}` into the **Inventory** store (Architecture §8),
transitioning `state` through `Pulling → Creating → Starting → HealthGating →
Running` (or `Failed`).
- **Reads:** `Package`/`PackageVersion` and config template from the
[Registry](02-package-registry.md); `Recipe` for constraints; `Secret`
references resolved via [Secrets](07-secrets-manager.md) at start time.
- **Module-local persistence:** an install-attempt journal (step, timestamp,
outcome) used for rollback and audit; the previous known-good spec is retained
to support Update Manager rollback.
- The stored `MCPInstance.config` is the **rendered template with secret values
redacted to references** — secret material is never persisted here.
## Dependencies
- [MCP Package Registry](02-package-registry.md) — resolved, pinned version + template.
- [Smart Recipes](15-smart-recipes.md) — config template + runtime constraints.
- [Secrets Manager](07-secrets-manager.md) — runtime secret injection.
- [Infrastructure Discovery](16-infrastructure-discovery.md) — discovers Runtimes/targets.
- [Health Monitoring](09-health-monitoring.md) — the health gate + post-install liveness.
- [Gateway](04-gateway.md) / [Dynamic Tool Registry](05-dynamic-tool-registry.md) —
consume the registered instance's tools.
- [Security](19-security.md) — sandbox defaults.
- [Update Manager](10-update-manager.md) — re-invokes Install for version upgrades.
## Failure modes & handling
- **Pull fails** (network, missing digest): mark `Failed`, no container created,
reconciler retries with backoff; nothing to roll back.
- **Create/start fails:** run `Rollback` — stop+remove any partial container,
release reserved secret mounts, revert Inventory to pre-install state.
- **Health gate times out:** treat as failed install; roll back the new container
so the Gateway never sees a broken instance. On an *update* re-install, restore
the previous known-good spec (see [Update Manager](10-update-manager.md)).
- **Duplicate on re-run (idempotency):** the label lookup adopts the existing
container; no second container, no second `MCPInstance`.
- **Crash mid-install:** the install-attempt journal + label lookup let the next
reconcile cycle resume or clean up orphaned containers (garbage-collect
containers labelled with an `instance-id` absent from Inventory).
- **Runtime unavailable** (e.g. Docker socket down): install is deferred, not
errored permanently; data plane is unaffected (Architecture §5).
## Security notes
- Containers are **sandboxed by default** (Architecture §9): least privilege,
`NoNewPrivs`, read-only rootfs where possible, no host network unless the
recipe explicitly requires it with a recorded reason.
- **Secrets injected at runtime only** — env vars or tmpfs file mounts populated
at `Start`; never in the image, the stored `MCPInstance.config`, or logs.
- Only **digest-pinned, signature-verified** versions are installed; the pin is
re-validated against the Registry before pull.
- Every install/rollback is an audited mutating control-plane action
(Architecture §9) with principal (reconciler/user), target, and outcome.
## Open questions
- How much Docker installer semantics map cleanly to a K8s operator vs. a
separate controller (Architecture §12 "K8s runtime parity")?
- Concurrency/locking model when two reconcile cycles race on the same
`instance-id` — advisory lock in Inventory vs. Runtime-level create idempotency.
- Orphan GC cadence and safety margin before removing an unrecognized
Nexus-labelled container.
- Whether the health gate criteria live in the Recipe or default from the Package
`config_schema`.
## Milestone
Phase 2 — *Discover → install*. Installer v1 + Docker Runtime satisfies the phase
exit criteria: a Postgres container appears on the LAN and, within one reconcile
cycle, `postgres.query` is live at the [Gateway](04-gateway.md) with no human
action — idempotently, with rollback on failure.

169
docs/modules/04-gateway.md Normal file
View File

@@ -0,0 +1,169 @@
# Gateway
> Module 04 · Plane: Data · Roadmap phase: 1
> Part of [MCP Nexus architecture](../ARCHITECTURE.md).
## Purpose
The Gateway is **the single MCP endpoint every AI agent connects to** —
`https://mcp.company.lan`. It embodies tenet #1 ("one endpoint, many
agents"): to an agent the Gateway *is* an MCP server; upstream, together with
the **Router** (the "MCP Router + Aggregator" of Architecture §4), it acts as
an MCP client to every managed
[`MCPInstance`](../ARCHITECTURE.md). It dynamically **aggregates** all healthy
upstream MCP servers behind one address and **namespaces** their tools so
`homeassistant.turn_on`, `postgres.query`, and `github.create_issue` never
collide.
The Gateway owns the *edge and transport*. It does not own tool ownership or
dispatch (that is the Router) and it does not own the tool
catalog (that is the [Dynamic Tool Registry](05-dynamic-tool-registry.md)).
Keeping these three concerns separate is deliberate: the Gateway can be
restarted, load-balanced, and hardened without touching routing logic or the
catalog.
## Responsibilities
- Terminate TLS and accept agent connections over **Streamable HTTP** (the
current MCP HTTP transport) speaking **JSON-RPC 2.0**.
- Run the MCP server side of the `initialize` handshake: negotiate protocol
version, advertise server capabilities (`tools`, and later `resources` /
`prompts`), and open a session.
- Manage the per-session request lifecycle: parse JSON-RPC frames, correlate
ids, enforce timeouts, and stream partial/chunked results back to the agent.
- Delegate every request to [Auth](06-authentication.md) (who is calling?) and
[RBAC](08-rbac.md) (what may they see/call?) *before* it reaches the Router.
- Hand `tools/list` and `tools/call` to the Router, which resolves the owning
instance from the Dynamic Tool Registry and dispatches over the upstream
client pool.
- Maintain the **connection pool** to upstream MCP servers (stdio and
Streamable HTTP), reusing warm connections across agent requests.
- Ship the `nexus connect` **stdio↔HTTP bridge** so local agents (Claude
Desktop, Cursor) that only speak stdio can attach to the HTTP endpoint.
- Emit per-request audit and metrics events.
## Non-goals
- **Tool ownership / dispatch decisions** — owned by the Router.
- **The tool catalog and its freshness** — owned by the Dynamic Tool Registry.
- **AuthN/AuthZ policy** — the Gateway *calls* Auth/RBAC; it stores no policy.
- **Installing, healing, or updating instances** — control-plane concern; the
Gateway only consumes the healthy upstream set.
- **Secret storage or injection** — secrets reach MCP containers via the
[Secrets Manager](07-secrets-manager.md), never through the Gateway.
## Interfaces
```go
// Server is the agent-facing MCP server. One per Nexus process.
type Server interface {
// Serve binds the Streamable HTTP listener (TLS terminated here).
Serve(ctx context.Context, addr string, tls *tls.Config) error
// Handle processes one JSON-RPC request within an authenticated session
// and writes the response (possibly streamed) to w.
Handle(ctx context.Context, s *Session, req *jsonrpc.Request, w StreamWriter) error
}
// Session is a live agent connection after a successful initialize.
type Session struct {
ID string
Principal auth.Principal // resolved by Auth
Roles []rbac.Role // resolved by RBAC → visible tool set
Protocol string // negotiated MCP version
Caps ClientCaps
}
// StreamWriter delivers incremental JSON-RPC results / progress notifications.
type StreamWriter interface {
WriteChunk(any) error
Close(err error) error
}
// UpstreamPool hands the Router warm MCP client connections per instance.
type UpstreamPool interface {
Client(ctx context.Context, instanceID string) (mcpclient.Conn, release func(), error)
Drain(instanceID string) // called when Health marks an instance gone
}
```
HTTP / MCP endpoints exposed at the edge:
| Method / path | Purpose |
|---|---|
| `POST /mcp` | Streamable HTTP MCP endpoint; carries JSON-RPC `initialize`, `tools/list`, `tools/call`, notifications. |
| `GET /mcp` | Server→client stream channel (SSE-style) for long-lived sessions. |
| `GET /healthz` | Liveness of the edge (does not gate on upstream health). |
`nexus connect --url https://mcp.company.lan --token …` runs a local process
that presents stdio to the agent and proxies frames to `POST/GET /mcp`.
## Data
The Gateway is **mostly stateless** — it holds only in-memory session and pool
state, so replicas behind a load balancer are interchangeable (HA topology,
Architecture §10).
- **Session table (in-memory):** `sessionID → Session`, TTL-expired.
- **Connection pool (in-memory):** per-`MCPInstance` warm `mcpclient.Conn`s,
keyed by `instanceID`, with idle eviction and per-instance max.
- **Persisted:** nothing of its own. It *reads* the healthy instance set and
the RBAC-filtered tool view via the Registry/Router, and *writes* audit +
metrics events to the shared Audit/Metrics stores.
## Dependencies
- [Authentication](06-authentication.md) — resolves the calling principal.
- [RBAC](08-rbac.md) — resolves the visible/callable tool set per session.
- [Dynamic Tool Registry](05-dynamic-tool-registry.md) — source of the live
`{namespace}.{tool}` catalog and tool→instance mapping.
- [AI Agent Profiles](14-agent-profiles.md) — maps agent identity to roles.
- [Health Monitoring](09-health-monitoring.md) — signals which instances are
poolable; drives `UpstreamPool.Drain`.
- [Secrets Manager](07-secrets-manager.md) — ensures no secret transits the
edge in an agent-visible payload.
- [Metrics](18-metrics.md) / [Security](19-security.md) — audit, TLS, limits.
## Failure modes & handling
- **Upstream instance down mid-call:** the pool returns a transport error; the
Router surfaces a JSON-RPC error to the agent. The Gateway never hangs on a
dead upstream — every dispatch is deadline-bounded.
- **Upstream slow:** per-request deadline + circuit breaker per instance;
streamed progress keeps the agent connection alive until the deadline.
- **Auth/RBAC unavailable:** fail **closed** — reject with a JSON-RPC error;
never fall back to unauthenticated access (tenet #5).
- **Pool exhaustion:** bounded per-instance connections + a wait queue with a
timeout; over-limit calls get a retryable error, not unbounded fan-out.
- **Malformed JSON-RPC:** respond with the spec error code; drop the frame,
keep the session.
- **Bridge disconnect (`nexus connect`):** local process retries with backoff;
the HTTP session is TTL-reaped if the bridge does not resume.
## Security notes
- **TLS terminates here** (Architecture §9); internal calls to Router/Auth are
over localhost/socket.
- The Gateway enforces that a session may only ever see and invoke its
RBAC-visible tool set — requests are scoped *before* the Router dispatches,
so an agent cannot call a tool outside its role even by guessing its name.
- Per-session and per-principal **rate limiting** at the edge.
- No secret material is ever echoed back to an agent; error messages from
upstreams are sanitized.
- Every `tools/call` is audited with principal, resolved instance, and outcome.
## Open questions
- Session affinity vs. stateless replicas: do we need sticky sessions for
streamed calls behind the HA load balancer, or is a shared session store
sufficient?
- Should the `nexus connect` bridge be a subcommand of the main binary or a
separately distributed micro-binary (relates to Architecture §12)?
- Streamable HTTP resumability: how far do we implement MCP resumable streams
vs. requiring the agent to re-issue on reconnect?
## Milestone
Phase 1 (Walking skeleton). Ships with the MCP server edge, upstream client
pool, and `nexus connect`. **Exit:** register a filesystem MCP + one HTTP MCP
in YAML, connect Claude, see `filesystem.*` and `other.*` tools, and call one
successfully — the core value proposition proven end-to-end.

View File

@@ -0,0 +1,164 @@
# Dynamic Tool Registry
> Module 05 · Plane: Data · Roadmap phase: 1
> Part of [MCP Nexus architecture](../ARCHITECTURE.md).
## Purpose
The Dynamic Tool Registry is the **live, continuously-updated catalog of every
tool exposed by every healthy [`MCPInstance`](../ARCHITECTURE.md)**, each
namespaced as `{namespace}.{tool}` (Architecture §7). It is the authoritative
answer to two questions on the data-plane hot path:
1. *"What tools exist?"* — it serves `tools/list` to agents, filtered to the
caller's RBAC-visible set.
2. *"Who owns this tool?"* — it maintains the `tool → owning instance` mapping
the Router uses to dispatch `tools/call`.
It is **live** because upstream MCP servers appear, disappear, and change their
tool sets at runtime. The Registry reacts to instances coming and going (from
the [Installer](03-auto-installer.md) / [Health](09-health-monitoring.md)) and
to upstream `tools/list` results and `notifications/tools/list_changed`
notifications, keeping the aggregated catalog current without a restart.
## Responsibilities
- Aggregate the tool lists of all **healthy** instances into one namespaced
catalog; drop tools of unhealthy/absent instances immediately.
- Assign and enforce the `{namespace}.{tool}` naming and **collision rules**
(see [Data](#data)).
- Maintain the `{namespace}.{tool} → instanceID` dispatch map consumed by the
Router.
- Serve `tools/list` results and the `GET /tools` view, always filtered by the
caller's RBAC-resolved visibility.
- Subscribe to lifecycle events (instance added/removed, health transitions)
and to upstream `list_changed` notifications; refresh incrementally.
- Cache upstream tool schemas (input/output JSON Schema) so `tools/list` does
not fan out to every upstream on each agent call.
- Emit change events so the Gateway can forward `list_changed` to connected
agents whose visible set changed.
## Non-goals
- **Transport / sessions / the agent edge** — owned by the
[Gateway](04-gateway.md).
- **Dispatching `tools/call` and streaming results** — owned by the Router; the
Registry only supplies the ownership mapping.
- **Deciding what an agent may see** — the Registry *applies* an RBAC filter it
is given; policy lives in [RBAC](08-rbac.md).
- **Installing or health-checking instances** — control-plane modules feed the
Registry; it does not manage instance lifecycle.
## Interfaces
```go
// Registry is the live namespaced tool catalog. Concurrency-safe; reads are
// hot-path, writes are event-driven.
type Registry interface {
// List returns the tools visible to a principal, already RBAC-filtered.
List(ctx context.Context, view rbac.ToolView) ([]Tool, error)
// Resolve maps a namespaced tool name to its owning instance for dispatch.
Resolve(name string) (instanceID string, t Tool, ok bool)
// Subscribe streams catalog change events (tool added/removed/schema-changed).
Subscribe(ctx context.Context) (<-chan ChangeEvent, func())
}
// Ingest is the write side, driven by lifecycle + upstream events.
type Ingest interface {
// UpsertInstance (re)loads an instance's tools under its namespace.
UpsertInstance(ctx context.Context, inst MCPInstance, tools []UpstreamTool) error
// RemoveInstance drops all tools owned by an instance.
RemoveInstance(ctx context.Context, instanceID string) error
// OnListChanged handles an upstream notifications/tools/list_changed.
OnListChanged(ctx context.Context, instanceID string) error
}
type Tool struct {
Name string // "{namespace}.{tool}", e.g. "postgres.query"
Namespace string
Bare string // upstream tool name, e.g. "query"
InstanceID string
InputSchema json.RawMessage
Title, Description string
}
```
HTTP / MCP surface:
| Method / path | Purpose |
|---|---|
| MCP `tools/list` | Served via the Gateway; RBAC-filtered namespaced catalog. |
| MCP `notifications/tools/list_changed` | Forwarded to agents when their visible set changes. |
| `GET /tools` | Dashboard/API view of the full catalog with owning instance, namespace, health, and schema (control-plane, RBAC-guarded). |
## Data
- **Catalog (in-memory, source of truth at runtime):** `map[name]Tool`, plus a
`namespace → instanceID` index and reverse `instanceID → []name`.
- **Namespacing rule:** the namespace is the instance's stable, human-readable
name (from its Recipe/config, e.g. `homeassistant`, `postgres`, `github`).
The fully-qualified tool name is `{namespace}.{tool}`.
- **Collision rules:**
- Two instances may never share a namespace — enforced at instance
registration; a duplicate namespace is a config/reconcile error surfaced to
the operator, not silently merged.
- Within a namespace, upstream tool names are already unique per MCP server,
so `{namespace}.{tool}` is globally unique by construction.
- Bare tool names never reach the agent; only fully-qualified names are
exposed, so unrelated servers with the same tool name (`query`) never clash.
- **Persistence:** the catalog is *derived* and rebuilt from Inventory +
upstream `tools/list` on boot; it is a cache, not a store. Namespace
assignments live with the `MCPInstance` in Inventory.
## Dependencies
- [Gateway](04-gateway.md) — consumes `List` for `tools/list`, forwards
`list_changed`.
- [RBAC](08-rbac.md) — supplies the `ToolView` filter applied to every list.
- [Health Monitoring](09-health-monitoring.md) — health transitions add/remove
instances from the catalog.
- [Auto Installer](03-auto-installer.md) — instance create/destroy events.
- [AI Agent Profiles](14-agent-profiles.md) — identity → roles → tool view.
- [Service Graph](17-service-graph.md) — consumes the tool→instance mapping as
graph edges.
## Failure modes & handling
- **Upstream `tools/list` fails on ingest:** keep the last-known good tool set
for that instance, mark it stale, and retry with backoff; do not blank the
namespace on a transient error.
- **`list_changed` storm:** debounce/coalesce refreshes per instance so a
flapping upstream cannot thrash the hot path.
- **Instance vanishes:** `RemoveInstance` drops its tools atomically; in-flight
`Resolve` for a removed tool returns `ok=false`, and the Router returns a
clean JSON-RPC "unknown tool" error.
- **Namespace collision at registration:** reject the newer instance's tools,
emit an operator notification, and keep the existing namespace intact.
- **Cold start:** serve an empty catalog that fills as instances report; never
block the Gateway waiting for a full rebuild.
## Security notes
- Every `tools/list` and `GET /tools` response is RBAC-filtered — an agent
cannot even *see* a tool outside its role (Architecture §9, request scoping).
- Cached schemas contain no secret material; secrets live only in the running
MCP container (see [Secrets Manager](07-secrets-manager.md)).
- Namespaces are validated (charset/length) to prevent spoofing a well-known
namespace like `github`.
## Open questions
- Namespace assignment when two instances legitimately wrap the *same* service
type (two Postgres servers): auto-suffix (`postgres`, `postgres-2`) vs.
operator-chosen aliases?
- How aggressively to pre-fetch/refresh schemas vs. lazy-load on first
`tools/call`.
- Whether to surface tool *deprecation*/version metadata in the catalog for the
Update Manager and Dashboard.
## Milestone
Phase 1 (Walking skeleton), alongside the [Gateway](04-gateway.md). **Exit:**
with a filesystem MCP and one HTTP MCP registered in YAML, an agent's
`tools/list` shows `filesystem.*` and `other.*`, and `Resolve` maps a chosen
tool to its owning instance for a successful `tools/call`.

View File

@@ -0,0 +1,127 @@
# Authentication
> Module 06 · Plane: Data · Roadmap phase: 3
> Part of [MCP Nexus architecture](../ARCHITECTURE.md).
## Purpose
Authentication turns an inbound connection at the [API Gateway](04-gateway.md) into a **verified principal** — a stable identity for either a human user (dashboard/API operator) or an AI agent (Claude, Cursor, …). It answers exactly one question: *"Who is calling?"* It does **not** decide what they may do; that is [RBAC](08-rbac.md). AuthN is the seam through which every request passes before it can be scoped, routed, or audited (Architecture §5, §9). It is pluggable by tenet #3: API keys first (Phase 3), OAuth2/OIDC and JWT sessions next, and LDAP/Active Directory/SAML (enabling SSO) later.
## Responsibilities
- Own a set of pluggable **AuthN providers** that each attempt to resolve request credentials into a `Principal`. Starter provider: **API keys**. Follow-ons: OAuth2, OpenID Connect (authorization-code + PKCE), JWT session validation, then LDAP, Active Directory, and SAML (SSO/enterprise IdP federation).
- Extract credentials from the request in a provider-appropriate way: `Authorization: Bearer <token>`, `X-API-Key`, a session cookie, or an MCP `initialize` handshake carrying an agent credential.
- Establish and validate **sessions**: mint short-lived JWT access tokens (with a refresh path) after an interactive login; validate their signature, expiry, audience, and issuer on every request.
- Issue, hash, store, list, and **revoke API keys**; keys are stored only as salted hashes, shown in plaintext exactly once at creation.
- Attach the resolved `Principal` (with its provider, credential id, and asserted identity claims) to the request context so downstream RBAC, Router, and Audit all read one canonical identity.
- Map an agent's presented credential to its [AgentProfile](14-agent-profiles.md) identity so profile→role resolution can proceed.
- Emit an audit event for every authentication decision — success, failure, and revocation — with principal, provider, and source address (Architecture §9).
## Non-goals
- **Authorization.** Deciding which tools/namespaces a principal may see or call is [RBAC](08-rbac.md); this module only proves identity.
- **The agent-facing identity binding.** Modeling per-agent policy objects is [AI Agent Profiles](14-agent-profiles.md); AuthN merely resolves a credential to the profile's identity.
- **Credential storage for upstream services.** MCP-server secrets live in the [Secrets Manager](07-secrets-manager.md); AuthN secures *access to Nexus*, not Nexus's access to backends.
- **Transport security.** TLS termination is the [Gateway](04-gateway.md) / [Security](19-security.md) baseline.
- **Plugin transport.** Out-of-process auth-backend plugins are wired via the [Plugin System](11-plugin-system.md) in Phase 5; this module defines only the in-process interface.
## Interfaces
```go
// A resolved, verified identity. Attached to every request context after AuthN.
type Principal struct {
ID string // stable identity key, e.g. "user:alice", "agent:claude-01"
Kind PrincipalKind // KindUser | KindAgent | KindService
Provider string // "apikey", "oidc", "jwt", "ldap", "saml"
Claims map[string]string // sub, email, groups, agent_id, ...
Expires time.Time // zero for non-expiring credentials (e.g. API key)
}
type PrincipalKind string
const (
KindUser PrincipalKind = "user"
KindAgent PrincipalKind = "agent"
KindService PrincipalKind = "service"
)
// A pluggable authentication backend (tenet #3).
type Provider interface {
Name() string // "apikey", "oidc", ...
// Authenticate inspects the request; ok=false if this provider does not
// recognize the credential, err for a malformed/invalid one.
Authenticate(ctx context.Context, r *http.Request) (Principal, bool, error)
}
// The authenticator runs providers in order and returns the first match.
type Authenticator interface {
Register(p Provider) error
Resolve(ctx context.Context, r *http.Request) (Principal, error) // ErrUnauthenticated if none match
}
// API-key lifecycle. Keys are stored hashed; plaintext is returned once.
type KeyManager interface {
Issue(ctx context.Context, subject string, opts KeyOpts) (plaintext string, id string, err error)
Verify(ctx context.Context, presented string) (Principal, error)
Revoke(ctx context.Context, id string) error
List(ctx context.Context, subject string) ([]KeyMeta, error)
}
```
HTTP/API surface (behind the Gateway; TLS-terminated):
- `POST /api/v1/auth/login` — interactive login → sets session cookie / returns JWT.
- `POST /api/v1/auth/refresh` — exchange a refresh token for a fresh access token.
- `POST /api/v1/auth/logout` — invalidate the current session.
- `GET /api/v1/auth/oidc/callback` — OAuth2/OIDC authorization-code redirect target.
- `GET /api/v1/auth/saml/acs` — SAML assertion consumer service (SSO, later).
- `POST /api/v1/auth/keys` · `GET /api/v1/auth/keys` · `DELETE /api/v1/auth/keys/{id}` — API-key CRUD (RBAC-gated).
MCP handshake: an agent presents its credential in the `initialize` request (bearer token or API key header on the Streamable HTTP transport); the resolved `Principal` is pinned to the MCP session for its lifetime.
Request-resolution flow (data plane, per request):
```
1. Gateway receives request (MCP over Streamable HTTP, or dashboard/API call).
2. Authenticator.Resolve iterates registered providers in priority order:
apikey → jwt/session → oidc → (ldap/ad/saml later)
3. First provider that recognizes the credential returns (Principal, true).
- none recognize it → ErrUnauthenticated → 401 (audited)
- one recognizes but it's bad → error → 401 (audited, no fall-through)
4. Principal attached to request context (id, kind, provider, claims).
5. For agents: Principal.ID feeds AgentProfile.Identify (Module 14).
6. RBAC (Module 08) reads the Principal to scope the request.
7. Audit records the decision: {principal, provider, source_ip, outcome}.
```
## Data
- **Reads/writes** the **Identity** store (Architecture §8): `Users`, `AgentProfiles`, `Roles`, and **API keys** (stored as `{id, subject, hash, salt, created, last_used, expires, revoked}`).
- **Reads** OIDC/SAML provider config (issuer URL, client id, JWKS endpoint, metadata) from the **Config** store; the OAuth **client secret** is a ref into the [Secrets Manager](07-secrets-manager.md), never inlined.
- **Writes** authentication events to the **Audit** store (append-only).
- JWT signing keys live in the **Secrets** store; the JWKS for verifying externally issued tokens is fetched and cached from the IdP.
## Dependencies
- Sits behind the [Gateway](04-gateway.md), which invokes `Resolve` as the first middleware on the data-plane hot path.
- Feeds the `Principal` to [RBAC](08-rbac.md) for authorization and to [AI Agent Profiles](14-agent-profiles.md) for identity→profile binding.
- OAuth client secrets and JWT signing keys come from [Secrets](07-secrets-manager.md).
- Every decision is recorded per [Security](19-security.md) §9 audit rules.
- Backend providers are plugin points via the [Plugin System](11-plugin-system.md) (Phase 5).
## Failure modes & handling
| Failure | Behavior |
|---|---|
| No provider recognizes the credential | `ErrUnauthenticated` → `401`; audited as a failed attempt with source IP. |
| Malformed/expired JWT | `401`; client directed to `/auth/refresh` or re-login; never falls through to another provider as valid. |
| IdP (OIDC/SAML) unreachable | Interactive login fails closed; already-issued, still-valid JWTs continue to work until expiry (no data-plane outage). |
| JWKS fetch fails | Serve from cached keys; refuse tokens signed by unknown `kid` rather than trusting blindly. |
| API key leaked/compromised | Immediate `Revoke`; hash comparison is constant-time; keys are rate-limited and can carry an expiry. |
| Brute-force credential guessing | Per-source and per-subject rate limiting (Security §9) + exponential backoff; repeated failures alert via [Notifications](13-notifications.md). |
| Clock skew on token validation | Bounded leeway (e.g. ±60s) on `nbf`/`exp`; beyond that, reject. |
## Security notes
Honors Architecture §9. Credentials are never logged and never written into `DiscoveredResource`, config, or audit payloads — only credential **ids** and outcomes are recorded. API keys are stored as salted hashes (argon2id/bcrypt), compared in constant time, and shown in plaintext once. JWTs are short-lived, signed with a key from [Secrets](07-secrets-manager.md), and validated for signature, `iss`, `aud`, and expiry on every request. OAuth flows use authorization-code + PKCE; client secrets are secret-refs. All auth traffic is TLS-only (fail closed if terminated plaintext). Sessions bind to the `Principal` for their lifetime so a mid-session privilege change forces re-resolution. Fail closed everywhere: an unresolvable request is unauthenticated, never anonymous-with-access.
## Open questions
- Should agent credentials be first-class API keys scoped to an [AgentProfile](14-agent-profiles.md), or a distinct credential type with its own lifecycle?
- mTLS as an additional agent authentication provider for zero-shared-secret deployments?
- Session store: stateless JWT only, or a server-side session table to enable instant global revocation (at a hot-path lookup cost)?
- How to federate group/role claims from multiple IdPs into one [RBAC](08-rbac.md) role space without collisions?
- Do we support step-up auth (re-authentication) for high-risk control-plane mutations?
## Milestone
Delivered in **Phase 3** (Secure — data-plane hardening). Thin slice first: **API-key authentication** end to end — issue a key, present it on the MCP handshake and the control API, resolve a `Principal`, and audit the decision — so two differently keyed agents can be told apart. OIDC/OAuth + JWT sessions follow within the phase; LDAP/AD/SAML SSO are deferred (Roadmap "not early"). Contributes to the Phase 3 exit proof: two agents with different identities connect and are distinguished before RBAC scopes them.

View File

@@ -0,0 +1,127 @@
# Secrets Manager
> Module 07 · Plane: Cross-cutting · Roadmap phase: 3
> Part of [MCP Nexus architecture](../ARCHITECTURE.md).
## Purpose
The Secrets Manager is Nexus's encrypted vault for the credentials that MCP servers need to talk to real services: API keys, passwords, SSH keys, tokens, and certificates. It stores every `Secret` (Architecture §7) **encrypted at rest** via envelope encryption, resolves **secret references** carried by [Recipes](15-smart-recipes.md)/config templates at install time, and hands the plaintext to the [Auto Installer](03-auto-installer.md) for **runtime injection** into the target MCP container (env var or file mount). It embodies tenet #5 and Architecture §9: secrets are **never returned to agents, never written into config sent to agents, and never logged** — unless an explicit policy permits it.
## Responsibilities
- Store `Secret` objects encrypted at rest in the **Secrets** store (Architecture §8), keyed by a stable name/ref.
- Perform **envelope encryption**: a per-secret data-encryption key (DEK) encrypts the payload; the DEK is wrapped by a master **key-encryption key (KEK)**. The KEK provider is pluggable (tenet #3): local master key (from file/env/OS keyring) first, with cloud KMS / HashiCorp Vault adapters later.
- Resolve **secret refs** (e.g. `secret://homeassistant/token`) to plaintext **only** at install/inject time, for the Installer, in memory.
- Provide **runtime injection** material to the Installer: env pairs or tmpfs-backed file mounts, so the plaintext lands inside the sandboxed container and nowhere else.
- **Rotate** secrets and re-wrap DEKs on KEK rotation without re-encrypting every payload; version secrets so an in-flight instance keeps working during rotation.
- **Audit every access** — who/what resolved which secret ref, when, and for which instance — to the append-only log (Architecture §9).
- Enforce a **redaction/leak policy**: refuse to serialize plaintext into any agent-visible payload, dashboard response, or log line; return refs or masked values instead.
## Non-goals
- **Authenticating callers to Nexus.** That is [Authentication](06-authentication.md); this module secures the credentials Nexus *uses*, not access *to* Nexus.
- **Deciding which tools a principal may use.** That is [RBAC](08-rbac.md).
- **Creating containers or templating config.** The [Auto Installer](03-auto-installer.md) injects; [Recipes](15-smart-recipes.md) declare refs. This module only stores, encrypts, and resolves.
- **Being a general end-user password manager or a public KMS.** It is scoped to MCP-server credentials.
- **Discovering credentials.** Discovery methods reference secrets by ref (Module 01); they do not populate the vault automatically.
## Interfaces
```go
// A stored secret. Payload is always encrypted at rest; plaintext exists only in memory.
type Secret struct {
Ref string // "secret://homeassistant/token" — the reference key
Type SecretType // apikey | password | ssh-key | token | certificate
Version int // incremented on rotation
CreatedAt time.Time
Policy AccessPolicy // who/what may resolve; may plaintext ever be revealed?
}
type SecretType string
// The store never returns Secret.plaintext except via Resolve, and never logs it.
type Store interface {
Put(ctx context.Context, ref string, plaintext []byte, t SecretType, p AccessPolicy) error
// Resolve returns plaintext IN MEMORY for injection; every call is audited.
Resolve(ctx context.Context, ref string, requester Principal) ([]byte, error)
Rotate(ctx context.Context, ref string, newPlaintext []byte) (version int, err error)
Delete(ctx context.Context, ref string) error
List(ctx context.Context) ([]SecretMeta, error) // metadata only — never plaintext
}
// Envelope encryption: a KEK provider wraps/unwraps per-secret DEKs (tenet #3).
type KEKProvider interface {
Name() string // "local", "aws-kms", "vault", ...
WrapDEK(ctx context.Context, dek []byte) (wrapped []byte, keyID string, err error)
UnwrapDEK(ctx context.Context, wrapped []byte, keyID string) (dek []byte, err error)
}
// What the Installer receives — injection material, not persisted plaintext.
type Injection struct {
Env map[string]string // VAR -> value (in-memory, container env)
Files map[string][]byte // mountPath -> content (tmpfs file mount)
}
type Injector interface {
// Resolve all refs in a rendered config into injection material for one instance.
Materialize(ctx context.Context, refs []string, instanceID string) (Injection, error)
}
```
HTTP/API surface (control-plane API, behind Auth, RBAC-gated; Admin/operator only):
- `POST /api/v1/secrets` — create a secret (write-only; plaintext in, never echoed back).
- `GET /api/v1/secrets` · `GET /api/v1/secrets/{ref}` — **metadata only** (type, version, created, last-accessed); never plaintext.
- `POST /api/v1/secrets/{ref}/rotate` — rotate; returns new version, not the value.
- `DELETE /api/v1/secrets/{ref}` — delete (guarded against deleting in-use refs).
- `POST /api/v1/secrets/rekey` — rotate the KEK and re-wrap all DEKs.
There is **no** endpoint, MCP tool, or dashboard view that returns a stored plaintext to a caller by default.
## Data
- **Reads/writes** the **Secrets** store (Architecture §8): `{ref, type, version, ciphertext, wrapped_dek, kek_key_id, iv, policy, created_at, last_accessed}` plus envelope-key metadata. Only ciphertext + wrapped DEKs are persisted.
- The **master KEK** is *not* stored in the database; it comes from the configured `KEKProvider` (local key material outside the DB, or an external KMS handle).
- **Reads** secret refs declared by [Recipes](15-smart-recipes.md)/config templates ([Registry](02-package-registry.md)) at install time.
- **Writes** every access to the **Audit** store (append-only) — ref, requester, instance, outcome; never the value.
Secret-ref resolution + injection flow (install time, control plane):
```
1. Reconciler decides to install an MCPInstance; Recipe/config template
carries refs, e.g. env HASS_TOKEN = secret://homeassistant/token.
2. Installer calls Injector.Materialize(refs, instanceID).
3. For each ref: Store.Resolve(ref, requester)
a. load {ciphertext, wrapped_dek, kek_key_id} from the Secrets store
b. KEKProvider.UnwrapDEK(wrapped_dek, kek_key_id) → DEK (in memory)
c. decrypt ciphertext with DEK → plaintext (in memory only)
d. audit {ref, requester, instanceID, outcome} — never the value
4. Materialize returns Injection{Env, Files} (plaintext held transiently).
5. Installer creates the sandboxed container with env pairs / tmpfs mounts.
6. Plaintext buffers cleared; it now lives only inside the container.
```
## Dependencies
- Injected by the [Auto Installer](03-auto-installer.md) into sandboxed containers ([Runtime](19-security.md)) at create time.
- Referenced by [Recipes](15-smart-recipes.md) and [Package Registry](02-package-registry.md) config templates via secret refs.
- Access requests carry a `Principal` from [Authentication](06-authentication.md); resolution policy may consult [RBAC](08-rbac.md).
- KEK providers are plugin points via the [Plugin System](11-plugin-system.md) (KMS/Vault adapters, Phase 5).
- Rotation events surface through [Notifications](13-notifications.md); leak-prevention is part of [Security](19-security.md) defense-in-depth.
## Failure modes & handling
| Failure | Behavior |
|---|---|
| KEK provider (KMS) unavailable | Cannot unwrap DEKs → new installs needing secrets **fail closed** and are retried; already-running instances keep their injected material (data plane unaffected). |
| Secret ref not found at install | Installer receives a typed `ErrSecretNotFound`; the instance is not created with a blank/placeholder credential; reconciler flags it for review. |
| Attempt to log/serialize plaintext | Blocked by a redaction wrapper; the value is masked (`****`) and a leak-attempt warning is audited. |
| Rotation mid-use | Versioned secrets: the running instance keeps its injected version; new/restarted instances get the new version; old version retired after drain. |
| KEK compromise suspected | `rekey` re-wraps all DEKs under a new KEK; payloads need not be re-encrypted; event audited + alerted. |
| Corrupt ciphertext / failed decrypt | Fail closed; surface as an integrity error; never return partial/garbage plaintext. |
| Delete of an in-use ref | Rejected unless forced; forcing warns that bound instances will fail on next restart. |
## Security notes
Honors Architecture §9 and tenet #5 as its core contract. **Encryption at rest** is envelope-based: unique per-secret DEK, DEK wrapped by a pluggable KEK/KMS; the master key never lives in the database. Plaintext exists **only transiently in memory** during `Resolve`/`Materialize` and inside the target sandboxed container (preferably tmpfs file mounts over env vars, since env is visible to child processes and some inspection paths). Secrets are **never** returned to agents, never placed in config sent to agents, and never logged — a redaction layer enforces this at the serialization boundary, and any attempt is audited. Every resolution is audited with requester + target instance. Access is RBAC/policy-gated; reading a stored plaintext back out is not a supported operation by default (write-and-inject, not read-back).
## Open questions
- Default injection mechanism: tmpfs file mounts vs env vars — trade off broad MCP-server compatibility against env-visibility risk.
- Policy grammar for the rare "agent may read this secret" exception — how is it expressed and how loudly is it audited?
- KEK rotation cadence and whether re-keying should be automatic on a schedule.
- Do we support external secret *sources* (read-through to Vault/KMS at inject time) in addition to storing our own ciphertext?
- Certificate lifecycle: does the Secrets Manager track expiry and drive renewal, or is that the [Update Manager](10-update-manager.md)?
## Milestone
Delivered in **Phase 3** (Secure). Thin slice first: local-KEK envelope encryption, `Put`/`Resolve`/`Rotate`, secret-ref resolution, and **runtime injection into a sandboxed container** by the Installer, with full audit and hard redaction on logs/agent payloads. KMS/Vault KEK providers are Phase 5 plugins. Contributes directly to the Phase 3 exit proof: a secret-backed MCP works without the secret ever appearing in any agent-visible payload or log.

129
docs/modules/08-rbac.md Normal file
View File

@@ -0,0 +1,129 @@
# RBAC
> Module 08 · Plane: Data · Roadmap phase: 3
> Part of [MCP Nexus architecture](../ARCHITECTURE.md).
## Purpose
RBAC (role-based access control) answers the question [Authentication](06-authentication.md) does not: *"Now that we know who is calling, what may they see and do?"* It maps a verified `Principal` to a set of **roles**, and each role to a set of allowed MCP **namespaces and tools**. The resolution collapses to one artifact the hot path cares about: the principal's **allowed tool set**. This is the concrete realization of Architecture §9 "request scoping" and tenet #5 — an agent may not call, *or even see*, a tool outside its role.
Example roles ship as a starter pack, each exposing a **different** slice of the tool catalog:
| Role | Sees (illustrative namespaces/tools) |
|---|---|
| Developer | `github.*`, `postgres.*`, `filesystem.*`, `docker.*` |
| Home Automation | `homeassistant.*`, `frigate.*`, `mqtt.*` |
| Networking | `unifi.*`, `pihole.*`, `traefik.*` |
| Infrastructure | `proxmox.*`, `docker.*`, `prometheus.*`, `grafana.*` |
| Finance | `finance.*`, curated read-only reporting tools |
| Admin | `*` (all namespaces) + control-plane mutations |
## Responsibilities
- Own the **Role** and **grant** model (Architecture §7): a role is a named set of allow rules over namespaces/tools; principals hold one or more roles (directly, or via IdP group claims).
- **Resolve** a `Principal` → roles → the union of allowed `{namespace}.{tool}` patterns → a concrete **allowed tool set**, evaluated against the current [Dynamic Tool Registry](05-dynamic-tool-registry.md).
- Provide two enforcement primitives the [Gateway](04-gateway.md)/Router calls on the hot path:
- **`Filter`** — at `tools/list`, return only the tools the principal may see.
- **`Can`** — at `tools/call`, decide whether this exact tool invocation is permitted.
- Support **wildcard and namespace-level grants** (`homeassistant.*`) as well as tool-level grants (`postgres.query`) and explicit denies (deny overrides allow).
- Cache resolved tool sets per principal with invalidation on role change, tool-catalog change, or session end.
- Reconcile with [AI Agent Profiles](14-agent-profiles.md): a profile is a binding of an agent identity to roles; RBAC is the engine that turns those roles into the visible set.
- Audit every enforcement decision, especially denials, with principal, tool, and outcome (Architecture §9).
## Non-goals
- **Authentication.** Establishing the `Principal` is [Authentication](06-authentication.md); RBAC trusts the resolved identity.
- **The agent-facing policy object.** Per-agent identity binding is [AI Agent Profiles](14-agent-profiles.md); RBAC is the underlying role engine both users and profiles share.
- **Secret access policy.** Whether a secret may be revealed is enforced by the [Secrets Manager](07-secrets-manager.md) (though it may consult roles); RBAC governs tool visibility/invocation.
- **Multi-tenancy above roles.** Projects/tenants are deferred (Architecture §12); v1 is RBAC only.
- **Row/field-level data authz inside a tool's result.** RBAC gates the *tool*, not the upstream service's internal permissions.
## Interfaces
```go
// A single allow/deny rule over the namespaced tool space.
type Rule struct {
Effect Effect // Allow | Deny (Deny wins on conflict)
Pattern string // "homeassistant.*", "postgres.query", "*"
}
type Effect string
const (
Allow Effect = "allow"
Deny Effect = "deny"
)
// A Role is a named set of rules (Architecture §7).
type Role struct {
Name string
Rules []Rule
}
// The allowed set resolved for a principal at a point in time.
type AllowedSet struct {
Principal string
Tools map[string]bool // fully-qualified "ns.tool" -> allowed
Roles []string
ResolvedAt time.Time
}
type Enforcer interface {
// Resolve principal -> roles -> concrete allowed tool set (against the live catalog).
Resolve(ctx context.Context, p Principal) (AllowedSet, error)
// Filter is called at tools/list: keep only visible tools.
Filter(ctx context.Context, p Principal, tools []Tool) ([]Tool, error)
// Can is called at tools/call: may this principal invoke this tool?
Can(ctx context.Context, p Principal, tool string) (bool, error)
}
type RoleStore interface {
Upsert(ctx context.Context, r Role) error
Get(ctx context.Context, name string) (Role, error)
List(ctx context.Context) ([]Role, error)
AssignRoles(ctx context.Context, principal string, roles []string) error
}
```
HTTP/API surface (control-plane API, behind Auth, RBAC-gated — Admin only):
- `GET /api/v1/roles` · `POST /api/v1/roles` · `PUT /api/v1/roles/{name}` · `DELETE /api/v1/roles/{name}`.
- `POST /api/v1/principals/{id}/roles` — assign/unassign roles to a user or agent profile.
- `GET /api/v1/principals/{id}/allowed` — debug view: the resolved allowed tool set.
MCP data-plane integration (not new endpoints — hooks inside the Router):
- On `tools/list` → `Enforcer.Filter` removes disallowed tools **before the response is serialized**.
- On `tools/call` → `Enforcer.Can` gates invocation; a disallowed call returns a JSON-RPC error (as if the tool did not exist), never reaching the upstream client.
## Data
- **Reads/writes** `Role` definitions and principal→role assignments in the **Identity** store (Architecture §8), alongside `Users` and `AgentProfiles`.
- **Reads** the live tool catalog from the [Dynamic Tool Registry](05-dynamic-tool-registry.md) to expand wildcard grants into concrete `ns.tool` entries.
- **Writes** authorization decisions (esp. denials) to the **Audit** store.
- Module-local: an in-memory, TTL'd cache of `AllowedSet` per principal, invalidated on role edits and catalog changes via the event bus.
## Dependencies
- Consumes the `Principal` from [Authentication](06-authentication.md).
- Enforced inside the [Gateway](04-gateway.md)/Router at both `tools/list` and `tools/call`.
- Expands grants against the [Dynamic Tool Registry](05-dynamic-tool-registry.md).
- Bound to agents by [AI Agent Profiles](14-agent-profiles.md) (identity → roles).
- Managed through the [Web Dashboard](12-web-dashboard.md) (roles, assignments).
- Decisions audited per [Security](19-security.md).
## Failure modes & handling
| Failure | Behavior |
|---|---|
| Principal has no roles | Empty allowed set → sees zero tools, can call nothing. **Fail closed** (default deny). |
| Role references a namespace/tool that no longer exists | Wildcard resolves against the live catalog; stale explicit grants are simply inert (no error, no phantom access). |
| Allow and Deny both match a tool | **Deny wins** — explicit denies always override allows. |
| Tool catalog changes mid-session | Cached `AllowedSet` invalidated on the catalog-change event; next `tools/list`/`tools/call` re-resolves. New tools are not auto-visible until re-resolution. |
| Role edited while an agent is connected | Cache invalidated; the change takes effect on the next request without dropping the connection. |
| RBAC store unavailable | Fail closed — deny rather than fall back to allow; surfaced via [Notifications](13-notifications.md). |
| A denied `tools/call` | Returned as a JSON-RPC method-not-found / not-permitted error; never forwarded upstream; audited. |
## Security notes
Honors Architecture §9 and tenet #5. Enforcement is **dual-gate and mandatory**: `tools/list` filtering means a disallowed tool is invisible (no information leak about what exists), and `tools/call` gating means even a guessed tool name cannot be invoked — the two together defeat both enumeration and direct-call bypass. Default is **deny**: absence of a grant is denial, and explicit `Deny` overrides any allow. Wildcards expand against the live catalog so a newly installed tool is not silently exposed to an over-broad role without re-resolution. All role mutations and all denials are audited with principal + tool. RBAC is enforced in the Router (server side), never the client — the agent's view is a *consequence* of enforcement, not the enforcement itself.
## Open questions
- Role composition: flat roles only, or hierarchical/inheritable roles for large orgs?
- Do we need per-tool **parameter** constraints (e.g. `postgres.query` read-only), or is tool-level granularity enough for v1? (Leaning: tool-level v1.)
- How do IdP group claims map onto Nexus roles — static mapping table, or claim-expression rules?
- Should denies be expressible per-principal (an override) as well as per-role?
- Time-bounded / just-in-time role grants for elevated access windows?
## Milestone
Delivered in **Phase 3** (Secure). Thin slice first: the starter role pack (Developer, Home Automation, Admin, …), principal→role assignment, and **dual-gate enforcement in the Router** at `tools/list` and `tools/call` with default-deny. Directly powers the Phase 3 exit proof: two agents with different roles connect and each sees a *different* tool set, with disallowed tools neither visible nor callable.

View File

@@ -0,0 +1,124 @@
# Health Monitoring
> Module 09 · Plane: Control · Roadmap phase: 4
> Part of [MCP Nexus architecture](../ARCHITECTURE.md).
## Purpose
Keep every managed [`MCPInstance`](../ARCHITECTURE.md) alive and serving. Health Monitoring continuously probes each instance/container, tracks its lifecycle state (`running`, `offline`, `updating`, `restarting`, `quarantined`), measures latency and error rates, and **automatically heals** unhealthy instances via restart-with-backoff, escalating to quarantine after repeated failure. It is the `act`/`record` feedback arm of the reconciliation loop (Architecture §6): it turns observed instance health into events the reconciler, Router, Update Manager, and Notifications react to. A failure of the control plane must never take down the data plane, so Health only *signals* the Router which instances are poolable — it never sits in the hot path.
## Responsibilities
- Run **liveness** and **readiness** probes against every `MCPInstance` on a per-instance schedule: liveness = "the process/container is up"; readiness = "it completes an MCP `initialize` + `tools/list` within budget."
- Collect health signals from three layers: **container** (Docker inspect: state, restart count, OOM/exit codes), **transport** (MCP client connect success, RTT), and **application** (probe tool-call latency, JSON-RPC error rate over a sliding window).
- Own the per-instance **health state machine** and drive transitions; persist current state + last transition into Inventory.
- Execute the **restart policy**: on failed liveness, restart the instance via the Runtime with exponential backoff and jitter; **quarantine** after N consecutive failed restarts.
- Emit health-transition events on the event bus so: the [Router](04-gateway.md) drains dead instances, the [Update Manager](10-update-manager.md) can roll back a bad update, and [Notifications](13-notifications.md) fire on `offline`/`quarantined`.
- Expose current + historical health via the control API for the [Dashboard](12-web-dashboard.md) and feed gauges/counters to [Metrics](18-metrics.md).
## Non-goals
- **Does not pull, create, or destroy containers** — it *requests* restart/recreate through the Runtime + [Auto Installer](03-auto-installer.md); Runtime owns container lifecycle.
- **Does not decide desired state.** Whether an instance *should* exist is the reconciler's call from Recipes ⨯ Inventory; Health only reports and heals what exists.
- **Does not route traffic.** It signals poolability; the [Gateway](04-gateway.md)/Router owns `UpstreamPool.Drain`.
- **Does not define probe transport plugins** — those interfaces are generalized by the [Plugin System](11-plugin-system.md) in Phase 5.
- **Does not send messages** — it emits events; delivery is [Notifications](13-notifications.md).
## Interfaces
```go
// HealthState is the per-instance lifecycle state.
type HealthState string
const (
StateRunning HealthState = "running"
StateOffline HealthState = "offline" // liveness failing
StateUpdating HealthState = "updating" // owned transiently by Update Manager
StateRestarting HealthState = "restarting" // heal in progress
StateQuarantined HealthState = "quarantined" // gave up after repeated restarts
)
// Probe is a pluggable health check against one instance.
type Probe interface {
Name() string // "container", "mcp-readiness", "tool-latency"
Check(ctx context.Context, inst *MCPInstance) ProbeResult
}
type ProbeResult struct {
Healthy bool
Latency time.Duration
ErrRate float64 // over the probe's sliding window
Reason string // human-readable on failure
Kind ProbeKind // Liveness | Readiness
}
// Monitor owns probing, the state machine, and the restart policy.
type Monitor interface {
Register(p Probe) error
Watch(ctx context.Context) error // periodic + edge-triggered
Status(instanceID string) (HealthReport, bool) // current snapshot
Quarantine(ctx context.Context, instanceID, reason string) error
Release(ctx context.Context, instanceID string) error // un-quarantine
}
// RestartPolicy computes backoff and the quarantine threshold.
type RestartPolicy struct {
Base time.Duration // e.g. 1s
Max time.Duration // cap, e.g. 5m
Multiplier float64 // e.g. 2.0
Jitter float64 // 0..1
MaxRestarts int // consecutive failures → quarantine
HealthyReset time.Duration // uptime after which the failure count resets
}
```
Internal HTTP (control API, RBAC-guarded; not agent-facing):
- `GET /api/v1/health/instances` — list instances with state, latency, error rate, restart count.
- `GET /api/v1/health/instances/{id}` — full `HealthReport` + transition history.
- `POST /api/v1/health/instances/{id}/restart` — operator-forced restart.
- `POST /api/v1/health/instances/{id}/quarantine` · `/release` — manual quarantine controls.
## State machine
```
probe ok
┌──────────────────────────────┐
▼ │
running ──liveness fail──▶ offline ──heal──▶ restarting ──ready──▶ running
│ │ │
│ update begins │ │ restart fails × MaxRestarts
▼ ▼ ▼
updating ──(rollback path)──▶ ◀──────────── quarantined ──operator release──▶ restarting
```
`updating` is entered/left by the [Update Manager](10-update-manager.md); Health suppresses auto-restart while an instance is `updating` and instead reports readiness so the Updater can decide **rollback on unhealthy**. `quarantined` instances are drained from the Router and require operator release (or a successful update) to re-enter the loop.
## Data
- **Writes** `health` + `state` on each `MCPInstance` in the **Inventory** store (§8), plus an append-only transition log (state, reason, timestamp) for the dashboard timeline.
- **Reads** instance definitions from Inventory and probe/policy config from the **Config** store; probe credentials (if any) by ref from [Secrets](07-secrets-manager.md).
- Module-local in-memory ring buffers hold recent latency/error samples per instance for sliding-window rate computation (survive-restart not required; recomputed on boot).
## Dependencies
- [Auto Installer](03-auto-installer.md) + Runtime — executes restart/recreate on Health's request.
- [Gateway](04-gateway.md) / Router — consumes health events via `UpstreamPool.Drain`; stops routing to non-`running` instances.
- [Update Manager](10-update-manager.md) — reads readiness to gate promotion and trigger **rollback on unhealthy**; owns the `updating` state.
- [Notifications](13-notifications.md) — fires on `offline`/`quarantined`/recovered transitions.
- [Metrics](18-metrics.md) — exports state, latency, error-rate, and restart-count series.
- [Plugin System](11-plugin-system.md) — generalizes the `Probe` interface for custom checks.
## Failure modes & handling
| Failure | Behavior |
|---|---|
| Instance fails liveness | Transition `running→offline`, drain from Router, restart with backoff (`restarting`). |
| Restarts keep failing | After `MaxRestarts` consecutive failures → `quarantined`; stop restarting; notify; require operator/update to recover. |
| Flapping (ready→fail→ready) | Backoff prevents restart storms; failure count only resets after `HealthyReset` sustained uptime. |
| Probe itself times out / Docker API down | Probe error is not the same as instance-down: mark `unknown`, retry with backoff, do **not** restart on probe infrastructure failure alone. |
| Update in progress | Auto-restart suppressed; Health reports readiness to the Updater, which owns rollback. |
| Health monitor crashes | Data plane unaffected (Router keeps last-known poolable set until TTL); on restart, probes re-establish state from Inventory. |
| False-positive readiness | Application probe (real tool-call latency + error rate), not just container-up, reduces "up but broken" routing. |
## Security notes
Honors Architecture §9. Probes are least-privilege: readiness uses the same sandboxed MCP client path as the Router, never a privileged shell into the container. Probe credentials come from [Secrets](07-secrets-manager.md) by ref and are never logged or written into `HealthReport`. Failure `Reason` strings are sanitized before they reach the dashboard or notifications so upstream error payloads cannot leak secrets or internal topology. Every restart/quarantine/release is a mutating action and is **audited** with principal (or `system:health`) + target + outcome. Restart requests to the Runtime are the only privileged action and flow through the same audited Installer path.
## Open questions
- Should backoff/quarantine thresholds be per-`type` (a flaky IoT MCP vs. a rack Postgres MCP differ), overridable per Recipe?
- Do we expose raw probe latency histograms to the dashboard, or only rolled-up percentiles (topology-leak risk, per Discovery's parallel question)?
- Quarantine auto-recovery: should Nexus periodically re-probe quarantined instances with a long backoff, or strictly require operator/update release?
- How do readiness budgets interact with legitimately slow-starting MCP servers (model loads, large indexes)?
## Milestone
Delivered in **Phase 4** (Operate / day-2). Thin slice: container + MCP-readiness probes, the five-state machine, exponential-backoff auto-restart, and quarantine-after-N, wired to Router drain and Notifications. Exit proof (shared with the phase goal): kill a managed MCP container → it is auto-restarted, a Discord alert fires, and the dashboard reflects `running→offline→restarting→running` in real time.

View File

@@ -0,0 +1,167 @@
# Update Manager
> Module 10 · Plane: Control · Roadmap phase: 4
> Part of [MCP Nexus architecture](../ARCHITECTURE.md).
## Purpose
The Update Manager keeps managed `MCPInstance`s current and safe. It watches
upstream sources for newer versions of the packages an instance was installed
from — Docker image tags/digests, GitHub releases, and
[Registry](02-package-registry.md) index updates — decides whether to act based
on **policy** (automatic, manual-approval-gated, or pinned), and applies updates
as **re-installs** through the [Auto Installer](03-auto-installer.md). Crucially,
it keeps the previous known-good spec so any update can be **rolled back** safely.
It is a day-2 control-plane reconciler (Architecture §5): a failure here degrades
*management currency*, never the data plane. Per tenet #2 it proposes and applies
state transitions through the loop, it does not run ad-hoc upgrade scripts.
## Responsibilities
- **Watch** upstream sources per instance and detect available updates:
- Docker image **tags/digests** (a new digest behind a tracked tag).
- **GitHub releases** for `source=github` packages.
- **Registry index updates** (a package's recommended version advanced).
- Evaluate each candidate against the instance's **update policy**:
`automatic`, `manual` (approval-gated), or `pinned` (never auto-update).
- Enforce **version pinning**: a pinned instance is skipped entirely (still
reported as "update available").
- Apply an approved update by driving a **re-install** at the new digest via the
[Installer](03-auto-installer.md), preserving the same `bound_resource_uuid`.
- Retain the **previous known-good `MCPInstance` spec** and perform **rollback**
when the new version fails to become/stay healthy.
- Coordinate with [Health Monitoring](09-health-monitoring.md) to gate and
validate updates, and with [Notifications](13-notifications.md) to alert on
available / applied / failed / rolled-back updates.
- Record all decisions and transitions to Audit (Architecture §9).
## Non-goals
- Performing the actual pull/create/start — delegated to the
[Installer](03-auto-installer.md) and Runtime.
- Choosing/verifying package signatures or digests — owned by the
[Registry](02-package-registry.md).
- Restarting crashed-but-same-version instances — that is
[Health Monitoring](09-health-monitoring.md) (self-healing), not an update.
- Deciding which package a service uses ([Recipes](15-smart-recipes.md)).
## Interfaces
```go
type UpdatePolicy string
const (
PolicyAuto UpdatePolicy = "automatic" // apply without human action
PolicyManual UpdatePolicy = "manual" // detect, then wait for approval
PolicyPinned UpdatePolicy = "pinned" // never auto-update
)
// Everything-is-a-plugin (tenet #3): each upstream is a watcher.
type UpdateWatcher interface {
Name() string // "docker-digest", "github-release", "registry-index"
// Return a candidate if a newer version exists for this instance.
Check(ctx context.Context, inst MCPInstance) (Candidate, bool, error)
}
type Candidate struct {
InstanceID string
From PackageVersion // current (digest-pinned)
To PackageVersion // proposed (digest-pinned)
Source string // watcher name
Notes string // e.g. GitHub release notes
}
type UpdateManager interface {
Scan(ctx context.Context) ([]Candidate, error) // periodic + event-driven
Approve(ctx context.Context, candidateID string) error
Apply(ctx context.Context, c Candidate) (MCPInstance, error) // re-install
Rollback(ctx context.Context, instanceID string) error // to known-good
SetPolicy(ctx context.Context, instanceID string, p UpdatePolicy) error
Pin(ctx context.Context, instanceID string, version string) error
}
```
HTTP/API surface (behind Auth + [RBAC](08-rbac.md)):
- `GET /api/v1/updates` — available/pending candidates across instances.
- `POST /api/v1/updates/{id}/approve` — approve a manual-gated candidate.
- `POST /api/v1/updates/{id}/apply` — force-apply now (audited).
- `POST /api/v1/instances/{id}/rollback` — roll back to previous known-good.
- `PUT /api/v1/instances/{id}/update-policy` — set policy / pin a version.
## Data
- **Reads:** `MCPInstance` (current `package`, `version`, digest) from
**Inventory**; `Package`/`PackageVersion` from the **Registry** store
(Architecture §8); instance `health` from Health Monitoring.
- **Writes:** on apply, the Installer updates the `MCPInstance` (new `version`,
`container_ref`); the Update Manager updates `state` around the transition
(`Updating`, `RollingBack`).
- **Module-local persistence:**
- `update_policy` per instance (`automatic|manual|pinned`, optional pinned
version/digest).
- `previous_good_spec` — the last healthy `MCPInstance` spec (package, version,
digest, rendered config-with-secret-refs) for rollback.
- `update_candidate` rows (from→to, source, status: available/approved/applied/
failed/rolled-back) and per-watcher cursors/ETags to avoid re-alerting.
## Dependencies
- [Auto Installer](03-auto-installer.md) — updates are re-installs of a new version.
- [MCP Package Registry](02-package-registry.md) — resolves/validates target digests.
- [Health Monitoring](09-health-monitoring.md) — gates the new version; signals rollback.
- [Notifications](13-notifications.md) — alerts on available/applied/failed/rolled-back.
- [Secrets Manager](07-secrets-manager.md) — re-injected at the re-install (unchanged).
- [Plugin System](11-plugin-system.md) — `UpdateWatcher` is a plugin point.
- [RBAC](08-rbac.md) — who may approve/apply/pin/rollback.
## Failure modes & handling
- **New version unhealthy:** the Installer's health gate fails, or Health
Monitoring reports the freshly updated instance unhealthy within a validation
window → automatic **rollback** to `previous_good_spec`; fire a failed-update
notification.
- **Watcher/upstream unreachable:** skip silently (best-effort), keep the last
cursor; repeated failures alert but never block other updates or the data plane.
- **Registry can't validate target digest:** candidate is rejected (not applied);
reported as `failed` with reason.
- **Pinned instance with newer upstream:** never applied; surfaced as
"update available (pinned)" only.
- **Rollback itself fails:** instance marked degraded, high-severity notification,
left for Health Monitoring quarantine + human intervention.
- **Concurrent apply + heal:** updates are serialized per `instance-id`
(shared lock with the Installer) so an update and a restart don't race.
- **Approval race / stale candidate:** applying a candidate whose `From` no longer
matches the live instance is rejected and re-scanned.
## Security notes
- Only **digest-pinned, signature-verified** target versions are applied; the
Registry re-verifies before the Installer pulls (Architecture §9).
- Secrets are re-injected at runtime by the Installer on the re-install — never
copied into `previous_good_spec` as plaintext; the retained spec holds secret
*references* only.
- Approve/apply/pin/rollback are audited mutating actions with principal, target,
from→to versions, and outcome.
- Auto-update policy is itself an RBAC-guarded setting; a compromised auto policy
is a supply-chain risk, so downgrades/policy changes are audited.
## Open questions
- Health **validation window** length after an update before declaring success —
fixed, or per-recipe?
- Should GitHub-release watching drive version *selection* directly, or only
notify the Registry to advance its index (single source of truth)?
- Batch/canary updates across many instances of the same package vs. one-at-a-time.
- Retention depth for `previous_good_spec` — just N−1, or a short history for
multi-step rollback?
- Maintenance windows / update scheduling to avoid disrupting active agent calls.
## Milestone
Phase 4 — *Operate (day-2)*. Update Manager ships with Docker-digest, GitHub-release,
and Registry-index watchers; automatic/manual/pinned policies; and health-gated
rollback. It contributes to the Phase 4 exit story alongside
[Health](09-health-monitoring.md), [Metrics](18-metrics.md),
[Notifications](13-notifications.md), and the [Dashboard](12-web-dashboard.md).

View File

@@ -0,0 +1,106 @@
# Plugin System
> Module 11 · Plane: Cross-cutting · Roadmap phase: 5
> Part of [MCP Nexus architecture](../ARCHITECTURE.md).
## Purpose
The extensibility backbone that makes tenet #3 real — **"everything is a plugin."** The Nexus core knows only *interfaces*; concrete discovery methods, fingerprinters, installers, package providers, notifiers, auth backends, dashboards, and tool transformers are swappable implementations. This module does not invent those interfaces — each owning module already defines its own (`discovery.Method`, `notify.Notifier`, `health.Probe`, …). Module 11 **generalizes** them into one plugin contract, lifecycle, discovery/registration mechanism, and a recommended **out-of-process transport** so third parties can extend Nexus, in any language, without forking or recompiling the core.
## Responsibilities
- Define the common **plugin contract**: a lifecycle (`register → configure → start → health → stop`) and metadata (name, type, version, capabilities) every plugin implements regardless of its type.
- Enumerate and load the supported **plugin types**, each mapping to an interface already owned by a core module: discovery methods, fingerprinters, installers, package providers, notifiers, auth backends, dashboards, and tool transformers.
- Provide the **out-of-process transport** (gRPC over a handshake-negotiated local socket, in the style of `hashicorp/go-plugin`) so plugins run as isolated subprocesses — a crashing or malicious plugin cannot take down the core.
- Support **in-process (compiled-in) plugins** too, for first-party built-ins and performance-sensitive paths — the transport is an implementation detail behind the same contract.
- Own **versioning & compatibility**: a semver'd plugin API, capability negotiation on the handshake, and refusal to load incompatible plugins.
- Manage plugin **discovery, install, config, and health**: a plugin registry, per-plugin config (secrets by ref), and health surfacing to the [Dashboard](12-web-dashboard.md).
## Non-goals
- **Does not define the per-type interfaces' semantics.** What a `Method` or `Notifier` *means* is owned by [Discovery](01-discovery-engine.md), [Notifications](13-notifications.md), etc. Module 11 provides the wrapper, transport, and lifecycle.
- **Not the MCP tool plugins agents call.** Those are upstream `MCPInstance`s managed by the control loop. This is about extending *Nexus itself*.
- **Not a general sandbox runtime** — container sandboxing of MCP servers is [Security](19-security.md) + Runtime; plugin isolation here is process-level (separate proc, dropped privileges, scoped host interface).
- **Not a package registry.** Distributing/signing plugin artifacts leans on [Package Registry](02-package-registry.md) + [Security](19-security.md) supply-chain rules.
## Interfaces
```go
// Plugin is the universal contract every plugin implements, regardless of type.
type Plugin interface {
// Info returns identity + declared capabilities used for compatibility checks.
Info() PluginInfo
// Configure applies validated config; secrets arrive as refs, resolved by host.
Configure(ctx context.Context, cfg map[string]any) error
// Health reports readiness so the dashboard/plugin manager can surface it.
Health(ctx context.Context) error
// Stop releases resources; the host may kill the subprocess after a deadline.
Stop(ctx context.Context) error
}
type PluginInfo struct {
Name string
Type PluginType // Discovery, Fingerprinter, Installer, PackageProvider,
// Notifier, AuthBackend, Dashboard, ToolTransformer
Version string // plugin's own version
APIVer string // Nexus plugin-API semver it was built against
Caps []string
}
// Host is the surface the core exposes back to a plugin (least-privilege).
type Host interface {
Recorder() metrics.Recorder // scoped metrics
Secret(ref string) (string, error) // resolve a Secret by ref only
Logger() Logger // structured, plugin-namespaced
EventBus() EventEmitter // emit typed events (RBAC-scoped)
}
// Manager loads, negotiates, and supervises plugins (in- or out-of-process).
type Manager interface {
Load(ctx context.Context, spec PluginSpec) (handle PluginHandle, err error)
// TypedDispatch adapts a loaded plugin to the owning module's interface,
// e.g. as a discovery.Method or notify.Notifier, over the transport.
Bind(handle PluginHandle) (any, error)
List() []PluginStatus
Unload(ctx context.Context, name string) error
}
```
Transport: a `go-plugin`-style handshake starts the subprocess, negotiates the API version + a shared secret, and multiplexes typed **gRPC** services (one service definition per plugin type) over a local Unix socket. Each core module ships the gRPC ⇆ Go-interface adapter for its type.
Lifecycle: `Load` (spawn + handshake + version check) → `Configure` (validated config, secrets by ref) → `Bind` (adapt to the owning module's typed interface) → *serving* (health-polled) → `Unload`/`Stop` (graceful, then kill after a deadline). A plugin that fails `Health` is restarted with backoff; the owning module keeps functioning on its built-in implementations while a plugin is down, so an extension is never a single point of failure for a core capability.
Internal HTTP (control API, RBAC-guarded):
- `GET /api/v1/plugins` — installed plugins, type, version, health, config status.
- `POST /api/v1/plugins` — install/register a plugin (artifact ref + config).
- `POST /api/v1/plugins/{name}/reload` · `DELETE /api/v1/plugins/{name}`.
## Data
- **Writes** the plugin registry (installed plugins, type, version, artifact digest, enabled state) to the **Config**/Registry store (§8).
- **Reads** plugin config from Config and plugin credentials by ref from [Secrets](07-secrets-manager.md) — never handed the raw vault.
- Plugin subprocess state (PID, socket, handshake) is module-local runtime state, rebuilt on restart from the registry.
## Dependencies
- Generalizes interfaces owned by: [Discovery](01-discovery-engine.md) (`Method`, `Fingerprinter`), [Auto Installer](03-auto-installer.md) (installer/runtime adapters), [Package Registry](02-package-registry.md) (package providers), [Notifications](13-notifications.md) (`Notifier`), [Authentication](06-authentication.md) (auth backends), [Web Dashboard](12-web-dashboard.md) (panels), [Health](09-health-monitoring.md) (`Probe`), and the Router ([tool transformers](05-dynamic-tool-registry.md)).
- [Secrets Manager](07-secrets-manager.md) — resolves plugin credential refs.
- [Security](19-security.md) — signature verification + isolation policy for plugin artifacts.
- [Metrics](18-metrics.md) — plugins emit through a scoped `Recorder`.
## Failure modes & handling
| Failure | Behavior |
|---|---|
| Plugin subprocess crashes | Isolated by the transport; the Manager marks it unhealthy, restarts with backoff, and the owning module falls back to built-ins — the core stays up. |
| Version incompatibility | Handshake rejects a plugin built against an incompatible API semver; it is not loaded; surfaced in dashboard. |
| Plugin hangs / slow RPC | Every host↔plugin call is deadline-bounded; a hung plugin is killed after a grace period, not awaited. |
| Malicious/greedy plugin | Runs out-of-process with dropped privileges and only the least-privilege `Host` surface (no direct DB, secrets only by ref); resource limits applied. |
| Bad plugin config | `Configure` validates and rejects atomically; the plugin stays in its prior good state or disabled. |
| Unsigned/tampered artifact | [Security](19-security.md) supply-chain check refuses to install (signature + digest pin). |
## Security notes
Honors Architecture §9 and tenet #5. Out-of-process isolation is the core safety property: plugins get **no ambient authority** — only the narrow `Host` interface (scoped metrics, secrets by ref, namespaced logging, RBAC-scoped event emit), never the database, the raw vault, or the host network unless policy grants it. Plugin artifacts are **signature-verified and digest-pinned** before load (supply chain). Installing/reloading/removing a plugin is a mutating, [RBAC](08-rbac.md)-gated, audited action. Plugin logs and emitted events are sanitized so a plugin cannot exfiltrate secrets or forge another principal's identity.
## Open questions
- Transport: standardize on `hashicorp/go-plugin` directly, or a thin in-house gRPC handshake to avoid the dependency surface?
- Do we support **WASM** plugins (sandboxed, portable) as a third mode alongside in-process and subprocess?
- How rich should the `Host` surface be before it becomes an attack surface — where is the line between capability and least privilege?
- Compatibility policy: strict semver gate, or capability-negotiation that allows partial feature sets?
- Plugin distribution/governance — reuse the community registry + recipe trust model (Architecture §12)?
## Milestone
Delivered in **Phase 5** (Extend & scale). The per-type interfaces exist earlier (each module defines its own from its phase); Phase 5 promotes them to a stable, semver'd plugin API with the out-of-process gRPC transport, plugin manager, and registry. Exit proof (shared with the phase goal): a **community-authored discovery plugin + recipe installs a new MCP with no core changes**, running as an isolated subprocess.

View File

@@ -0,0 +1,108 @@
# Web Dashboard
> Module 12 · Plane: Data · Roadmap phase: 4
> Part of [MCP Nexus architecture](../ARCHITECTURE.md).
## Purpose
The human control surface for Nexus: a modern **React** UI that shows what has been discovered, what is installed, whether it is healthy, what changed, and who may access what — with real-time updates and dark mode. Per tenet #4, the compiled UI is **embedded in the Go binary via `embed.FS`** and served by the same process on the same port, so there is no separate frontend to deploy, build, or version-skew. It is a *data-plane* surface for operators (distinct from the agent MCP edge) and sits behind the same [Auth](06-authentication.md) + [RBAC](08-rbac.md) as everything else — the dashboard is a privileged client, not a bypass.
## Responsibilities
- Serve the embedded SPA (`embed.FS`) plus a versioned, RBAC-guarded **backend API** the UI consumes; API and assets are one binary, one origin.
- Provide **real-time updates** over WebSocket (with SSE fallback): live instance state, discovery events, logs, health transitions, and update progress stream without polling.
- Render the full **page map** (below), each backed by the owning module's control API — the dashboard aggregates, it does not own the data.
- Authenticate operators (session/JWT from [Auth](06-authentication.md)) and render only what the operator's [role](08-rbac.md) permits — hide pages/actions the role can't perform.
- Stream container **logs** and expose action buttons (restart, quarantine, approve update, test notification) that call the owning module's API — every action audited.
- Support **dark mode** (default) and be responsive for tablet/desktop operator use.
- Provide the operator workflows that close the loop by hand when automation is below the confidence threshold (tenet #6): approve a low-confidence discovery, promote a gated update, release a quarantined instance, edit a recipe match.
- Surface **build/version info** and per-module health so an operator can tell at a glance whether the control plane itself is healthy, independent of any single MCP instance.
## Non-goals
- **Owns no domain data or business logic.** Every panel reads/writes through another module's API; the dashboard holds only view state.
- **Not the agent endpoint.** Agents speak MCP to the [Gateway](04-gateway.md); the dashboard is for humans and never routes tool calls.
- **Not an auth provider.** It consumes [Auth](06-authentication.md)/[RBAC](08-rbac.md); it defines no policy.
- **Not a separate service.** No standalone Node server in production — assets are embedded; a dev proxy is a build-time convenience only.
- **Not a plugin host runtime** — dashboard *plugins* (custom panels) are a [Plugin System](11-plugin-system.md) type; the shell just mounts them.
## Interfaces
```go
//go:embed all:dist
var uiFS embed.FS
// DashboardServer mounts the SPA + the aggregating backend API on the shared mux.
type DashboardServer interface {
// Mount registers static asset routes (from uiFS) and /api/v1/* handlers,
// all wrapped by the Auth + RBAC middleware chain.
Mount(r Router, mw ...Middleware) error
// Realtime upgrades a request to WS/SSE and streams events the caller
// is authorized to see (RBAC-filtered event bus fan-out).
Realtime(w http.ResponseWriter, r *http.Request) error
}
```
Real-time transport:
- `GET /api/v1/stream` — WebSocket (SSE fallback at `GET /api/v1/events`); server pushes RBAC-filtered `Event`s from the bus (instance state, discovery, logs, update progress). Client subscribes to topics; server enforces visibility.
Real-time model: the browser opens one multiplexed connection and subscribes to topics (`instances`, `discovery`, `logs:{id}`, `updates`). The server tees the internal event bus, applies the connection's RBAC filter, and fans out only authorized events — the same event stream Notifications and Metrics consume, so the dashboard shows exactly the transitions the rest of the system acted on, with no separate polling loop drifting out of sync.
Asset pipeline: the React app is built (`vite build`) into `dist/` at compile time and embedded with `//go:embed`. There is no runtime Node process; a Vite dev server proxying to a running `nexus` is a developer convenience only. The binary is the single deployable artifact (tenet #4).
Backend API surface (each page consumes the owning module's control API):
| Page | Backing API / module |
|---|---|
| Dashboard (overview) | aggregate of Health, Metrics, Inventory |
| Discovered Services | [Discovery](01-discovery-engine.md) `/discovery/resources` |
| Installed MCPs | [Installer](03-auto-installer.md) / Inventory `/instances` |
| Containers | Runtime `/containers` |
| Logs | Runtime log stream (via `/stream`) |
| Metrics | [Metrics](18-metrics.md) `/metrics` + Grafana links |
| Health | [Health](09-health-monitoring.md) `/health/instances` |
| Updates | [Update Manager](10-update-manager.md) `/updates` |
| Settings | Config store `/settings` |
| Secrets | [Secrets](07-secrets-manager.md) `/secrets` (refs only) |
| Users | [Auth](06-authentication.md) `/users` |
| Roles / Permissions | [RBAC](08-rbac.md) `/roles`, `/permissions` |
| Discovery (config) | [Discovery](01-discovery-engine.md) `/discovery/methods` |
| Plugins | [Plugin System](11-plugin-system.md) `/plugins` |
| Registry | [Package Registry](02-package-registry.md) `/packages`, `/recipes` |
| Agents | [AI Agent Profiles](14-agent-profiles.md) `/agents` |
| Notifications | [Notifications](13-notifications.md) `/notifications/channels`, `/rules` |
| Permissions | [RBAC](08-rbac.md) `/permissions` (effective grant inspector) |
Pages fall into three groups: **observe** (Dashboard, Discovered Services, Installed MCPs, Containers, Logs, Metrics, Health), **govern** (Users, Roles, Permissions, Agents, Secrets), and **operate** (Updates, Discovery config, Plugins, Registry, Settings, Notifications). The nav and the realtime subscription set are both derived from the operator's role, so an operator only loads and streams the groups they may act on.
## Data
- **Owns no persistent store.** All reads/writes proxy to module APIs; the browser holds ephemeral view/query-cache state only.
- Static assets (the React `dist/`) are compiled into the binary via `embed.FS` at build time — versioned with the binary, never fetched at runtime.
- Real-time state is derived from the shared event bus, RBAC-filtered per connection.
## Dependencies
- [Authentication](06-authentication.md) — operator login / sessions.
- [RBAC](08-rbac.md) — page/action/data visibility; drives the realtime fan-out filter.
- Every module with a control API (Discovery, Registry, Installer, Health, Updates, Secrets, Notifications, Agents, Plugins, Metrics) — the dashboard is their aggregating client.
- [Metrics](18-metrics.md) — the Metrics page and links to shipped Grafana dashboards.
- [Plugin System](11-plugin-system.md) — custom dashboard panels are a plugin type.
## Failure modes & handling
| Failure | Behavior |
|---|---|
| A backing module API is down | The affected panel shows a scoped error/empty state; the rest of the dashboard keeps working (no all-or-nothing render). |
| WebSocket drops | Client reconnects with backoff and falls back to SSE, then to periodic REST polling as last resort. |
| Operator lacks a role for a page | The page/action is hidden **and** the API rejects (defense in depth); RBAC is enforced server-side, not just in the UI. |
| Stale asset vs. API version | Assets are embedded with the binary, so UI and API versions can't skew; a build-info banner surfaces the version. |
| Large log stream | Streaming with backpressure + client-side windowing; server caps per-connection throughput. |
| Session expiry mid-session | 401 triggers a re-auth flow; unsaved form state is preserved where feasible. |
| Control plane degraded (reconciler down) | The dashboard clearly flags degraded management while the data-plane panels (Gateway/Router health) keep reporting — the split of Architecture §5 is visible, not hidden. |
| Browser clock skew / event replay | Events carry server timestamps + monotonic sequence; the client orders by sequence, not local time. |
## Security notes
Honors Architecture §9. The dashboard is served over the same **TLS**-terminated edge and is fully behind [Auth](06-authentication.md) + [RBAC](08-rbac.md) — there is no anonymous view. RBAC is enforced **server-side** for every API call and for the realtime fan-out; hiding a control in the UI is UX, never the security boundary. The Secrets page shows **references and metadata only** — secret values are never sent to the browser (tenet #5). Every mutating action from the UI is audited with the operator principal. Standard web hardening applies: CSRF protection on state-changing requests, strict CSP, same-origin API, and no third-party asset CDNs (everything embedded).
## Open questions
- WebSocket vs. SSE as the primary transport — WS is bidirectional but SSE is simpler behind proxies; ship both with WS default?
- How much log history to buffer server-side for the Logs page before requiring the operator to pull from container/host?
- Do dashboard plugins (custom panels) load as sandboxed iframes/web-components, or as build-time-composed federated modules?
- Mobile: a responsive read-mostly view for on-call, or defer entirely to notifications?
## Milestone
Delivered in **Phase 4** (Operate / day-2). Thin slice: the embedded React shell with Dashboard, Discovered Services, Installed MCPs, Health, Logs, and Updates pages, live WebSocket updates, dark mode, behind Auth/RBAC. Exit proof (shared with the phase goal): kill a managed MCP container → the Health/Installed pages reflect `offline→restarting→running` in real time without a manual refresh.

View File

@@ -0,0 +1,111 @@
# Notifications
> Module 13 · Plane: Cross-cutting · Roadmap phase: 4
> Part of [MCP Nexus architecture](../ARCHITECTURE.md).
## Purpose
Tell a human when something meaningful happens. Notifications turns internal events — a new service discovered, an update available or applied, a failed install, a service gone offline, a security alert — into messages delivered over **pluggable notifiers**: Discord, Slack, Email, Pushover, ntfy, Telegram, and Signal. It owns the `notify` stage of the reconciliation loop (Architecture §6): the reconciler and other modules only *emit typed events*; this module decides what is worth telling whom, over which channel, how often, and in what words.
## Responsibilities
- Define a typed **event model** and subscribe to the event bus for notification-worthy transitions.
- Implement a **Notifier plugin interface** with built-in notifiers for Discord, Slack, Email (SMTP), Pushover, ntfy, Telegram, and Signal.
- Own the **event→channel routing/subscription** model: rules that map (event type, severity, source) → one or more configured channels.
- Assign and honor **severity levels** (`info`, `warning`, `error`, `critical`) with per-channel minimum-severity filters.
- Provide **throttling, dedup, and grouping**: collapse repeated/flapping events, rate-limit per channel, and batch bursts into digests.
- **Template** messages per event type with a per-channel renderer (rich embeds for Discord/Slack, plain text for ntfy/SMS-style, subject+body for email).
- Record delivery outcome (sent/failed/suppressed) to Audit and emit `nexus_notifications_sent_total` to [Metrics](18-metrics.md).
## Non-goals
- **Does not detect conditions.** It never polls health or updates itself — it reacts to events emitted by [Health](09-health-monitoring.md), [Update Manager](10-update-manager.md), [Discovery](01-discovery-engine.md), [Installer](03-auto-installer.md), and [Security](19-security.md).
- **Not metric alerting.** Threshold-on-series alerting is Prometheus Alertmanager against [Metrics](18-metrics.md); this module is event-driven.
- **Not a message queue / inbox.** Fire-and-forward with bounded retry; it is not durable messaging or a ticketing system.
- **Not the plugin transport.** The out-of-process notifier transport is generalized by the [Plugin System](11-plugin-system.md) in Phase 5; here notifiers are in-process interfaces.
- **Does not store secrets** — channel credentials (webhook URLs, SMTP creds, bot tokens) come by ref from [Secrets](07-secrets-manager.md).
## Interfaces
```go
// Notifier is a pluggable delivery channel.
type Notifier interface {
Name() string // "discord", "slack", "email", "ntfy", "telegram", "pushover", "signal"
Configure(ctx context.Context, cfg ChannelConfig) error // creds via Secret refs
Send(ctx context.Context, msg Message) error
Health(ctx context.Context) error // reachability check for the dashboard
}
// Event is what the rest of Nexus emits onto the bus.
type Event struct {
Type EventType // ServiceDiscovered, UpdateAvailable, UpdateApplied,
// InstallFailed, ServiceOffline, SecurityAlert, ...
Severity Severity // Info | Warning | Error | Critical
Source string // module + subject, e.g. "health/postgres-01"
Subject string // short title
Fields map[string]any // structured detail for templating
Time time.Time
DedupKey string // events sharing a key collapse within a window
}
// Router decides which channels an event reaches and applies throttling.
type Router interface {
Register(n Notifier) error
Route(ctx context.Context, e Event) error // subscription match → throttle → render → Send
Test(ctx context.Context, channel string) error // "send test notification"
}
// Template renders an Event into a channel-specific Message.
type Template interface {
Render(e Event, channel string) (Message, error)
}
```
Internal HTTP (control API, RBAC-guarded):
- `GET/POST /api/v1/notifications/channels` — list/configure channels (creds by Secret ref).
- `POST /api/v1/notifications/channels/{name}/test` — send a test message.
- `GET/POST /api/v1/notifications/rules` — manage event→channel subscription rules.
- `GET /api/v1/notifications/history` — recent delivery log (from Audit).
### Templating
Each event type has a base template rendered per channel: rich embeds (title, color-by-severity, fields, action link back to the [Dashboard](12-web-dashboard.md)) for Discord/Slack; `subject` + body for email; a compact single line for ntfy/Pushover/Telegram/Signal. Templates are overridable in config and rendered from the `Event.Fields` map, so a new event type ships a default template without code changes. Rendering is isolated from delivery — a template error degrades to a safe plain-text fallback rather than dropping the alert.
### Event → severity defaults
| Event | Default severity |
|---|---|
| `ServiceDiscovered` | info |
| `UpdateAvailable` | info |
| `UpdateApplied` | info / warning (if rollback) |
| `InstallFailed` | error |
| `ServiceOffline` / quarantined | error / critical |
| `SecurityAlert` | critical |
## Data
- **Reads** channel + rule config from the **Config** store; channel credentials by ref from [Secrets](07-secrets-manager.md) (never persisted here, never logged).
- **Writes** delivery outcomes to the **Audit** store (§8) — append-only: event, channels, sent/failed/suppressed, timestamp.
- Module-local in-memory state: dedup windows, per-channel rate-limit token buckets, and pending-digest buffers (rebuilt on restart; at-most-once semantics on crash).
## Dependencies
- Consumes events from [Health](09-health-monitoring.md), [Update Manager](10-update-manager.md), [Discovery](01-discovery-engine.md), [Auto Installer](03-auto-installer.md), and [Security](19-security.md).
- [Secrets Manager](07-secrets-manager.md) — channel credentials by ref.
- [Plugin System](11-plugin-system.md) — the `Notifier` interface is one of its first-class plugin types; third-party notifiers land Phase 5.
- [Metrics](18-metrics.md) — `nexus_notifications_sent_total{channel,severity,outcome}`.
- [Web Dashboard](12-web-dashboard.md) — channel config, rule editing, and history UI.
## Failure modes & handling
| Failure | Behavior |
|---|---|
| A channel is unreachable (webhook 5xx, SMTP down) | Bounded retry with backoff; then mark delivery `failed` in Audit; surface channel health in dashboard. Never block the emitter. |
| Notification storm (flapping instance) | Dedup by `DedupKey` within a window + per-channel rate limiting; bursts collapse into a digest. |
| A notifier plugin hangs | `Send` is deadline-bounded; a slow/hung channel is isolated and does not delay others (per-channel workers). |
| Misconfigured channel (bad token) | `Configure`/`Test` validates on save; runtime failures degrade to `failed` + dashboard warning, not a crash. |
| Nexus restarts mid-burst | At-most-once: pending in-memory digests may be lost; durable audit records what *was* sent. Critical events are sent eagerly, not batched. |
| Secret rotation | Channel re-reads cred by ref on next send; no restart required. |
## Security notes
Honors Architecture §9. Channel credentials (webhook URLs, bot tokens, SMTP passwords) are [Secrets](07-secrets-manager.md) referenced by ref, never stored in config sent to the dashboard and never logged. Message bodies are **sanitized before send**: no secret material, no raw agent payloads, and security-alert messages avoid leaking exploit detail to low-trust channels. Configuring/testing channels and editing rules are mutating actions, gated by [RBAC](08-rbac.md) and audited. `SecurityAlert` routing should prefer channels with delivery guarantees (email/Pushover) over best-effort chat webhooks.
## Open questions
- Do we support **acknowledgement / two-way** interactions (e.g. approve an update from a Slack button), or keep notifications strictly one-way in v1?
- Digest cadence: fixed windows, or adaptive based on event rate?
- Per-user vs. per-system channels — should individual [users](12-web-dashboard.md) subscribe personal channels, or are channels a system-wide concern in v1?
- Signal delivery requires a linked device/`signal-cli` sidecar — bundle guidance or treat as advanced/optional?
## Milestone
Delivered in **Phase 4** (Operate / day-2). Thin slice: the `Notifier` interface with Discord + Email built in, event subscriptions for `ServiceDiscovered`/`InstallFailed`/`ServiceOffline`/`UpdateApplied`/`SecurityAlert`, severity filtering, and basic dedup/throttle. Exit proof (shared with the phase goal): kill a managed MCP container → Health emits `ServiceOffline` → a Discord alert fires within one cycle.

View File

@@ -0,0 +1,120 @@
# AI Agent Profiles
> Module 14 · Plane: Data · Roadmap phase: 3
> Part of [MCP Nexus architecture](../ARCHITECTURE.md).
## Purpose
An **AgentProfile** (Architecture §7) is the agent-facing binding of *identity → roles → visible tool set*. Where [Authentication](06-authentication.md) proves *who* is connecting and [RBAC](08-rbac.md) is the general role→tool engine, an AgentProfile is the first-class object that ties a specific AI agent — Claude, Cursor, VSCode, OpenWebUI, Codex, Aider, Continue.dev — to the roles it holds and therefore the exact tools it may see and call at the [Gateway](04-gateway.md). It is what makes "one endpoint, many agents" (tenet #1) concrete: every connecting agent resolves to a profile, and the profile decides its worldview.
## Responsibilities
- Own the **AgentProfile** model: `{id, identity, allowed_roles}` (Architecture §7), plus display metadata (agent kind, description, owner) and an association to one or more credentials.
- **Identify** a connecting agent: map the credential presented at the MCP `initialize` handshake (an API key, token, or client identifier resolved by [Authentication](06-authentication.md)) to exactly one AgentProfile.
- **Resolve** the profile to a concrete visible tool set by handing its `allowed_roles` to the [RBAC](08-rbac.md) `Enforcer` — the profile does not compute tool visibility itself; it *binds* the identity to roles that RBAC expands.
- Provide the Router with the profile's resolved `AllowedSet` so `tools/list` and `tools/call` are scoped per agent (enforcement stays in RBAC/Router, per tenet #5).
- Maintain per-profile operational metadata: last-seen, connection count, and (optionally) per-profile rate-limit / quota hints for [Security](19-security.md).
- Support **provisioning**: create a profile, mint/attach its agent credential (via [Authentication](06-authentication.md)), and assign roles — the operator workflow for onboarding a new agent.
## Non-goals
- **Authenticating the credential.** Verifying the presented secret is [Authentication](06-authentication.md); the profile consumes the resolved `Principal`.
- **Computing tool visibility.** The allow/deny expansion is [RBAC](08-rbac.md); the profile only supplies the roles.
- **Being a human-user object.** Human operators are `Users` in the Identity store; profiles model *agents*. Both share the RBAC engine, but a profile is the agent binding specifically (this is how it differs from raw RBAC).
- **Per-tool parameter policy.** Fine-grained call constraints, if any, belong to [RBAC](08-rbac.md)/[Security](19-security.md).
- **Managing agent-side config.** Nexus does not push settings into Claude/Cursor; it presents an endpoint and scopes it.
## Interfaces
```go
// The agent-facing binding of identity -> roles (Architecture §7).
type AgentProfile struct {
ID string // "profile:claude-primary"
Identity string // stable agent identity key, matches Principal.ID for KindAgent
Kind string // "claude" | "cursor" | "vscode" | "openwebui" | "codex" | "aider" | "continue"
DisplayName string
AllowedRoles []string // role names expanded by RBAC
CredentialIDs []string // API keys / tokens that resolve to this profile
CreatedAt time.Time
LastSeen time.Time
}
type ProfileStore interface {
Upsert(ctx context.Context, p AgentProfile) error
Get(ctx context.Context, id string) (AgentProfile, error)
// ByCredential maps an authenticated credential -> its profile (identify step).
ByCredential(ctx context.Context, credentialID string) (AgentProfile, error)
List(ctx context.Context) ([]AgentProfile, error)
}
// Resolver ties identify -> roles -> visible tool set via RBAC.
type Resolver interface {
// Identify: authenticated Principal (an agent) -> its profile.
Identify(ctx context.Context, p Principal) (AgentProfile, error)
// VisibleTools: profile's roles expanded by the RBAC Enforcer against the live catalog.
VisibleTools(ctx context.Context, prof AgentProfile) (AllowedSet, error)
}
```
HTTP/API surface (control-plane API, behind Auth, RBAC-gated; Admin/operator):
- `GET /api/v1/agents` · `POST /api/v1/agents` — list / create profiles.
- `GET /api/v1/agents/{id}` · `PUT /api/v1/agents/{id}` · `DELETE /api/v1/agents/{id}`.
- `POST /api/v1/agents/{id}/roles` — set the profile's `allowed_roles`.
- `POST /api/v1/agents/{id}/credentials` — mint/attach an agent credential (delegates to [Authentication](06-authentication.md)).
- `GET /api/v1/agents/{id}/tools` — debug: the resolved visible tool set for this profile.
MCP data-plane integration: on `initialize`, the Router calls `Identify(Principal)` to pin the profile to the session; `tools/list`/`tools/call` are then scoped by the profile's `VisibleTools` (enforced in RBAC/Router).
Identify → resolve flow (per agent session):
```
1. Agent (e.g. Cursor) opens an MCP session and sends `initialize`
carrying its credential (API key / token on Streamable HTTP).
2. Authentication (Module 06) verifies it → Principal{Kind: agent, ID}.
3. Resolver.Identify(Principal) → ProfileStore.ByCredential → AgentProfile
- no match → fail closed (empty set) unless a guest profile is set.
4. Resolver.VisibleTools(profile): hand profile.AllowedRoles to the RBAC
Enforcer, which expands them against the live tool catalog (Module 05)
→ AllowedSet (the agent's worldview).
5. Profile + AllowedSet pinned to the MCP session.
6. tools/list → RBAC.Filter(AllowedSet); tools/call → RBAC.Can(...).
7. last_seen updated; provisioning/role changes audited (Module 19).
```
Example: profile `claude-primary` (identity `agent:claude-01`) holds roles
`[Developer, Home Automation]`; RBAC expands these to `github.*`, `postgres.*`,
`filesystem.*`, `homeassistant.*` — so this agent sees exactly those namespaces
and nothing from Networking/Finance/Admin.
## Data
- **Reads/writes** `AgentProfile` rows in the **Identity** store (Architecture §8), alongside `Users`, `Roles`, and API keys; a credential→profile index supports the identify step.
- **Reads** the [RBAC](08-rbac.md) role definitions to expand `allowed_roles`, and the [Dynamic Tool Registry](05-dynamic-tool-registry.md) (via RBAC) for the live catalog.
- **Writes** profile provisioning and role-binding changes to the **Audit** store.
- Module-local: `last_seen`/connection counters updated on the hot path (cheap, async).
## Dependencies
- Credential verification and the `Principal` come from [Authentication](06-authentication.md).
- Role expansion and enforcement come from [RBAC](08-rbac.md); profiles supply the roles.
- Visibility is scoped against the [Dynamic Tool Registry](05-dynamic-tool-registry.md).
- Enforced at the [Gateway](04-gateway.md)/Router on the hot path.
- Managed through the [Web Dashboard](12-web-dashboard.md) (agent onboarding).
- Provisioning + access audited per [Security](19-security.md).
## Failure modes & handling
| Failure | Behavior |
|---|---|
| Credential resolves to no profile | Treated as an unrecognized agent → no roles → empty visible set (**fail closed**); connection allowed only if a default/guest profile is explicitly configured. |
| One credential maps to two profiles | Rejected at provisioning; the credential→profile index is unique. Ambiguity is a configuration error, not a runtime coin-flip. |
| Profile references a deleted role | Missing role contributes nothing (no phantom access); resolution logs the dangling reference. |
| Roles changed while agent connected | RBAC cache invalidated; next `tools/list`/`tools/call` reflects the new set without dropping the session. |
| Profile deleted while agent connected | Session's pinned profile is invalidated; subsequent requests fail closed and the agent must reconnect. |
| Identity store unavailable | Fail closed — unresolved profile means no tools; surfaced via [Notifications](13-notifications.md). |
## Security notes
Honors Architecture §9 and tenets #1/#5. A profile is the point where an opaque agent connection becomes a **scoped, named identity** — everything downstream (visibility, invocation, audit) keys off it. Default is **fail closed**: an agent whose credential maps to no profile, or a profile with no roles, sees nothing. The profile never widens access on its own; it can only reference roles, and RBAC enforces them server-side, so a compromised/misconfigured agent cannot self-elevate. Every profile lifecycle event (create, role change, credential attach/detach) is audited with principal + target. Per-profile rate-limit hints feed the [Security](19-security.md) baseline so one noisy agent cannot starve others (tenet #1's "many agents" fairness).
## Open questions
- Should a profile map to exactly one credential or many (e.g. Claude Desktop + Claude Code sharing one profile vs. distinct profiles)?
- Do we ship a curated starter profile pack (Claude, Cursor, VSCode, …) with sensible default roles, or start every profile empty?
- Is there value in profile-level tool *overrides* on top of roles, or does that erode the clean identity→roles→tools chain?
- How do profiles interact with per-agent quotas/budgets (token usage) tracked in [Metrics](18-metrics.md)?
- Guest/anonymous profile: supported at all, or is every agent required to be provisioned first?
## Milestone
Delivered in **Phase 3** (Secure). Thin slice first: the `AgentProfile` model, credential→profile `Identify`, and role expansion via [RBAC](08-rbac.md) so a connecting agent resolves to a scoped visible tool set at the Router. Directly powers the Phase 3 exit proof: two agents (distinct profiles, distinct roles) connect and each sees a different tool set — the profile is what binds each agent to its slice.

View File

@@ -0,0 +1,123 @@
# Smart Recipes
> Module 15 · Plane: Control · Roadmap phase: 2
> Part of [MCP Nexus architecture](../ARCHITECTURE.md).
## Purpose
Provide declarative **YAML rules that map a fingerprint to an MCP install** — no code required to support a new service. A recipe binds `match` → `install` → `config`, and is the **desired-state input** to the reconciler: given a `DiscoveredResource` (from [Discovery Engine](01-discovery-engine.md)) and a `Package` (from the [Registry](02-package-registry.md)), a matching recipe resolves into a concrete `MCPInstance` spec for the [Auto Installer](03-auto-installer.md).
## Responsibilities
- Define the **Recipe schema** (Architecture §7 `Recipe`): `match`, `install`, `config`, plus metadata (`name`, `version`, `priority`).
- **Match** a `DiscoveredResource` against recipe predicates (fingerprint `type`, open `ports`, `http_title`, capabilities, version constraints, evidence fields) and produce a **match score**.
- **Resolve precedence** when multiple recipes match a single resource, deterministically.
- **Render config templates** — substitute variables (`{{ip}}`, `{{port}}`, `{{hostname}}`, `{{version}}`) and **secret references** into the config a `Package`'s `config_schema` expects.
- Reference **Registry packages** (by name + optional version constraint) and **Secrets** (by ref, never inline).
- Ship a **starter recipe pack** and load user/community recipes from the Config store.
- Validate recipes against package `config_schema` and report actionable errors.
## Non-goals
- **Does not discover or fingerprint** — it consumes `DiscoveredResource`s from [Discovery Engine](01-discovery-engine.md).
- **Does not resolve or verify packages** — that is the [Registry](02-package-registry.md) (signatures, digests, versions).
- **Does not install or run** the resulting spec — the [Auto Installer](03-auto-installer.md) + Runtime do.
- **Does not store or decrypt secrets** — it emits secret *refs*; [Secrets](07-secrets-manager.md) injects values at container runtime.
- **Does not own the reconcile loop** — it is a pure(ish) resolver the reconciler calls.
## Interfaces
```go
// A Recipe is a declarative match+install+config rule (Architecture §7).
type Recipe struct {
Name string
Version string
Priority int // tie-breaker; higher wins
Match MatchSpec // predicates over a DiscoveredResource
Install InstallSpec // package ref (docker_image / package name)
Config map[string]any // templated values -> Package.config_schema
}
// The engine matches, scores, and resolves recipes into instance specs.
type Engine interface {
Load(ctx context.Context) error // from Config store
Match(r DiscoveredResource) []Scored // sorted, best first
Resolve(r DiscoveredResource, pick Recipe) (MCPInstanceSpec, error)
Validate(rec Recipe, pkg Package) error // against config_schema
}
type Scored struct {
Recipe Recipe
Score float64 // match strength, 0.0–1.0
}
// Resolve renders templates + secret refs into a spec the Installer consumes.
type MCPInstanceSpec struct {
Package PackageRef // resolved via Registry
Config map[string]any // templates rendered; secrets as refs
SecretRefs []SecretRef // injected at runtime, never inlined
BoundResource string // DiscoveredResource.uuid
}
```
## Matching, scoring & precedence
- **Matching:** every predicate in `match` must hold (AND semantics). Predicates: `fingerprint` (type), `ports`, `http_title` (regex/substring), `capabilities`, `version` (semver constraint), and arbitrary `evidence.*` equality.
- **Scoring:** a matched recipe's score combines predicate **specificity** (more/stricter predicates → higher) with the resource's own discovery **confidence**. This favors precise recipes over broad catch-alls.
- **Precedence** when several recipes match one resource, in order:
1. Highest explicit `priority`.
2. Then highest match score (specificity × confidence).
3. Then most specific version constraint.
4. Then recipe `name` (stable, deterministic) as final tie-break.
- The reconciler acts on the winner only if `confidence ≥ threshold` (tenet #6); below threshold it queues for human approval rather than auto-installing.
## Example recipe — Home Assistant
```yaml
name: home-assistant
version: 1.2.0
priority: 50
match:
fingerprint: home-assistant # DiscoveredResource.type
ports: [8123]
http_title: "Home Assistant" # substring/regex over evidence.http_title
capabilities: [rest-api]
install:
package: mcp-home-assistant # resolved via Registry (name + constraint)
version: ">=0.4 <1.0"
config:
base_url: "http://{{ip}}:{{port}}" # {{port}} -> 8123 from the resource
token: "{{secret:home-assistant/llat}}" # secret ref, injected at runtime
verify_tls: false
```
Given a discovered Home Assistant at `192.168.1.20:8123`, this resolves to an `MCPInstanceSpec` for package `mcp-home-assistant`, config `base_url=http://192.168.1.20:8123`, and a `SecretRef` to `home-assistant/llat` — the token value is never written into the spec or shown to agents.
## Data
- **Reads** `Recipe`s from the **Registry**/**Config** stores (§8; recipes are syncable from remote indexes like packages).
- **Reads** `DiscoveredResource`s from **Inventory** and `Package` metadata (incl. `config_schema`) from **Registry**.
- **Produces** an `MCPInstanceSpec` consumed by the reconciler/Installer to create an `MCPInstance` (§7). It does not itself persist instances.
- References `Secret`s by ref only.
## Dependencies
- Input from [Discovery Engine](01-discovery-engine.md) (`DiscoveredResource` + confidence).
- Package resolution via [MCP Package Registry](02-package-registry.md).
- Secret refs resolved at runtime by [Secrets](07-secrets-manager.md).
- Output consumed by the [Auto Installer](03-auto-installer.md).
- User-authored/community recipes distributed under the [Plugin System](11-plugin-system.md) / community registry governance (Phase 5).
## Failure modes & handling
| Failure | Behavior |
|---|---|
| No recipe matches a resource | Resource stays visible in Inventory; no install; optionally surfaced as "no recipe" for a human to author one. |
| Multiple recipes match | Deterministic precedence (priority → score → version → name) picks one; the discarded matches are recorded for transparency. |
| Referenced package not found in Registry | Resolution error; spec not produced; flagged for operator; reconciler retries after next Registry sync. |
| Config template references a missing variable/secret | Validation fails **before** install; recipe rejected with the exact missing key; no partial install. |
| Recipe config violates package `config_schema` | Rejected at `Validate`; never handed to the Installer. |
| Confidence below threshold | Match computed but held for human-in-the-loop approval (tenet #6). |
## Security notes
Honors Architecture §9: secrets appear in recipes **only as refs** (`{{secret:...}}`) and are injected into the MCP container at runtime by [Secrets](07-secrets-manager.md) — never inlined into config, never persisted in the rendered spec, never returned to agents or logged. Community/user recipes are untrusted input: they are schema-validated and their referenced packages are signature-verified by the Registry before any install (supply-chain rule). Applying a recipe is an audited, mutating action. Recipes cannot grant a container more privilege than the package/runtime policy allows.
## Open questions
- Do we allow OR/`anyOf` predicate groups, or keep strict AND for predictability?
- Should scoring weights (specificity vs confidence) be tunable per deployment?
- Community recipe trust model: signing, review, namespacing (ties to Architecture §12).
- Multi-instance: when one host exposes several matching services, how do recipes express "one instance per resource" vs "one shared instance"?
- Template language scope — keep it to safe variable substitution, or allow limited expressions?
## Milestone
Delivered in **Phase 2** (Discover → install), as **Smart Recipes v1: match→install→config rules + a starter recipe pack**. Thin slice that lands first: strict-AND matching with the precedence rules above and safe `{{variable}}`/`{{secret:...}}` substitution, proven by the Postgres recipe that turns a discovered Postgres into a live `postgres.query` at the Gateway within one reconcile cycle. OR-predicates, tunable scoring, and community distribution follow later.

View File

@@ -0,0 +1,98 @@
# Infrastructure Discovery
> Module 16 · Plane: Control · Roadmap phase: 2
> Part of [MCP Nexus architecture](../ARCHITECTURE.md).
## Purpose
Discover and connect to the **compute substrates** Nexus can deploy MCP containers onto — and enumerate the workloads already running on them. Where credentials allow, auto-connect and register each substrate as a **Runtime** target the [Auto Installer](03-auto-installer.md) can use. This is the "where can things run?" half of discovery, complementing [Discovery Engine](01-discovery-engine.md)'s "what services exist?".
## Responsibilities
- Detect and probe substrates: **Docker**, **Docker Compose** projects, **Kubernetes**, **Nomad**, **Proxmox**, **LXC**, **VirtualBox**, **VMware**, **Hyper-V**, and **bare metal** (SSH-reachable hosts).
- **Auto-connect** where credentials/sockets are available (local Docker socket, in-cluster K8s service account, `KUBECONFIG`, Proxmox API token, SSH keys), with graceful degradation to "detected but not connected" otherwise.
- Register each connected substrate as a **Runtime target** — capabilities (can it run OCI containers? host-network? privileged?), capacity hints, and a health status the reconciler/Installer can place against.
- Enumerate existing workloads on each substrate (containers, pods, VMs, jails) and emit them so Discovery Engine's fingerprinters can classify them (a running container is both a *host workload* here and a candidate *service* there).
- Track substrate lifecycle: appear, become reachable/unreachable, drain, disappear — emitted on the shared event bus.
## Non-goals
- **Does not fingerprint application services** into product types — it emits raw workloads for [Discovery Engine](01-discovery-engine.md) to classify.
- **Does not create/start MCP containers** — it provides the Runtime *targets*; the [Auto Installer](03-auto-installer.md) does placement and lifecycle.
- **Does not implement the Runtime abstraction** (create/exec/logs/destroy). It *populates* Runtime targets; the Runtime interface itself is owned alongside the Installer.
- **Does not manage substrate credentials at rest** — those live in [Secrets](07-secrets-manager.md).
- **Does not schedule/bin-pack** beyond exposing capability + capacity hints (advanced placement is future work).
## How the two discovery modules relate
Both are Control-plane, Phase 2, and **share the event bus and Inventory store**. The split is by *target of discovery*:
| | Module 1 — Discovery Engine | Module 16 — Infrastructure Discovery |
|---|---|---|
| Finds | Application **services** (Postgres, Home Assistant…) | Compute **substrates**/hosts (Docker, K8s, Proxmox…) |
| Emits | `DiscoveredResource` (a thing to *manage via* MCP) | `RuntimeTarget` (a place to *run* MCP containers) + raw workloads |
| Feeds | [Recipes](15-smart-recipes.md) → what MCP to install | [Installer](03-auto-installer.md) → where to install it |
They cooperate: Module 16 connects a Docker host, enumerates its containers, and forwards each as an Observation; Module 1's Docker fingerprinter then classifies those containers into typed `DiscoveredResource`s. Conversely, when the Installer needs to place an `MCPInstance`, it picks from the `RuntimeTarget`s Module 16 registered.
## Interfaces
```go
// A prober detects and (optionally) connects to one class of substrate.
type SubstrateProber interface {
Name() string // "docker", "kubernetes", "proxmox", ...
Detect(ctx context.Context, scope Scope) ([]Candidate, error)
Connect(ctx context.Context, c Candidate, creds SecretRef) (RuntimeTarget, error)
}
// RuntimeTarget is a connected substrate the Installer can place onto.
type RuntimeTarget struct {
ID string
Kind string // "docker" | "kubernetes" | "proxmox" | ...
Endpoint string // socket / api url / ssh host
Capabilities RuntimeCaps // OCI, host-network, privileged, gpu...
Capacity CapacityHint // cpu/mem/limits, best-effort
Health Health
Workloads []WorkloadRef // existing containers/pods/vms/jails
}
type Registry interface {
Register(t RuntimeTarget) error
Targets(ctx context.Context, filter TargetFilter) ([]RuntimeTarget, error)
Watch(ctx context.Context) (<-chan TargetEvent, error)
}
```
Internal HTTP (RBAC-guarded, not agent-facing):
- `GET /api/v1/infra/targets` — list runtime targets + health/capabilities.
- `POST /api/v1/infra/connect` — attempt connection to a detected candidate with a secret ref.
- `GET /api/v1/infra/targets/{id}/workloads` — enumerate existing workloads.
## Data
- **Writes** `RuntimeTarget` records and their enumerated `WorkloadRef`s into the **Inventory** store (§8), alongside `DiscoveredResource`s.
- **Reads** substrate connection settings from **Config**; substrate credentials from **Secrets** (by ref only).
- Emits workloads as Observations consumed by [Discovery Engine](01-discovery-engine.md).
## Dependencies
- Feeds the Runtime targeting used by the [Auto Installer](03-auto-installer.md).
- Shares event bus + Inventory with [Discovery Engine](01-discovery-engine.md).
- Prober interfaces provided by the [Plugin System](11-plugin-system.md); new runtime adapters (containerd, K8s operator, LXC) widen this in Phase 5.
- Credentials via [Secrets](07-secrets-manager.md); connect actions gated by [RBAC](08-rbac.md) and recorded in Audit.
- Substrates + workloads render as nodes in the [Service Graph](17-service-graph.md).
## Failure modes & handling
| Failure | Behavior |
|---|---|
| Substrate detected but no credentials | Recorded as `detected/unconnected`; surfaced for operator to attach a secret. No auto-connect. |
| Credentials invalid / connection refused | Target marked `unreachable` with reason; backoff retry; Installer excludes it from placement. |
| Substrate becomes unreachable | Bound `MCPInstance`s flagged for the reconciler; target drained, not deleted, until a grace period elapses. |
| Docker socket present but daemon dead | Distinguished from "no Docker" — reported as `error`, not absent, to avoid flapping. |
| Ambiguous host (both K8s node and Docker host) | Multiple targets registered, deduped by endpoint natural key; capabilities merged. |
| Privileged/host-network capability requested but substrate can't offer it | Capability advertised as false; Installer refuses placement rather than degrading isolation. |
## Security notes
Honors Architecture §9: substrate credentials live only in [Secrets](07-secrets-manager.md), are injected at connect time, and are never logged or written into Inventory. Auto-connect is **least-privilege** — Nexus requests the minimum scope (e.g. a read+deploy K8s role, not cluster-admin) and prefers local sockets over broad network API tokens. Connecting to and enumerating a substrate is a mutating, audited action. Capability flags (privileged, host-network) are explicit so the Installer can honor the sandboxing defaults; a substrate that cannot sandbox is not silently used.
## Open questions
- Is `RuntimeTarget` a distinct first-class domain object, or a specialization of `DiscoveredResource` with `source=infra`? (Leaning: distinct, sharing the store.)
- How much capacity/placement intelligence belongs here vs the Installer/reconciler?
- Hyper-V/VMware/VirtualBox are explicitly *not early* (Roadmap "not early") — ship as Phase 5 plugins?
- Credential discovery UX: how far do we go auto-detecting `KUBECONFIG`/SSH configs before prompting?
## Milestone
Delivered in **Phase 2** (Discover → install). Thin slice that lands first: **local Docker socket** detection + auto-connect, registered as a `RuntimeTarget`, with existing container enumeration forwarded to [Discovery Engine](01-discovery-engine.md). This is the substrate the Phase 2 exit proof installs onto. Kubernetes, Proxmox, SSH/bare-metal, and the VM hypervisors follow as additional probers.

View File

@@ -0,0 +1,182 @@
# Service Graph
> Module 17 · Plane: Cross-cutting · Roadmap phase: 5
> Part of [MCP Nexus architecture](../ARCHITECTURE.md).
## Purpose
The Service Graph builds and exposes a **live dependency graph** of everything
Nexus mediates, from the agent all the way down to the physical device:
```
Agent → Gateway → MCP server → real service → device
Claude → Gateway → Home Assistant MCP → Home Assistant → ESPHome → Light
```
It answers questions no single module can on its own: *"If Home Assistant goes
down, which agents and tools break?"*, *"What does `github.create_issue`
actually talk to?"*, *"Show me the full path from this agent to that light."*
The graph is **derived and observational** — it is assembled from
[Inventory](../ARCHITECTURE.md), routing data (the
[Dynamic Tool Registry](05-dynamic-tool-registry.md) tool→instance mapping),
and [Discovery](01-discovery-engine.md) relationships. It is explicitly **not
in the request hot path**: it never gates or slows a `tools/call`. It is a
read-only lens over data other modules already own.
## Responsibilities
- Assemble a typed node/edge graph spanning agents, the Gateway,
`MCPInstance`s, discovered real services, and downstream devices.
- **Infer edges** from existing data: routing (tool→instance), instance
binding (`bound_resource_uuid`), and Discovery-observed relationships
(service→device, service→dependency).
- Keep the graph current by subscribing to the same lifecycle/health/routing
events other data-plane modules emit.
- Expose a **query API** (subgraph, neighbors, paths, blast-radius) for the
Dashboard and external tooling.
- Feed the [Dashboard](12-web-dashboard.md) an interactive, real-time
visualization with health overlays.
- Annotate nodes/edges with health, confidence, and last-seen so the graph
doubles as an impact/topology map.
## Non-goals
- **Serving requests or routing** — that is the [Gateway](04-gateway.md) and
Router; the Service Graph observes them, never intercepts.
- **Being a source of truth** — it derives from Inventory / Registry /
Discovery and holds no independent authoritative state.
- **Discovering infrastructure** — that is [Discovery](01-discovery-engine.md)
and [Infra Discovery](16-infrastructure-discovery.md); the graph consumes
their output.
- **Alerting/notifying** — health transitions are owned by
[Health](09-health-monitoring.md) and [Notifications](13-notifications.md);
the graph only visualizes state.
## Interfaces
```go
type NodeType string
const (
NodeAgent NodeType = "agent" // AgentProfile
NodeGateway NodeType = "gateway" // the single edge
NodeInstance NodeType = "instance" // MCPInstance (the MCP server)
NodeService NodeType = "service" // DiscoveredResource (real service)
NodeDevice NodeType = "device" // leaf device (e.g. a light)
)
type EdgeType string
const (
EdgeConnects EdgeType = "connects" // agent → gateway (session)
EdgeRoutes EdgeType = "routes" // gateway → instance (tool dispatch)
EdgeBinds EdgeType = "binds" // instance → service (bound_resource_uuid)
EdgeDependsOn EdgeType = "depends_on" // service → service/device (discovery)
)
type Node struct {
ID string
Type NodeType
Label string
Health string // healthy | degraded | down | unknown
Confidence float64 // for discovery-derived nodes/edges
Ref string // uuid/instanceID/profileID it mirrors
}
type Edge struct {
From, To string
Type EdgeType
Health string
Since time.Time
}
// Graph is the read-only query surface. Rebuilt from events, never on the
// request path.
type Graph interface {
Snapshot(ctx context.Context, f Filter) (Nodes []Node, Edges []Edge, err error)
Neighbors(ctx context.Context, id string, depth int) ([]Node, []Edge, error)
Path(ctx context.Context, from, to string) ([]Edge, error) // e.g. agent→device
BlastRadius(ctx context.Context, id string) ([]Node, error) // what breaks if id fails
Subscribe(ctx context.Context) (<-chan Delta, func()) // live updates
}
```
HTTP / API surface (all RBAC-guarded, control-plane):
| Method / path | Purpose |
|---|---|
| `GET /graph` | Full or filtered graph snapshot (JSON nodes+edges). |
| `GET /graph/nodes/{id}/neighbors?depth=n` | Local subgraph around a node. |
| `GET /graph/path?from=&to=` | Concrete dependency path, e.g. agent→light. |
| `GET /graph/nodes/{id}/blast-radius` | Downstream/upstream impact set. |
| `GET /graph/stream` | SSE stream of graph deltas for live rendering. |
## Data
- **Derived graph (in-memory, cached):** nodes and typed edges, rebuilt from
source modules and updated incrementally from their event streams.
- **Edge inference sources:**
- `connects` — from active Gateway sessions (`AgentProfile` → Gateway).
- `routes` — from the Dynamic Tool Registry `{namespace}.{tool} → instanceID`
map (Gateway → `MCPInstance`).
- `binds` — from `MCPInstance.bound_resource_uuid` (instance →
`DiscoveredResource`).
- `depends_on` — from Discovery relationships and capability probing
(service → service, service → device); carries a **confidence** score
(tenet #6).
- **Persistence:** optional snapshotting to the Config/Inventory store for
historical topology; the runtime graph is a rebuildable projection, not a
primary store.
## Dependencies
- [Dynamic Tool Registry](05-dynamic-tool-registry.md) — `routes` edges.
- [Gateway](04-gateway.md) — `connects` edges (live sessions).
- [Discovery Engine](01-discovery-engine.md) /
[Infra Discovery](16-infrastructure-discovery.md) — services, devices, and
`depends_on` relationships with confidence.
- [Health Monitoring](09-health-monitoring.md) — node/edge health overlays.
- [AI Agent Profiles](14-agent-profiles.md) — agent nodes.
- [RBAC](08-rbac.md) — filters the graph to what the viewer may see.
- [Web Dashboard](12-web-dashboard.md) — consumer of the visualization.
## Failure modes & handling
- **Stale source data:** nodes/edges carry `last-seen`; the graph marks
entries `unknown` rather than deleting them on a single missed event, and
reconciles on periodic resync.
- **Source module unavailable:** the graph degrades gracefully — it renders the
partial graph it can derive and flags missing regions, never blocking.
- **Low-confidence inferred edges:** `depends_on` edges below the confidence
threshold are rendered distinctly (dashed) and excluded from
`BlastRadius` unless explicitly requested (tenet #6).
- **Large graphs:** the query API paginates/filters and supports depth-bounded
neighbor queries so the Dashboard never fetches the whole graph at once.
- **Event lag:** because it is off the hot path, transient inconsistency is
acceptable and self-heals on the next resync.
## Security notes
- The graph is **RBAC-filtered per viewer**: a user only sees agents,
instances, and services their role permits (Architecture §9). Blast-radius
and path queries respect the same filter.
- No secrets or credentials appear on nodes/edges — the graph references
services and devices by identity, never by connection secret.
- Graph queries are audited like other control-plane reads.
## Open questions
- How deep does device-level (`NodeDevice`) resolution go — do we model
ESPHome→individual-entity, or stop at the integration boundary?
- Do we retain historical topology snapshots for time-travel/diff views, and if
so where (Inventory vs. a dedicated store)?
- Should `depends_on` inference incorporate observed call traffic (from
Metrics) in addition to static discovery relationships?
## Milestone
Phase 5 (Extend & scale), alongside the [Plugin System](11-plugin-system.md)
and HA work. **Exit contribution:** the Dashboard renders a live
agent→gateway→MCP→service→device graph with health overlays and supports
path and blast-radius queries over the running fleet.

106
docs/modules/18-metrics.md Normal file
View File

@@ -0,0 +1,106 @@
# Metrics
> Module 18 · Plane: Cross-cutting · Roadmap phase: 4
> Part of [MCP Nexus architecture](../ARCHITECTURE.md).
## Purpose
Make Nexus observable. Metrics provides **Prometheus exposition** at `/metrics` and ships a set of **Grafana dashboards** (as JSON) so an operator can see request volume, latency, error rates, token usage, per-agent activity, installed-MCP and discovery counts, container/resource usage, and instance states at a glance. It is strictly **observational**: it records what other modules do and never changes behavior. Per tenet #4, exposition runs in-process in the single binary — no sidecar required — and Prometheus/Grafana remain optional external tools you point at Nexus.
## Responsibilities
- Own the process-wide metrics **registry** and the `GET /metrics` Prometheus text-exposition endpoint (OpenMetrics-compatible).
- Define and enforce the **metric taxonomy**: stable names, unit suffixes, and a controlled label set (bounded cardinality) that every module emits against.
- Provide thin emit helpers (counters, gauges, histograms) so modules record metrics without depending on the Prometheus client directly (keeps the client swappable and cardinality reviewable).
- Track the standard series: agent requests, tool-call latency, error rates, token usage, per-agent activity, installed MCP count, discovery counts, container/resource usage (CPU/mem per instance), and `MCPInstance` state gauges.
- Ship a curated **Grafana dashboard set** as JSON under `deploy/grafana/`, embedded and downloadable from the [Dashboard](12-web-dashboard.md).
- Expose Go runtime + build-info metrics (`go_*`, `nexus_build_info`) for baseline health.
## Non-goals
- **Not alerting.** Thresholds/alert routing are [Notifications](13-notifications.md) (event-driven) and Prometheus Alertmanager (metric-driven, external). Metrics only exposes series.
- **Not the audit log.** Audit is the immutable per-action record (§8); metrics are aggregate, lossy, and sampled. They answer different questions.
- **Not tracing.** Distributed tracing (OpenTelemetry spans) is a possible later addition, out of scope here.
- **Not health decisions.** [Health Monitoring](09-health-monitoring.md) owns state transitions and restart policy; it *emits* state gauges here.
- **Not long-term storage.** Nexus exposes; Prometheus scrapes and stores.
## Interfaces
```go
// Recorder is the narrow surface every module uses to emit metrics.
// Backed by the Prometheus client; kept minimal so the backend is swappable.
type Recorder interface {
Counter(name string, labels Labels) Counter
Gauge(name string, labels Labels) Gauge
Histogram(name string, labels Labels) Observer
}
type Labels map[string]string
type Observer interface{ Observe(v float64) }
type Counter interface{ Inc(); Add(float64) }
type Gauge interface{ Set(float64); Inc(); Dec() }
// Timer is sugar for latency histograms: defer t.ObserveDuration().
func StartTimer(h Observer) Timer
```
HTTP endpoint:
- `GET /metrics` — Prometheus/OpenMetrics exposition (behind Auth in hardened deployments; scrape token supported).
### Metric taxonomy (naming: `nexus_<subsystem>_<name>_<unit>`)
| Metric | Type | Key labels | Source module |
|---|---|---|---|
| `nexus_gateway_requests_total` | counter | `method`, `outcome` | [Gateway](04-gateway.md) |
| `nexus_toolcall_latency_seconds` | histogram | `namespace`, `tool`, `outcome` | [Router](05-dynamic-tool-registry.md) |
| `nexus_toolcall_errors_total` | counter | `namespace`, `tool`, `code` | Router |
| `nexus_tokens_total` | counter | `agent`, `direction` (in/out) | [Agent Profiles](14-agent-profiles.md) |
| `nexus_agent_activity_total` | counter | `agent`, `action` | Agent Profiles |
| `nexus_instances` | gauge | `state`, `package` | [Health](09-health-monitoring.md) |
| `nexus_mcp_installed` | gauge | `package` | [Installer](03-auto-installer.md) |
| `nexus_discovered_resources` | gauge | `type`, `confidence_bucket` | [Discovery](01-discovery-engine.md) |
| `nexus_container_cpu_seconds_total` | counter | `instance`, `package` | Runtime |
| `nexus_container_memory_bytes` | gauge | `instance`, `package` | Runtime |
| `nexus_updates_available` | gauge | `package` | [Update Manager](10-update-manager.md) |
| `nexus_notifications_sent_total` | counter | `channel`, `severity`, `outcome` | [Notifications](13-notifications.md) |
| `nexus_build_info` | gauge (=1) | `version`, `commit`, `go_version` | core |
**Cardinality rule:** high-cardinality identities (`agent`, `instance`) are allowed only on low-frequency series or via a bounded allowlist; free-form user input is never a label value. `confidence_bucket` discretizes the 0–1 score to keep Discovery cardinality flat.
**Method:** the request-path series follow the **RED** method (Rate, Errors, Duration) so the Overview dashboard reads as one story per subsystem; resource series follow **USE** (Utilization, Saturation, Errors) for containers/hosts. Histograms use fixed, documented buckets tuned to MCP tool-call latencies (sub-ms probes to multi-second model calls) so percentiles are comparable across instances and Grafana panels don't need per-panel re-bucketing.
## Data
- **Owns no store.** State lives in the in-memory Prometheus registry, reset on process restart (Prometheus persists the scraped history).
- **Reads** nothing persistent of its own; other modules push values through `Recorder`.
- Grafana dashboard JSON is a static, versioned asset shipped with the binary (`deploy/grafana/*.json`), not runtime state.
## Dependencies
- Emitted into by nearly every module (Gateway, Router, Discovery, Installer, Health, Update Manager, Notifications, Runtime, Agent Profiles) — it is cross-cutting by design.
- [Web Dashboard](12-web-dashboard.md) — the Metrics page renders selected series and links to the shipped Grafana dashboards.
- [Security](19-security.md) — `/metrics` auth + scrape-token policy.
- [Plugin System](11-plugin-system.md) — plugins receive a scoped `Recorder` so their metrics land in the same taxonomy.
## Grafana dashboard set
Shipped as JSON, importable or auto-provisioned:
1. **Overview** — request rate, p50/p95/p99 tool-call latency, error ratio, live instance count by state, updates available.
2. **Gateway & Router** — throughput, latency heatmap per namespace, error codes, connection-pool saturation.
3. **Agents** — per-agent activity, token usage (in/out), top tools per agent.
4. **Control plane** — discovery counts by type/confidence, installs over time, reconcile outcomes.
5. **Resources** — per-instance CPU/memory, container restarts (from Health), host headroom.
## Failure modes & handling
| Failure | Behavior |
|---|---|
| A module emits an unregistered/misnamed metric | Central taxonomy + emit helpers reject at registration; CI lints metric names. |
| Label cardinality explosion | Bounded label allowlist; free-form values are hashed/bucketed or dropped; cardinality budget checked in tests. |
| `/metrics` scrape is slow/large | Exposition is O(series); registry size is bounded by the taxonomy, so scrape cost stays flat. |
| Metrics registry contention under load | Client uses lock-free counters; emit is non-blocking and never on the request critical path's error path. |
| Prometheus not deployed | No effect — exposition is passive; Nexus functions fully without a scraper. |
## Security notes
Honors Architecture §9. `/metrics` can leak operational shape (instance names, package versions, agent identities), so in hardened deployments it is placed **behind Auth** and/or restricted to a scrape token / internal listener, never exposed alongside the public agent edge. No secret values, tokens, or raw request payloads are ever used as metric names or label values. Per-agent series use stable opaque IDs, not human PII. Access to `/metrics` and dashboard JSON is subject to [RBAC](08-rbac.md) when served through the dashboard.
## Open questions
- Do we adopt OpenTelemetry as the emit API now (metrics + future traces) or stay on the Prometheus client and bridge later?
- Auto-provision Grafana via its API on first run, or ship JSON for manual import only?
- Per-agent token accounting: exact per-call, or sampled to bound cardinality at high agent counts?
- Should `/metrics` default to authenticated even in the single-binary homelab topology, or open-on-loopback?
## Milestone
Delivered in **Phase 4** (Operate / day-2), though `/metrics` exists from **Phase 0** (exit criteria: `nexus serve` exposes `/healthz` and `/metrics`). Phase 4 fills the full taxonomy and ships the Grafana dashboard set. Exit proof (with Health + Notifications): killing a managed container is visible as an `nexus_instances{state="offline"}` blip and a container-restart counter increment on the Overview dashboard.

120
docs/modules/19-security.md Normal file
View File

@@ -0,0 +1,120 @@
# Security
> Module 19 · Plane: Cross-cutting · Roadmap phase: 3
> Part of [MCP Nexus architecture](../ARCHITECTURE.md).
## Purpose
Security is the umbrella model that makes tenet #5 ("secure by default") real and gives Architecture §9 its implementation. It is not one component but the **defense-in-depth composition** of the security-relevant modules — [Authentication](06-authentication.md), [RBAC](08-rbac.md), [Secrets](07-secrets-manager.md), [AI Agent Profiles](14-agent-profiles.md) — plus the cross-cutting controls no single module owns: **container sandboxing**, **append-only audit**, **signed-package/supply-chain verification**, **rate limiting**, and **TLS everywhere**. This doc defines those cross-cutting controls and how the layers stack so a failure of any one layer is not a breach. Reference: Architecture §9.
## Responsibilities
- **Sandbox every MCP instance.** Each `MCPInstance` runs in its own container with least privilege: dropped Linux capabilities, `no-new-privileges`, non-root user, **read-only rootfs** where the image allows, seccomp/AppArmor profiles, constrained networking (no host network unless the recipe demands it), and CPU/memory limits.
- **Container isolation defaults.** Define the baseline security context the [Auto Installer](03-auto-installer.md)/Runtime applies to every container so sandboxing is the default, not an opt-in.
- **Audit every action.** Provide the append-only audit sink and schema for *both* control-plane mutations and *every* agent tool call, recorded with principal, target, and outcome (Architecture §9).
- **Supply-chain verification.** Own the trust model for signed packages: verify signatures before install and **pin by digest** (never by mutable tag), in concert with the [Package Registry](02-package-registry.md).
- **Rate limiting.** Provide per-principal / per-profile / per-tool rate limits at the [Gateway](04-gateway.md) so one agent cannot starve others or abuse an upstream (tenet #1 fairness).
- **TLS everywhere.** Define TLS termination at the Gateway and localhost/socket-only internal component calls.
- **Compose the layers.** Specify how AuthN → Profiles → RBAC → Secrets → sandbox → audit stack into defense-in-depth, and how a bypass of one is contained by the next.
## Non-goals
- **Implementing the identity/policy engines.** AuthN ([06](06-authentication.md)), RBAC ([08](08-rbac.md)), Secrets ([07](07-secrets-manager.md)), Profiles ([14](14-agent-profiles.md)) own their mechanisms; this module composes them and owns the *cross-cutting* controls.
- **Running containers.** The [Auto Installer](03-auto-installer.md)/Runtime creates containers; this module defines the security context they must apply.
- **Being a SIEM.** Nexus emits an append-only audit log and metrics; long-term aggregation/alerting integrates via [Notifications](13-notifications.md) and external tooling.
- **Guaranteeing upstream service security.** Nexus sandboxes the MCP server; the real service behind it (Postgres, Home Assistant) has its own posture.
## Interfaces
```go
// The security context every managed container must be created with (Runtime applies it).
type SandboxSpec struct {
ReadOnlyRootfs bool
RunAsNonRoot bool
DropCaps []string // e.g. ["ALL"]; AddCaps only if a recipe justifies it
AddCaps []string
NoNewPrivs bool
Seccomp string // profile name/path
AppArmor string
Network NetMode // "none" | "bridge-scoped" | "host" (host requires justification)
CPU/*limits*/ Resource
Memory Resource
}
// Append-only audit sink (Architecture §8 Audit store). Records are immutable.
type AuditSink interface {
// Record is called for every mutating control-plane action AND every agent tool call.
Record(ctx context.Context, e AuditEvent) error
Query(ctx context.Context, f AuditFilter) ([]AuditEvent, error) // read-only
}
type AuditEvent struct {
Time time.Time
Principal string // who (user/agent/service)
Action string // "tools/call", "role.update", "secret.resolve", ...
Target string // "postgres.query", "role:developer", "secret://..."
Outcome string // "allow" | "deny" | "success" | "error"
Meta map[string]string // never contains secret plaintext
}
// Supply-chain verification (with the Package Registry).
type Verifier interface {
VerifySignature(ctx context.Context, ref string, sig Signature) error
ResolveDigest(ctx context.Context, ref string) (digest string, err error) // pin, never tag
}
// Rate limiting at the Gateway.
type RateLimiter interface {
Allow(ctx context.Context, key string) (bool, RetryAfter) // key = principal|profile|tool
}
```
HTTP/API surface:
- `GET /api/v1/audit?principal=&action=&target=&from=&to=` — query the append-only log (RBAC-gated; read-only, no delete/edit endpoint by design).
- `GET /api/v1/security/posture` — summary: TLS status, sandbox defaults, unsigned-package count, recent denials.
- `GET /metrics` — security-relevant counters (auth failures, denials, rate-limit drops) for [Metrics](18-metrics.md).
## Data
- **Writes/reads** the **Audit** store (Architecture §8) — append-only; no update/delete path exists in code, and integrity may be reinforced by a hash chain over events.
- **Reads** signature/trust anchors and digests from the [Package Registry](02-package-registry.md) before install.
- **Reads** the `SandboxSpec` defaults from the **Config** store; recipes may request (justified) deviations.
- Emits security counters to [Metrics](18-metrics.md) and security transitions to [Notifications](13-notifications.md).
## Dependencies
- Composes [Authentication](06-authentication.md), [RBAC](08-rbac.md), [Secrets](07-secrets-manager.md), and [AI Agent Profiles](14-agent-profiles.md) into defense-in-depth.
- Sandbox spec applied by the [Auto Installer](03-auto-installer.md)/Runtime.
- Signature/digest verification with the [Package Registry](02-package-registry.md).
- Rate limiting + TLS at the [Gateway](04-gateway.md).
- Security events flow to [Notifications](13-notifications.md) and counters to [Metrics](18-metrics.md).
## Failure modes & handling
| Failure | Behavior |
|---|---|
| Unsigned or signature-invalid package | Blocked before install; never becomes an `MCPInstance`; audited + alerted (tenet #5). |
| Image only offers a mutable tag | Resolved to a digest and pinned; if no digest can be resolved, install fails closed. |
| Container cannot run read-only rootfs | Recipe must declare the writable paths (tmpfs/volume) explicitly; unjustified full-write rootfs is refused. |
| Recipe requests host network / extra caps | Requires explicit justification in the recipe; flagged in `security/posture`; audited on install. |
| Audit sink unavailable | Mutations/tool calls fail closed (no silent unaudited actions) or buffer to a durable local queue — never proceed unlogged. |
| Rate limit exceeded | Request rejected with `429`/retry-after; drop counted; repeated abuse alerts. |
| TLS misconfigured / plaintext | Gateway refuses to serve the data plane in the clear; fails closed at boot. |
| A single layer bypassed (e.g. a tool leaks a name) | Next layer contains it: RBAC still denies the call, Secrets still redacts, audit still records — no single failure is a breach. |
## Security notes
This module *is* the security notes for the system; the layering is the point (Architecture §9):
1. **Transport** — TLS terminates at the Gateway; internal calls are localhost/socket only.
2. **Identity** — [Authentication](06-authentication.md) resolves every request to a `Principal`; fail closed if unresolved.
3. **Binding** — [AI Agent Profiles](14-agent-profiles.md) map an agent identity to roles.
4. **Authorization** — [RBAC](08-rbac.md) dual-gates `tools/list` and `tools/call`; default deny.
5. **Secrets** — [Secrets](07-secrets-manager.md) inject at runtime, never to agents/logs; envelope-encrypted at rest.
6. **Isolation** — every MCP runs least-privilege, sandboxed, network-constrained.
7. **Supply chain** — signature-verified, digest-pinned packages only.
8. **Observability** — every action append-only audited; rate-limited; metered.
Each layer assumes the ones above it can fail. Secure defaults are non-optional: unauthenticated ⇒ no access, ungranted ⇒ no tool, unsigned ⇒ no install, unlogged ⇒ no action.
## Open questions
- Audit integrity: is a hash-chained/tamper-evident log worth the write cost for v1, or is append-only-by-code enough?
- Do we ship default seccomp/AppArmor profiles per MCP archetype, or a single conservative baseline?
- Rate-limit granularity default: per-principal, per-profile, per-tool, or a composite — and where are limits configured?
- How do we attest the sandbox actually applied (verify the running container's security context vs. the spec)?
- Signing/trust roots for the community recipe & package ecosystem (mirrors Architecture §12).
## Milestone
Delivered in **Phase 3** (Secure). Thin slice first: **container sandboxing defaults** (non-root, dropped caps, read-only rootfs where possible, scoped network), **append-only audit** of every tool call + control-plane mutation, **signed-package verification with digest pinning**, **rate limiting**, and **TLS termination** at the Gateway — composing AuthN/RBAC/Secrets/Profiles into defense-in-depth. Underwrites the whole Phase 3 exit proof: differentiated multi-agent access with a secret-backed MCP that never leaks the secret to any agent-visible payload or log.

View File

@@ -0,0 +1,48 @@
# Future Vision
> Module 20 · Plane: — · Roadmap phase: ongoing
> Part of [MCP Nexus architecture](../ARCHITECTURE.md).
## Purpose
This is the north star the other nineteen modules serve. Where the engineering docs describe *how* Nexus reconciles discovery, install, health, and routing, this one describes *why it matters* and what the finished thing feels like — so every design decision can be checked against a single question: **does this move us closer to AI infrastructure that configures itself?**
MCP Nexus aims to be the **operating system for AI infrastructure**. Not another MCP server, and not a directory of servers — the layer that makes an entire environment's worth of services instantly, safely usable by any AI agent, with zero manual MCP wiring.
## The North Star
You install Nexus on one host. It scans the environment — the LAN, the Docker daemon, the Kubernetes cluster, the homelab rack, the enterprise subnet. It finds every compatible service: Home Assistant, Postgres, Grafana, Ollama, UniFi, GitHub, the NAS, the reverse proxy. For each, it downloads the correct MCP package, configures it against the discovered service, secures it, keeps it healthy, and keeps it updated. Then it exposes **one endpoint**.
Every AI agent you own — Claude, Cursor, an internal copilot, a fleet of autonomous workers — points at that one endpoint and, in that instant, gains governed access to the *entire* infrastructure. No per-agent MCP config. No hand-maintained server lists. No copy-pasted credentials. The agent asks for what it's allowed to do; Nexus decides, routes, and audits.
## The guiding metaphor: plug and play for AI tools
Think about plugging a USB device into a computer. You don't hunt for a driver, edit a config file, or restart. The OS detects the device, identifies it, loads the right driver, and it just works.
**Nexus is that experience for AI tools.** Stand up a new service on your network and — the way a USB device is detected, matched to a driver, and made available — Nexus discovers it, matches it to a recipe, installs the right MCP "driver," and the corresponding tools appear behind the single endpoint. The agent that connected yesterday can use the new service today without anyone touching its config. Unplug the service and the tools drain away just as cleanly.
That is the whole promise in one sentence: **infrastructure that AI agents can use should configure itself, the way peripherals already do.**
## What "done" feels like
- **For the homelabber:** install one binary, open the dashboard, watch your whole rack light up as discovered services become live tools. Point Claude Desktop at Nexus once; never edit an MCP config again.
- **For the enterprise operator:** a single, audited, least-privilege gateway between every AI agent and every internal system. Roles decide who sees what; secrets never touch an agent; every tool call is logged. Turning on a new team of agents is a role assignment, not an integration project.
- **For the agent:** one endpoint, a tool set that matches exactly what it's permitted to do, and tools that appear and disappear as the real world changes — no stale integrations, no missing capabilities.
- **For the ecosystem:** anyone can publish a discovery method, a recipe, or an MCP package, and every Nexus instance can pick it up — because everything is a plugin (tenet #3).
## Progression
The vision is not a leap; it is the roadmap's phases compounding:
1. **One endpoint proven** (Phases 1–2) — the core loop: discover a service and its MCP just appears behind one address. This is the differentiating claim, proven early.
2. **Safe for many agents** (Phase 3) — auth, RBAC, secrets, audit turn "it works" into "it's trustworthy in a shared environment."
3. **Runs unattended** (Phase 4) — health self-healing, updates, metrics, notifications, and the dashboard make it a system you can leave alone and still understand.
4. **Open and at scale** (Phase 5) — the plugin system and HA topologies turn Nexus from a product into a platform the community extends and enterprises depend on.
5. **The OS layer** (beyond) — richer autonomy: recipes that self-tune, discovery that reaches more substrates, a service graph that reasons about dependencies, and confidence-driven automation that safely does more on its own over time.
Each phase already ends in something demonstrable; the north star is simply what they add up to.
## Risks & unknowns
- **Discovery is probabilistic, not certain** (tenet #6). "Plug and play" must stay honest: high-confidence detections auto-act; ambiguous ones ask a human. Over-automating erodes trust faster than under-automating.
- **Recipe & package trust.** A self-installing system is only as safe as its supply chain. Community recipes and packages need signing, governance, and a real trust model (Architecture §12) before "downloads the correct package" can be fully hands-off.
- **The USB metaphor has limits.** Real services have credentials, network policy, and blast radius a USB stick never does. Security-by-default and least privilege are what keep the convenience from becoming a liability.
- **Breadth vs. depth.** The value scales with how many services are recognized — but every new fingerprinter/recipe is surface area to maintain. The plugin system exists so breadth can grow without the core growing.
- **Autonomy pacing.** How much Nexus should *do on its own* versus *propose* is a dial, not a constant; it should move toward more autonomy only as confidence, auditability, and rollback prove themselves.
## Milestone
Ongoing — this doc has no ship date; it is the standard the shippable phases are measured against. The nearest concrete embodiment is the Phase 2 exit: **start a service on the LAN and watch its tools appear behind the single endpoint with zero manual MCP config.** That is the north star in miniature. Everything after makes it secure, unattended, extensible, and universal.