docs: initial architecture and design for MCP Nexus

MCP Nexus is a self-hosted control plane for MCP servers — "Kubernetes for
MCP." Agents connect to one endpoint; Nexus discovers services, installs the
right MCP servers, aggregates them behind a namespaced router, secures access
with auth/RBAC, and heals/updates them via a reconcile loop.

This first commit is design-phase only (no runnable code yet):
- README.md            project front door + module map
- docs/ARCHITECTURE.md target design: tenets, two-plane split, reconcile
                       loop, domain types, storage, security, deployment
- docs/ROADMAP.md      phased delivery (foundations -> walking skeleton ->
                       discover+install -> secure -> operate -> extend/scale)
- docs/modules/01-20   one design doc per module, all cross-linked

Backend stack decision: Go (single static binary, embedded SQLite + embedded
React dashboard). Repo initialized in /root with a whitelist .gitignore so
only project files are tracked.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
drjones
2026-07-07 04:38:27 +00:00
commit 8d3ffef920
24 changed files with 3000 additions and 0 deletions

View File

@@ -0,0 +1,124 @@
# Health Monitoring
> Module 09 · Plane: Control · Roadmap phase: 4
> Part of [MCP Nexus architecture](../ARCHITECTURE.md).
## Purpose
Keep every managed [`MCPInstance`](../ARCHITECTURE.md) alive and serving. Health Monitoring continuously probes each instance/container, tracks its lifecycle state (`running`, `offline`, `updating`, `restarting`, `quarantined`), measures latency and error rates, and **automatically heals** unhealthy instances via restart-with-backoff, escalating to quarantine after repeated failure. It is the `act`/`record` feedback arm of the reconciliation loop (Architecture §6): it turns observed instance health into events the reconciler, Router, Update Manager, and Notifications react to. A failure of the control plane must never take down the data plane, so Health only *signals* the Router which instances are poolable — it never sits in the hot path.
## Responsibilities
- Run **liveness** and **readiness** probes against every `MCPInstance` on a per-instance schedule: liveness = "the process/container is up"; readiness = "it completes an MCP `initialize` + `tools/list` within budget."
- Collect health signals from three layers: **container** (Docker inspect: state, restart count, OOM/exit codes), **transport** (MCP client connect success, RTT), and **application** (probe tool-call latency, JSON-RPC error rate over a sliding window).
- Own the per-instance **health state machine** and drive transitions; persist current state + last transition into Inventory.
- Execute the **restart policy**: on failed liveness, restart the instance via the Runtime with exponential backoff and jitter; **quarantine** after N consecutive failed restarts.
- Emit health-transition events on the event bus so: the [Router](04-gateway.md) drains dead instances, the [Update Manager](10-update-manager.md) can roll back a bad update, and [Notifications](13-notifications.md) fire on `offline`/`quarantined`.
- Expose current + historical health via the control API for the [Dashboard](12-web-dashboard.md) and feed gauges/counters to [Metrics](18-metrics.md).
## Non-goals
- **Does not pull, create, or destroy containers** — it *requests* restart/recreate through the Runtime + [Auto Installer](03-auto-installer.md); Runtime owns container lifecycle.
- **Does not decide desired state.** Whether an instance *should* exist is the reconciler's call from Recipes ⨯ Inventory; Health only reports and heals what exists.
- **Does not route traffic.** It signals poolability; the [Gateway](04-gateway.md)/Router owns `UpstreamPool.Drain`.
- **Does not define probe transport plugins** — those interfaces are generalized by the [Plugin System](11-plugin-system.md) in Phase 5.
- **Does not send messages** — it emits events; delivery is [Notifications](13-notifications.md).
## Interfaces
```go
// HealthState is the per-instance lifecycle state.
type HealthState string
const (
StateRunning HealthState = "running"
StateOffline HealthState = "offline" // liveness failing
StateUpdating HealthState = "updating" // owned transiently by Update Manager
StateRestarting HealthState = "restarting" // heal in progress
StateQuarantined HealthState = "quarantined" // gave up after repeated restarts
)
// Probe is a pluggable health check against one instance.
type Probe interface {
Name() string // "container", "mcp-readiness", "tool-latency"
Check(ctx context.Context, inst *MCPInstance) ProbeResult
}
type ProbeResult struct {
Healthy bool
Latency time.Duration
ErrRate float64 // over the probe's sliding window
Reason string // human-readable on failure
Kind ProbeKind // Liveness | Readiness
}
// Monitor owns probing, the state machine, and the restart policy.
type Monitor interface {
Register(p Probe) error
Watch(ctx context.Context) error // periodic + edge-triggered
Status(instanceID string) (HealthReport, bool) // current snapshot
Quarantine(ctx context.Context, instanceID, reason string) error
Release(ctx context.Context, instanceID string) error // un-quarantine
}
// RestartPolicy computes backoff and the quarantine threshold.
type RestartPolicy struct {
Base time.Duration // e.g. 1s
Max time.Duration // cap, e.g. 5m
Multiplier float64 // e.g. 2.0
Jitter float64 // 0..1
MaxRestarts int // consecutive failures → quarantine
HealthyReset time.Duration // uptime after which the failure count resets
}
```
Internal HTTP (control API, RBAC-guarded; not agent-facing):
- `GET /api/v1/health/instances` — list instances with state, latency, error rate, restart count.
- `GET /api/v1/health/instances/{id}` — full `HealthReport` + transition history.
- `POST /api/v1/health/instances/{id}/restart` — operator-forced restart.
- `POST /api/v1/health/instances/{id}/quarantine` · `/release` — manual quarantine controls.
## State machine
```
probe ok
┌──────────────────────────────┐
▼ │
running ──liveness fail──▶ offline ──heal──▶ restarting ──ready──▶ running
│ │ │
│ update begins │ │ restart fails × MaxRestarts
▼ ▼ ▼
updating ──(rollback path)──▶ ◀──────────── quarantined ──operator release──▶ restarting
```
`updating` is entered/left by the [Update Manager](10-update-manager.md); Health suppresses auto-restart while an instance is `updating` and instead reports readiness so the Updater can decide **rollback on unhealthy**. `quarantined` instances are drained from the Router and require operator release (or a successful update) to re-enter the loop.
## Data
- **Writes** `health` + `state` on each `MCPInstance` in the **Inventory** store (§8), plus an append-only transition log (state, reason, timestamp) for the dashboard timeline.
- **Reads** instance definitions from Inventory and probe/policy config from the **Config** store; probe credentials (if any) by ref from [Secrets](07-secrets-manager.md).
- Module-local in-memory ring buffers hold recent latency/error samples per instance for sliding-window rate computation (survive-restart not required; recomputed on boot).
## Dependencies
- [Auto Installer](03-auto-installer.md) + Runtime — executes restart/recreate on Health's request.
- [Gateway](04-gateway.md) / Router — consumes health events via `UpstreamPool.Drain`; stops routing to non-`running` instances.
- [Update Manager](10-update-manager.md) — reads readiness to gate promotion and trigger **rollback on unhealthy**; owns the `updating` state.
- [Notifications](13-notifications.md) — fires on `offline`/`quarantined`/recovered transitions.
- [Metrics](18-metrics.md) — exports state, latency, error-rate, and restart-count series.
- [Plugin System](11-plugin-system.md) — generalizes the `Probe` interface for custom checks.
## Failure modes & handling
| Failure | Behavior |
|---|---|
| Instance fails liveness | Transition `running→offline`, drain from Router, restart with backoff (`restarting`). |
| Restarts keep failing | After `MaxRestarts` consecutive failures → `quarantined`; stop restarting; notify; require operator/update to recover. |
| Flapping (ready→fail→ready) | Backoff prevents restart storms; failure count only resets after `HealthyReset` sustained uptime. |
| Probe itself times out / Docker API down | Probe error is not the same as instance-down: mark `unknown`, retry with backoff, do **not** restart on probe infrastructure failure alone. |
| Update in progress | Auto-restart suppressed; Health reports readiness to the Updater, which owns rollback. |
| Health monitor crashes | Data plane unaffected (Router keeps last-known poolable set until TTL); on restart, probes re-establish state from Inventory. |
| False-positive readiness | Application probe (real tool-call latency + error rate), not just container-up, reduces "up but broken" routing. |
## Security notes
Honors Architecture §9. Probes are least-privilege: readiness uses the same sandboxed MCP client path as the Router, never a privileged shell into the container. Probe credentials come from [Secrets](07-secrets-manager.md) by ref and are never logged or written into `HealthReport`. Failure `Reason` strings are sanitized before they reach the dashboard or notifications so upstream error payloads cannot leak secrets or internal topology. Every restart/quarantine/release is a mutating action and is **audited** with principal (or `system:health`) + target + outcome. Restart requests to the Runtime are the only privileged action and flow through the same audited Installer path.
## Open questions
- Should backoff/quarantine thresholds be per-`type` (a flaky IoT MCP vs. a rack Postgres MCP differ), overridable per Recipe?
- Do we expose raw probe latency histograms to the dashboard, or only rolled-up percentiles (topology-leak risk, per Discovery's parallel question)?
- Quarantine auto-recovery: should Nexus periodically re-probe quarantined instances with a long backoff, or strictly require operator/update release?
- How do readiness budgets interact with legitimately slow-starting MCP servers (model loads, large indexes)?
## Milestone
Delivered in **Phase 4** (Operate / day-2). Thin slice: container + MCP-readiness probes, the five-state machine, exponential-backoff auto-restart, and quarantine-after-N, wired to Router drain and Notifications. Exit proof (shared with the phase goal): kill a managed MCP container → it is auto-restarted, a Discord alert fires, and the dashboard reflects `running→offline→restarting→running` in real time.