MCP Nexus is a self-hosted control plane for MCP servers — "Kubernetes for
MCP." Agents connect to one endpoint; Nexus discovers services, installs the
right MCP servers, aggregates them behind a namespaced router, secures access
with auth/RBAC, and heals/updates them via a reconcile loop.
This first commit is design-phase only (no runnable code yet):
- README.md project front door + module map
- docs/ARCHITECTURE.md target design: tenets, two-plane split, reconcile
loop, domain types, storage, security, deployment
- docs/ROADMAP.md phased delivery (foundations -> walking skeleton ->
discover+install -> secure -> operate -> extend/scale)
- docs/modules/01-20 one design doc per module, all cross-linked
Backend stack decision: Go (single static binary, embedded SQLite + embedded
React dashboard). Repo initialized in /root with a whitelist .gitignore so
only project files are tracked.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
9.7 KiB
Health Monitoring
Module 09 · Plane: Control · Roadmap phase: 4 Part of MCP Nexus architecture.
Purpose
Keep every managed MCPInstance alive and serving. Health Monitoring continuously probes each instance/container, tracks its lifecycle state (running, offline, updating, restarting, quarantined), measures latency and error rates, and automatically heals unhealthy instances via restart-with-backoff, escalating to quarantine after repeated failure. It is the act/record feedback arm of the reconciliation loop (Architecture §6): it turns observed instance health into events the reconciler, Router, Update Manager, and Notifications react to. A failure of the control plane must never take down the data plane, so Health only signals the Router which instances are poolable — it never sits in the hot path.
Responsibilities
- Run liveness and readiness probes against every
MCPInstanceon a per-instance schedule: liveness = "the process/container is up"; readiness = "it completes an MCPinitialize+tools/listwithin budget." - Collect health signals from three layers: container (Docker inspect: state, restart count, OOM/exit codes), transport (MCP client connect success, RTT), and application (probe tool-call latency, JSON-RPC error rate over a sliding window).
- Own the per-instance health state machine and drive transitions; persist current state + last transition into Inventory.
- Execute the restart policy: on failed liveness, restart the instance via the Runtime with exponential backoff and jitter; quarantine after N consecutive failed restarts.
- Emit health-transition events on the event bus so: the Router drains dead instances, the Update Manager can roll back a bad update, and Notifications fire on
offline/quarantined. - Expose current + historical health via the control API for the Dashboard and feed gauges/counters to Metrics.
Non-goals
- Does not pull, create, or destroy containers — it requests restart/recreate through the Runtime + Auto Installer; Runtime owns container lifecycle.
- Does not decide desired state. Whether an instance should exist is the reconciler's call from Recipes ⨯ Inventory; Health only reports and heals what exists.
- Does not route traffic. It signals poolability; the Gateway/Router owns
UpstreamPool.Drain. - Does not define probe transport plugins — those interfaces are generalized by the Plugin System in Phase 5.
- Does not send messages — it emits events; delivery is Notifications.
Interfaces
// HealthState is the per-instance lifecycle state.
type HealthState string
const (
StateRunning HealthState = "running"
StateOffline HealthState = "offline" // liveness failing
StateUpdating HealthState = "updating" // owned transiently by Update Manager
StateRestarting HealthState = "restarting" // heal in progress
StateQuarantined HealthState = "quarantined" // gave up after repeated restarts
)
// Probe is a pluggable health check against one instance.
type Probe interface {
Name() string // "container", "mcp-readiness", "tool-latency"
Check(ctx context.Context, inst *MCPInstance) ProbeResult
}
type ProbeResult struct {
Healthy bool
Latency time.Duration
ErrRate float64 // over the probe's sliding window
Reason string // human-readable on failure
Kind ProbeKind // Liveness | Readiness
}
// Monitor owns probing, the state machine, and the restart policy.
type Monitor interface {
Register(p Probe) error
Watch(ctx context.Context) error // periodic + edge-triggered
Status(instanceID string) (HealthReport, bool) // current snapshot
Quarantine(ctx context.Context, instanceID, reason string) error
Release(ctx context.Context, instanceID string) error // un-quarantine
}
// RestartPolicy computes backoff and the quarantine threshold.
type RestartPolicy struct {
Base time.Duration // e.g. 1s
Max time.Duration // cap, e.g. 5m
Multiplier float64 // e.g. 2.0
Jitter float64 // 0..1
MaxRestarts int // consecutive failures → quarantine
HealthyReset time.Duration // uptime after which the failure count resets
}
Internal HTTP (control API, RBAC-guarded; not agent-facing):
GET /api/v1/health/instances— list instances with state, latency, error rate, restart count.GET /api/v1/health/instances/{id}— fullHealthReport+ transition history.POST /api/v1/health/instances/{id}/restart— operator-forced restart.POST /api/v1/health/instances/{id}/quarantine·/release— manual quarantine controls.
State machine
probe ok
┌──────────────────────────────┐
▼ │
running ──liveness fail──▶ offline ──heal──▶ restarting ──ready──▶ running
│ │ │
│ update begins │ │ restart fails × MaxRestarts
▼ ▼ ▼
updating ──(rollback path)──▶ ◀──────────── quarantined ──operator release──▶ restarting
updating is entered/left by the Update Manager; Health suppresses auto-restart while an instance is updating and instead reports readiness so the Updater can decide rollback on unhealthy. quarantined instances are drained from the Router and require operator release (or a successful update) to re-enter the loop.
Data
- Writes
health+stateon eachMCPInstancein the Inventory store (§8), plus an append-only transition log (state, reason, timestamp) for the dashboard timeline. - Reads instance definitions from Inventory and probe/policy config from the Config store; probe credentials (if any) by ref from Secrets.
- Module-local in-memory ring buffers hold recent latency/error samples per instance for sliding-window rate computation (survive-restart not required; recomputed on boot).
Dependencies
- Auto Installer + Runtime — executes restart/recreate on Health's request.
- Gateway / Router — consumes health events via
UpstreamPool.Drain; stops routing to non-runninginstances. - Update Manager — reads readiness to gate promotion and trigger rollback on unhealthy; owns the
updatingstate. - Notifications — fires on
offline/quarantined/recovered transitions. - Metrics — exports state, latency, error-rate, and restart-count series.
- Plugin System — generalizes the
Probeinterface for custom checks.
Failure modes & handling
| Failure | Behavior |
|---|---|
| Instance fails liveness | Transition running→offline, drain from Router, restart with backoff (restarting). |
| Restarts keep failing | After MaxRestarts consecutive failures → quarantined; stop restarting; notify; require operator/update to recover. |
| Flapping (ready→fail→ready) | Backoff prevents restart storms; failure count only resets after HealthyReset sustained uptime. |
| Probe itself times out / Docker API down | Probe error is not the same as instance-down: mark unknown, retry with backoff, do not restart on probe infrastructure failure alone. |
| Update in progress | Auto-restart suppressed; Health reports readiness to the Updater, which owns rollback. |
| Health monitor crashes | Data plane unaffected (Router keeps last-known poolable set until TTL); on restart, probes re-establish state from Inventory. |
| False-positive readiness | Application probe (real tool-call latency + error rate), not just container-up, reduces "up but broken" routing. |
Security notes
Honors Architecture §9. Probes are least-privilege: readiness uses the same sandboxed MCP client path as the Router, never a privileged shell into the container. Probe credentials come from Secrets by ref and are never logged or written into HealthReport. Failure Reason strings are sanitized before they reach the dashboard or notifications so upstream error payloads cannot leak secrets or internal topology. Every restart/quarantine/release is a mutating action and is audited with principal (or system:health) + target + outcome. Restart requests to the Runtime are the only privileged action and flow through the same audited Installer path.
Open questions
- Should backoff/quarantine thresholds be per-
type(a flaky IoT MCP vs. a rack Postgres MCP differ), overridable per Recipe? - Do we expose raw probe latency histograms to the dashboard, or only rolled-up percentiles (topology-leak risk, per Discovery's parallel question)?
- Quarantine auto-recovery: should Nexus periodically re-probe quarantined instances with a long backoff, or strictly require operator/update release?
- How do readiness budgets interact with legitimately slow-starting MCP servers (model loads, large indexes)?
Milestone
Delivered in Phase 4 (Operate / day-2). Thin slice: container + MCP-readiness probes, the five-state machine, exponential-backoff auto-restart, and quarantine-after-N, wired to Router drain and Notifications. Exit proof (shared with the phase goal): kill a managed MCP container → it is auto-restarted, a Discord alert fires, and the dashboard reflects running→offline→restarting→running in real time.