Files
mcp-gateway-nexus/docs/modules/09-health-monitoring.md
drjones 8d3ffef920 docs: initial architecture and design for MCP Nexus
MCP Nexus is a self-hosted control plane for MCP servers — "Kubernetes for
MCP." Agents connect to one endpoint; Nexus discovers services, installs the
right MCP servers, aggregates them behind a namespaced router, secures access
with auth/RBAC, and heals/updates them via a reconcile loop.

This first commit is design-phase only (no runnable code yet):
- README.md            project front door + module map
- docs/ARCHITECTURE.md target design: tenets, two-plane split, reconcile
                       loop, domain types, storage, security, deployment
- docs/ROADMAP.md      phased delivery (foundations -> walking skeleton ->
                       discover+install -> secure -> operate -> extend/scale)
- docs/modules/01-20   one design doc per module, all cross-linked

Backend stack decision: Go (single static binary, embedded SQLite + embedded
React dashboard). Repo initialized in /root with a whitelist .gitignore so
only project files are tracked.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-07 04:38:27 +00:00

9.7 KiB
Raw Permalink Blame History

Health Monitoring

Module 09 · Plane: Control · Roadmap phase: 4 Part of MCP Nexus architecture.

Purpose

Keep every managed MCPInstance alive and serving. Health Monitoring continuously probes each instance/container, tracks its lifecycle state (running, offline, updating, restarting, quarantined), measures latency and error rates, and automatically heals unhealthy instances via restart-with-backoff, escalating to quarantine after repeated failure. It is the act/record feedback arm of the reconciliation loop (Architecture §6): it turns observed instance health into events the reconciler, Router, Update Manager, and Notifications react to. A failure of the control plane must never take down the data plane, so Health only signals the Router which instances are poolable — it never sits in the hot path.

Responsibilities

  • Run liveness and readiness probes against every MCPInstance on a per-instance schedule: liveness = "the process/container is up"; readiness = "it completes an MCP initialize + tools/list within budget."
  • Collect health signals from three layers: container (Docker inspect: state, restart count, OOM/exit codes), transport (MCP client connect success, RTT), and application (probe tool-call latency, JSON-RPC error rate over a sliding window).
  • Own the per-instance health state machine and drive transitions; persist current state + last transition into Inventory.
  • Execute the restart policy: on failed liveness, restart the instance via the Runtime with exponential backoff and jitter; quarantine after N consecutive failed restarts.
  • Emit health-transition events on the event bus so: the Router drains dead instances, the Update Manager can roll back a bad update, and Notifications fire on offline/quarantined.
  • Expose current + historical health via the control API for the Dashboard and feed gauges/counters to Metrics.

Non-goals

  • Does not pull, create, or destroy containers — it requests restart/recreate through the Runtime + Auto Installer; Runtime owns container lifecycle.
  • Does not decide desired state. Whether an instance should exist is the reconciler's call from Recipes ⨯ Inventory; Health only reports and heals what exists.
  • Does not route traffic. It signals poolability; the Gateway/Router owns UpstreamPool.Drain.
  • Does not define probe transport plugins — those interfaces are generalized by the Plugin System in Phase 5.
  • Does not send messages — it emits events; delivery is Notifications.

Interfaces

// HealthState is the per-instance lifecycle state.
type HealthState string

const (
    StateRunning     HealthState = "running"
    StateOffline     HealthState = "offline"     // liveness failing
    StateUpdating    HealthState = "updating"    // owned transiently by Update Manager
    StateRestarting  HealthState = "restarting"  // heal in progress
    StateQuarantined HealthState = "quarantined" // gave up after repeated restarts
)

// Probe is a pluggable health check against one instance.
type Probe interface {
    Name() string // "container", "mcp-readiness", "tool-latency"
    Check(ctx context.Context, inst *MCPInstance) ProbeResult
}

type ProbeResult struct {
    Healthy   bool
    Latency   time.Duration
    ErrRate   float64 // over the probe's sliding window
    Reason    string  // human-readable on failure
    Kind      ProbeKind // Liveness | Readiness
}

// Monitor owns probing, the state machine, and the restart policy.
type Monitor interface {
    Register(p Probe) error
    Watch(ctx context.Context) error                 // periodic + edge-triggered
    Status(instanceID string) (HealthReport, bool)   // current snapshot
    Quarantine(ctx context.Context, instanceID, reason string) error
    Release(ctx context.Context, instanceID string) error // un-quarantine
}

// RestartPolicy computes backoff and the quarantine threshold.
type RestartPolicy struct {
    Base       time.Duration // e.g. 1s
    Max        time.Duration // cap, e.g. 5m
    Multiplier float64       // e.g. 2.0
    Jitter     float64       // 0..1
    MaxRestarts int          // consecutive failures → quarantine
    HealthyReset time.Duration // uptime after which the failure count resets
}

Internal HTTP (control API, RBAC-guarded; not agent-facing):

  • GET /api/v1/health/instances — list instances with state, latency, error rate, restart count.
  • GET /api/v1/health/instances/{id} — full HealthReport + transition history.
  • POST /api/v1/health/instances/{id}/restart — operator-forced restart.
  • POST /api/v1/health/instances/{id}/quarantine · /release — manual quarantine controls.

State machine

                 probe ok
   ┌──────────────────────────────┐
   ▼                              │
 running ──liveness fail──▶ offline ──heal──▶ restarting ──ready──▶ running
   │                          │                   │
   │ update begins            │                   │ restart fails × MaxRestarts
   ▼                          ▼                   ▼
 updating ──(rollback path)──▶ ◀──────────── quarantined ──operator release──▶ restarting

updating is entered/left by the Update Manager; Health suppresses auto-restart while an instance is updating and instead reports readiness so the Updater can decide rollback on unhealthy. quarantined instances are drained from the Router and require operator release (or a successful update) to re-enter the loop.

Data

  • Writes health + state on each MCPInstance in the Inventory store (§8), plus an append-only transition log (state, reason, timestamp) for the dashboard timeline.
  • Reads instance definitions from Inventory and probe/policy config from the Config store; probe credentials (if any) by ref from Secrets.
  • Module-local in-memory ring buffers hold recent latency/error samples per instance for sliding-window rate computation (survive-restart not required; recomputed on boot).

Dependencies

  • Auto Installer + Runtime — executes restart/recreate on Health's request.
  • Gateway / Router — consumes health events via UpstreamPool.Drain; stops routing to non-running instances.
  • Update Manager — reads readiness to gate promotion and trigger rollback on unhealthy; owns the updating state.
  • Notifications — fires on offline/quarantined/recovered transitions.
  • Metrics — exports state, latency, error-rate, and restart-count series.
  • Plugin System — generalizes the Probe interface for custom checks.

Failure modes & handling

Failure Behavior
Instance fails liveness Transition running→offline, drain from Router, restart with backoff (restarting).
Restarts keep failing After MaxRestarts consecutive failures → quarantined; stop restarting; notify; require operator/update to recover.
Flapping (ready→fail→ready) Backoff prevents restart storms; failure count only resets after HealthyReset sustained uptime.
Probe itself times out / Docker API down Probe error is not the same as instance-down: mark unknown, retry with backoff, do not restart on probe infrastructure failure alone.
Update in progress Auto-restart suppressed; Health reports readiness to the Updater, which owns rollback.
Health monitor crashes Data plane unaffected (Router keeps last-known poolable set until TTL); on restart, probes re-establish state from Inventory.
False-positive readiness Application probe (real tool-call latency + error rate), not just container-up, reduces "up but broken" routing.

Security notes

Honors Architecture §9. Probes are least-privilege: readiness uses the same sandboxed MCP client path as the Router, never a privileged shell into the container. Probe credentials come from Secrets by ref and are never logged or written into HealthReport. Failure Reason strings are sanitized before they reach the dashboard or notifications so upstream error payloads cannot leak secrets or internal topology. Every restart/quarantine/release is a mutating action and is audited with principal (or system:health) + target + outcome. Restart requests to the Runtime are the only privileged action and flow through the same audited Installer path.

Open questions

  • Should backoff/quarantine thresholds be per-type (a flaky IoT MCP vs. a rack Postgres MCP differ), overridable per Recipe?
  • Do we expose raw probe latency histograms to the dashboard, or only rolled-up percentiles (topology-leak risk, per Discovery's parallel question)?
  • Quarantine auto-recovery: should Nexus periodically re-probe quarantined instances with a long backoff, or strictly require operator/update release?
  • How do readiness budgets interact with legitimately slow-starting MCP servers (model loads, large indexes)?

Milestone

Delivered in Phase 4 (Operate / day-2). Thin slice: container + MCP-readiness probes, the five-state machine, exponential-backoff auto-restart, and quarantine-after-N, wired to Router drain and Notifications. Exit proof (shared with the phase goal): kill a managed MCP container → it is auto-restarted, a Discord alert fires, and the dashboard reflects running→offline→restarting→running in real time.