MCP Nexus is a self-hosted control plane for MCP servers — "Kubernetes for
MCP." Agents connect to one endpoint; Nexus discovers services, installs the
right MCP servers, aggregates them behind a namespaced router, secures access
with auth/RBAC, and heals/updates them via a reconcile loop.
This first commit is design-phase only (no runnable code yet):
- README.md project front door + module map
- docs/ARCHITECTURE.md target design: tenets, two-plane split, reconcile
loop, domain types, storage, security, deployment
- docs/ROADMAP.md phased delivery (foundations -> walking skeleton ->
discover+install -> secure -> operate -> extend/scale)
- docs/modules/01-20 one design doc per module, all cross-linked
Backend stack decision: Go (single static binary, embedded SQLite + embedded
React dashboard). Repo initialized in /root with a whitelist .gitignore so
only project files are tracked.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
125 lines
9.7 KiB
Markdown
125 lines
9.7 KiB
Markdown
# Health Monitoring
|
||
|
||
> Module 09 · Plane: Control · Roadmap phase: 4
|
||
> Part of [MCP Nexus architecture](../ARCHITECTURE.md).
|
||
|
||
## Purpose
|
||
Keep every managed [`MCPInstance`](../ARCHITECTURE.md) alive and serving. Health Monitoring continuously probes each instance/container, tracks its lifecycle state (`running`, `offline`, `updating`, `restarting`, `quarantined`), measures latency and error rates, and **automatically heals** unhealthy instances via restart-with-backoff, escalating to quarantine after repeated failure. It is the `act`/`record` feedback arm of the reconciliation loop (Architecture §6): it turns observed instance health into events the reconciler, Router, Update Manager, and Notifications react to. A failure of the control plane must never take down the data plane, so Health only *signals* the Router which instances are poolable — it never sits in the hot path.
|
||
|
||
## Responsibilities
|
||
- Run **liveness** and **readiness** probes against every `MCPInstance` on a per-instance schedule: liveness = "the process/container is up"; readiness = "it completes an MCP `initialize` + `tools/list` within budget."
|
||
- Collect health signals from three layers: **container** (Docker inspect: state, restart count, OOM/exit codes), **transport** (MCP client connect success, RTT), and **application** (probe tool-call latency, JSON-RPC error rate over a sliding window).
|
||
- Own the per-instance **health state machine** and drive transitions; persist current state + last transition into Inventory.
|
||
- Execute the **restart policy**: on failed liveness, restart the instance via the Runtime with exponential backoff and jitter; **quarantine** after N consecutive failed restarts.
|
||
- Emit health-transition events on the event bus so: the [Router](04-gateway.md) drains dead instances, the [Update Manager](10-update-manager.md) can roll back a bad update, and [Notifications](13-notifications.md) fire on `offline`/`quarantined`.
|
||
- Expose current + historical health via the control API for the [Dashboard](12-web-dashboard.md) and feed gauges/counters to [Metrics](18-metrics.md).
|
||
|
||
## Non-goals
|
||
- **Does not pull, create, or destroy containers** — it *requests* restart/recreate through the Runtime + [Auto Installer](03-auto-installer.md); Runtime owns container lifecycle.
|
||
- **Does not decide desired state.** Whether an instance *should* exist is the reconciler's call from Recipes ⨯ Inventory; Health only reports and heals what exists.
|
||
- **Does not route traffic.** It signals poolability; the [Gateway](04-gateway.md)/Router owns `UpstreamPool.Drain`.
|
||
- **Does not define probe transport plugins** — those interfaces are generalized by the [Plugin System](11-plugin-system.md) in Phase 5.
|
||
- **Does not send messages** — it emits events; delivery is [Notifications](13-notifications.md).
|
||
|
||
## Interfaces
|
||
```go
|
||
// HealthState is the per-instance lifecycle state.
|
||
type HealthState string
|
||
|
||
const (
|
||
StateRunning HealthState = "running"
|
||
StateOffline HealthState = "offline" // liveness failing
|
||
StateUpdating HealthState = "updating" // owned transiently by Update Manager
|
||
StateRestarting HealthState = "restarting" // heal in progress
|
||
StateQuarantined HealthState = "quarantined" // gave up after repeated restarts
|
||
)
|
||
|
||
// Probe is a pluggable health check against one instance.
|
||
type Probe interface {
|
||
Name() string // "container", "mcp-readiness", "tool-latency"
|
||
Check(ctx context.Context, inst *MCPInstance) ProbeResult
|
||
}
|
||
|
||
type ProbeResult struct {
|
||
Healthy bool
|
||
Latency time.Duration
|
||
ErrRate float64 // over the probe's sliding window
|
||
Reason string // human-readable on failure
|
||
Kind ProbeKind // Liveness | Readiness
|
||
}
|
||
|
||
// Monitor owns probing, the state machine, and the restart policy.
|
||
type Monitor interface {
|
||
Register(p Probe) error
|
||
Watch(ctx context.Context) error // periodic + edge-triggered
|
||
Status(instanceID string) (HealthReport, bool) // current snapshot
|
||
Quarantine(ctx context.Context, instanceID, reason string) error
|
||
Release(ctx context.Context, instanceID string) error // un-quarantine
|
||
}
|
||
|
||
// RestartPolicy computes backoff and the quarantine threshold.
|
||
type RestartPolicy struct {
|
||
Base time.Duration // e.g. 1s
|
||
Max time.Duration // cap, e.g. 5m
|
||
Multiplier float64 // e.g. 2.0
|
||
Jitter float64 // 0..1
|
||
MaxRestarts int // consecutive failures → quarantine
|
||
HealthyReset time.Duration // uptime after which the failure count resets
|
||
}
|
||
```
|
||
|
||
Internal HTTP (control API, RBAC-guarded; not agent-facing):
|
||
- `GET /api/v1/health/instances` — list instances with state, latency, error rate, restart count.
|
||
- `GET /api/v1/health/instances/{id}` — full `HealthReport` + transition history.
|
||
- `POST /api/v1/health/instances/{id}/restart` — operator-forced restart.
|
||
- `POST /api/v1/health/instances/{id}/quarantine` · `/release` — manual quarantine controls.
|
||
|
||
## State machine
|
||
```
|
||
probe ok
|
||
┌──────────────────────────────┐
|
||
▼ │
|
||
running ──liveness fail──▶ offline ──heal──▶ restarting ──ready──▶ running
|
||
│ │ │
|
||
│ update begins │ │ restart fails × MaxRestarts
|
||
▼ ▼ ▼
|
||
updating ──(rollback path)──▶ ◀──────────── quarantined ──operator release──▶ restarting
|
||
```
|
||
`updating` is entered/left by the [Update Manager](10-update-manager.md); Health suppresses auto-restart while an instance is `updating` and instead reports readiness so the Updater can decide **rollback on unhealthy**. `quarantined` instances are drained from the Router and require operator release (or a successful update) to re-enter the loop.
|
||
|
||
## Data
|
||
- **Writes** `health` + `state` on each `MCPInstance` in the **Inventory** store (§8), plus an append-only transition log (state, reason, timestamp) for the dashboard timeline.
|
||
- **Reads** instance definitions from Inventory and probe/policy config from the **Config** store; probe credentials (if any) by ref from [Secrets](07-secrets-manager.md).
|
||
- Module-local in-memory ring buffers hold recent latency/error samples per instance for sliding-window rate computation (survive-restart not required; recomputed on boot).
|
||
|
||
## Dependencies
|
||
- [Auto Installer](03-auto-installer.md) + Runtime — executes restart/recreate on Health's request.
|
||
- [Gateway](04-gateway.md) / Router — consumes health events via `UpstreamPool.Drain`; stops routing to non-`running` instances.
|
||
- [Update Manager](10-update-manager.md) — reads readiness to gate promotion and trigger **rollback on unhealthy**; owns the `updating` state.
|
||
- [Notifications](13-notifications.md) — fires on `offline`/`quarantined`/recovered transitions.
|
||
- [Metrics](18-metrics.md) — exports state, latency, error-rate, and restart-count series.
|
||
- [Plugin System](11-plugin-system.md) — generalizes the `Probe` interface for custom checks.
|
||
|
||
## Failure modes & handling
|
||
| Failure | Behavior |
|
||
|---|---|
|
||
| Instance fails liveness | Transition `running→offline`, drain from Router, restart with backoff (`restarting`). |
|
||
| Restarts keep failing | After `MaxRestarts` consecutive failures → `quarantined`; stop restarting; notify; require operator/update to recover. |
|
||
| Flapping (ready→fail→ready) | Backoff prevents restart storms; failure count only resets after `HealthyReset` sustained uptime. |
|
||
| Probe itself times out / Docker API down | Probe error is not the same as instance-down: mark `unknown`, retry with backoff, do **not** restart on probe infrastructure failure alone. |
|
||
| Update in progress | Auto-restart suppressed; Health reports readiness to the Updater, which owns rollback. |
|
||
| Health monitor crashes | Data plane unaffected (Router keeps last-known poolable set until TTL); on restart, probes re-establish state from Inventory. |
|
||
| False-positive readiness | Application probe (real tool-call latency + error rate), not just container-up, reduces "up but broken" routing. |
|
||
|
||
## Security notes
|
||
Honors Architecture §9. Probes are least-privilege: readiness uses the same sandboxed MCP client path as the Router, never a privileged shell into the container. Probe credentials come from [Secrets](07-secrets-manager.md) by ref and are never logged or written into `HealthReport`. Failure `Reason` strings are sanitized before they reach the dashboard or notifications so upstream error payloads cannot leak secrets or internal topology. Every restart/quarantine/release is a mutating action and is **audited** with principal (or `system:health`) + target + outcome. Restart requests to the Runtime are the only privileged action and flow through the same audited Installer path.
|
||
|
||
## Open questions
|
||
- Should backoff/quarantine thresholds be per-`type` (a flaky IoT MCP vs. a rack Postgres MCP differ), overridable per Recipe?
|
||
- Do we expose raw probe latency histograms to the dashboard, or only rolled-up percentiles (topology-leak risk, per Discovery's parallel question)?
|
||
- Quarantine auto-recovery: should Nexus periodically re-probe quarantined instances with a long backoff, or strictly require operator/update release?
|
||
- How do readiness budgets interact with legitimately slow-starting MCP servers (model loads, large indexes)?
|
||
|
||
## Milestone
|
||
Delivered in **Phase 4** (Operate / day-2). Thin slice: container + MCP-readiness probes, the five-state machine, exponential-backoff auto-restart, and quarantine-after-N, wired to Router drain and Notifications. Exit proof (shared with the phase goal): kill a managed MCP container → it is auto-restarted, a Discord alert fires, and the dashboard reflects `running→offline→restarting→running` in real time.
|