# Health Monitoring > Module 09 · Plane: Control · Roadmap phase: 4 > Part of [MCP Nexus architecture](../ARCHITECTURE.md). ## Purpose Keep every managed [`MCPInstance`](../ARCHITECTURE.md) alive and serving. Health Monitoring continuously probes each instance/container, tracks its lifecycle state (`running`, `offline`, `updating`, `restarting`, `quarantined`), measures latency and error rates, and **automatically heals** unhealthy instances via restart-with-backoff, escalating to quarantine after repeated failure. It is the `act`/`record` feedback arm of the reconciliation loop (Architecture §6): it turns observed instance health into events the reconciler, Router, Update Manager, and Notifications react to. A failure of the control plane must never take down the data plane, so Health only *signals* the Router which instances are poolable — it never sits in the hot path. ## Responsibilities - Run **liveness** and **readiness** probes against every `MCPInstance` on a per-instance schedule: liveness = "the process/container is up"; readiness = "it completes an MCP `initialize` + `tools/list` within budget." - Collect health signals from three layers: **container** (Docker inspect: state, restart count, OOM/exit codes), **transport** (MCP client connect success, RTT), and **application** (probe tool-call latency, JSON-RPC error rate over a sliding window). - Own the per-instance **health state machine** and drive transitions; persist current state + last transition into Inventory. - Execute the **restart policy**: on failed liveness, restart the instance via the Runtime with exponential backoff and jitter; **quarantine** after N consecutive failed restarts. - Emit health-transition events on the event bus so: the [Router](04-gateway.md) drains dead instances, the [Update Manager](10-update-manager.md) can roll back a bad update, and [Notifications](13-notifications.md) fire on `offline`/`quarantined`. - Expose current + historical health via the control API for the [Dashboard](12-web-dashboard.md) and feed gauges/counters to [Metrics](18-metrics.md). ## Non-goals - **Does not pull, create, or destroy containers** — it *requests* restart/recreate through the Runtime + [Auto Installer](03-auto-installer.md); Runtime owns container lifecycle. - **Does not decide desired state.** Whether an instance *should* exist is the reconciler's call from Recipes ⨯ Inventory; Health only reports and heals what exists. - **Does not route traffic.** It signals poolability; the [Gateway](04-gateway.md)/Router owns `UpstreamPool.Drain`. - **Does not define probe transport plugins** — those interfaces are generalized by the [Plugin System](11-plugin-system.md) in Phase 5. - **Does not send messages** — it emits events; delivery is [Notifications](13-notifications.md). ## Interfaces ```go // HealthState is the per-instance lifecycle state. type HealthState string const ( StateRunning HealthState = "running" StateOffline HealthState = "offline" // liveness failing StateUpdating HealthState = "updating" // owned transiently by Update Manager StateRestarting HealthState = "restarting" // heal in progress StateQuarantined HealthState = "quarantined" // gave up after repeated restarts ) // Probe is a pluggable health check against one instance. type Probe interface { Name() string // "container", "mcp-readiness", "tool-latency" Check(ctx context.Context, inst *MCPInstance) ProbeResult } type ProbeResult struct { Healthy bool Latency time.Duration ErrRate float64 // over the probe's sliding window Reason string // human-readable on failure Kind ProbeKind // Liveness | Readiness } // Monitor owns probing, the state machine, and the restart policy. type Monitor interface { Register(p Probe) error Watch(ctx context.Context) error // periodic + edge-triggered Status(instanceID string) (HealthReport, bool) // current snapshot Quarantine(ctx context.Context, instanceID, reason string) error Release(ctx context.Context, instanceID string) error // un-quarantine } // RestartPolicy computes backoff and the quarantine threshold. type RestartPolicy struct { Base time.Duration // e.g. 1s Max time.Duration // cap, e.g. 5m Multiplier float64 // e.g. 2.0 Jitter float64 // 0..1 MaxRestarts int // consecutive failures → quarantine HealthyReset time.Duration // uptime after which the failure count resets } ``` Internal HTTP (control API, RBAC-guarded; not agent-facing): - `GET /api/v1/health/instances` — list instances with state, latency, error rate, restart count. - `GET /api/v1/health/instances/{id}` — full `HealthReport` + transition history. - `POST /api/v1/health/instances/{id}/restart` — operator-forced restart. - `POST /api/v1/health/instances/{id}/quarantine` · `/release` — manual quarantine controls. ## State machine ``` probe ok ┌──────────────────────────────┐ ▼ │ running ──liveness fail──▶ offline ──heal──▶ restarting ──ready──▶ running │ │ │ │ update begins │ │ restart fails × MaxRestarts ▼ ▼ ▼ updating ──(rollback path)──▶ ◀──────────── quarantined ──operator release──▶ restarting ``` `updating` is entered/left by the [Update Manager](10-update-manager.md); Health suppresses auto-restart while an instance is `updating` and instead reports readiness so the Updater can decide **rollback on unhealthy**. `quarantined` instances are drained from the Router and require operator release (or a successful update) to re-enter the loop. ## Data - **Writes** `health` + `state` on each `MCPInstance` in the **Inventory** store (§8), plus an append-only transition log (state, reason, timestamp) for the dashboard timeline. - **Reads** instance definitions from Inventory and probe/policy config from the **Config** store; probe credentials (if any) by ref from [Secrets](07-secrets-manager.md). - Module-local in-memory ring buffers hold recent latency/error samples per instance for sliding-window rate computation (survive-restart not required; recomputed on boot). ## Dependencies - [Auto Installer](03-auto-installer.md) + Runtime — executes restart/recreate on Health's request. - [Gateway](04-gateway.md) / Router — consumes health events via `UpstreamPool.Drain`; stops routing to non-`running` instances. - [Update Manager](10-update-manager.md) — reads readiness to gate promotion and trigger **rollback on unhealthy**; owns the `updating` state. - [Notifications](13-notifications.md) — fires on `offline`/`quarantined`/recovered transitions. - [Metrics](18-metrics.md) — exports state, latency, error-rate, and restart-count series. - [Plugin System](11-plugin-system.md) — generalizes the `Probe` interface for custom checks. ## Failure modes & handling | Failure | Behavior | |---|---| | Instance fails liveness | Transition `running→offline`, drain from Router, restart with backoff (`restarting`). | | Restarts keep failing | After `MaxRestarts` consecutive failures → `quarantined`; stop restarting; notify; require operator/update to recover. | | Flapping (ready→fail→ready) | Backoff prevents restart storms; failure count only resets after `HealthyReset` sustained uptime. | | Probe itself times out / Docker API down | Probe error is not the same as instance-down: mark `unknown`, retry with backoff, do **not** restart on probe infrastructure failure alone. | | Update in progress | Auto-restart suppressed; Health reports readiness to the Updater, which owns rollback. | | Health monitor crashes | Data plane unaffected (Router keeps last-known poolable set until TTL); on restart, probes re-establish state from Inventory. | | False-positive readiness | Application probe (real tool-call latency + error rate), not just container-up, reduces "up but broken" routing. | ## Security notes Honors Architecture §9. Probes are least-privilege: readiness uses the same sandboxed MCP client path as the Router, never a privileged shell into the container. Probe credentials come from [Secrets](07-secrets-manager.md) by ref and are never logged or written into `HealthReport`. Failure `Reason` strings are sanitized before they reach the dashboard or notifications so upstream error payloads cannot leak secrets or internal topology. Every restart/quarantine/release is a mutating action and is **audited** with principal (or `system:health`) + target + outcome. Restart requests to the Runtime are the only privileged action and flow through the same audited Installer path. ## Open questions - Should backoff/quarantine thresholds be per-`type` (a flaky IoT MCP vs. a rack Postgres MCP differ), overridable per Recipe? - Do we expose raw probe latency histograms to the dashboard, or only rolled-up percentiles (topology-leak risk, per Discovery's parallel question)? - Quarantine auto-recovery: should Nexus periodically re-probe quarantined instances with a long backoff, or strictly require operator/update release? - How do readiness budgets interact with legitimately slow-starting MCP servers (model loads, large indexes)? ## Milestone Delivered in **Phase 4** (Operate / day-2). Thin slice: container + MCP-readiness probes, the five-state machine, exponential-backoff auto-restart, and quarantine-after-N, wired to Router drain and Notifications. Exit proof (shared with the phase goal): kill a managed MCP container → it is auto-restarted, a Discord alert fires, and the dashboard reflects `running→offline→restarting→running` in real time.