Files
mcp-gateway-nexus/docs/modules/09-health-monitoring.md
drjones 8d3ffef920 docs: initial architecture and design for MCP Nexus
MCP Nexus is a self-hosted control plane for MCP servers — "Kubernetes for
MCP." Agents connect to one endpoint; Nexus discovers services, installs the
right MCP servers, aggregates them behind a namespaced router, secures access
with auth/RBAC, and heals/updates them via a reconcile loop.

This first commit is design-phase only (no runnable code yet):
- README.md            project front door + module map
- docs/ARCHITECTURE.md target design: tenets, two-plane split, reconcile
                       loop, domain types, storage, security, deployment
- docs/ROADMAP.md      phased delivery (foundations -> walking skeleton ->
                       discover+install -> secure -> operate -> extend/scale)
- docs/modules/01-20   one design doc per module, all cross-linked

Backend stack decision: Go (single static binary, embedded SQLite + embedded
React dashboard). Repo initialized in /root with a whitelist .gitignore so
only project files are tracked.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-07 04:38:27 +00:00

125 lines
9.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Health Monitoring
> Module 09 · Plane: Control · Roadmap phase: 4
> Part of [MCP Nexus architecture](../ARCHITECTURE.md).
## Purpose
Keep every managed [`MCPInstance`](../ARCHITECTURE.md) alive and serving. Health Monitoring continuously probes each instance/container, tracks its lifecycle state (`running`, `offline`, `updating`, `restarting`, `quarantined`), measures latency and error rates, and **automatically heals** unhealthy instances via restart-with-backoff, escalating to quarantine after repeated failure. It is the `act`/`record` feedback arm of the reconciliation loop (Architecture §6): it turns observed instance health into events the reconciler, Router, Update Manager, and Notifications react to. A failure of the control plane must never take down the data plane, so Health only *signals* the Router which instances are poolable — it never sits in the hot path.
## Responsibilities
- Run **liveness** and **readiness** probes against every `MCPInstance` on a per-instance schedule: liveness = "the process/container is up"; readiness = "it completes an MCP `initialize` + `tools/list` within budget."
- Collect health signals from three layers: **container** (Docker inspect: state, restart count, OOM/exit codes), **transport** (MCP client connect success, RTT), and **application** (probe tool-call latency, JSON-RPC error rate over a sliding window).
- Own the per-instance **health state machine** and drive transitions; persist current state + last transition into Inventory.
- Execute the **restart policy**: on failed liveness, restart the instance via the Runtime with exponential backoff and jitter; **quarantine** after N consecutive failed restarts.
- Emit health-transition events on the event bus so: the [Router](04-gateway.md) drains dead instances, the [Update Manager](10-update-manager.md) can roll back a bad update, and [Notifications](13-notifications.md) fire on `offline`/`quarantined`.
- Expose current + historical health via the control API for the [Dashboard](12-web-dashboard.md) and feed gauges/counters to [Metrics](18-metrics.md).
## Non-goals
- **Does not pull, create, or destroy containers** — it *requests* restart/recreate through the Runtime + [Auto Installer](03-auto-installer.md); Runtime owns container lifecycle.
- **Does not decide desired state.** Whether an instance *should* exist is the reconciler's call from Recipes ⨯ Inventory; Health only reports and heals what exists.
- **Does not route traffic.** It signals poolability; the [Gateway](04-gateway.md)/Router owns `UpstreamPool.Drain`.
- **Does not define probe transport plugins** — those interfaces are generalized by the [Plugin System](11-plugin-system.md) in Phase 5.
- **Does not send messages** — it emits events; delivery is [Notifications](13-notifications.md).
## Interfaces
```go
// HealthState is the per-instance lifecycle state.
type HealthState string
const (
StateRunning HealthState = "running"
StateOffline HealthState = "offline" // liveness failing
StateUpdating HealthState = "updating" // owned transiently by Update Manager
StateRestarting HealthState = "restarting" // heal in progress
StateQuarantined HealthState = "quarantined" // gave up after repeated restarts
)
// Probe is a pluggable health check against one instance.
type Probe interface {
Name() string // "container", "mcp-readiness", "tool-latency"
Check(ctx context.Context, inst *MCPInstance) ProbeResult
}
type ProbeResult struct {
Healthy bool
Latency time.Duration
ErrRate float64 // over the probe's sliding window
Reason string // human-readable on failure
Kind ProbeKind // Liveness | Readiness
}
// Monitor owns probing, the state machine, and the restart policy.
type Monitor interface {
Register(p Probe) error
Watch(ctx context.Context) error // periodic + edge-triggered
Status(instanceID string) (HealthReport, bool) // current snapshot
Quarantine(ctx context.Context, instanceID, reason string) error
Release(ctx context.Context, instanceID string) error // un-quarantine
}
// RestartPolicy computes backoff and the quarantine threshold.
type RestartPolicy struct {
Base time.Duration // e.g. 1s
Max time.Duration // cap, e.g. 5m
Multiplier float64 // e.g. 2.0
Jitter float64 // 0..1
MaxRestarts int // consecutive failures → quarantine
HealthyReset time.Duration // uptime after which the failure count resets
}
```
Internal HTTP (control API, RBAC-guarded; not agent-facing):
- `GET /api/v1/health/instances` — list instances with state, latency, error rate, restart count.
- `GET /api/v1/health/instances/{id}` — full `HealthReport` + transition history.
- `POST /api/v1/health/instances/{id}/restart` — operator-forced restart.
- `POST /api/v1/health/instances/{id}/quarantine` · `/release` — manual quarantine controls.
## State machine
```
probe ok
┌──────────────────────────────┐
▼ │
running ──liveness fail──▶ offline ──heal──▶ restarting ──ready──▶ running
│ │ │
│ update begins │ │ restart fails × MaxRestarts
▼ ▼ ▼
updating ──(rollback path)──▶ ◀──────────── quarantined ──operator release──▶ restarting
```
`updating` is entered/left by the [Update Manager](10-update-manager.md); Health suppresses auto-restart while an instance is `updating` and instead reports readiness so the Updater can decide **rollback on unhealthy**. `quarantined` instances are drained from the Router and require operator release (or a successful update) to re-enter the loop.
## Data
- **Writes** `health` + `state` on each `MCPInstance` in the **Inventory** store (§8), plus an append-only transition log (state, reason, timestamp) for the dashboard timeline.
- **Reads** instance definitions from Inventory and probe/policy config from the **Config** store; probe credentials (if any) by ref from [Secrets](07-secrets-manager.md).
- Module-local in-memory ring buffers hold recent latency/error samples per instance for sliding-window rate computation (survive-restart not required; recomputed on boot).
## Dependencies
- [Auto Installer](03-auto-installer.md) + Runtime — executes restart/recreate on Health's request.
- [Gateway](04-gateway.md) / Router — consumes health events via `UpstreamPool.Drain`; stops routing to non-`running` instances.
- [Update Manager](10-update-manager.md) — reads readiness to gate promotion and trigger **rollback on unhealthy**; owns the `updating` state.
- [Notifications](13-notifications.md) — fires on `offline`/`quarantined`/recovered transitions.
- [Metrics](18-metrics.md) — exports state, latency, error-rate, and restart-count series.
- [Plugin System](11-plugin-system.md) — generalizes the `Probe` interface for custom checks.
## Failure modes & handling
| Failure | Behavior |
|---|---|
| Instance fails liveness | Transition `running→offline`, drain from Router, restart with backoff (`restarting`). |
| Restarts keep failing | After `MaxRestarts` consecutive failures → `quarantined`; stop restarting; notify; require operator/update to recover. |
| Flapping (ready→fail→ready) | Backoff prevents restart storms; failure count only resets after `HealthyReset` sustained uptime. |
| Probe itself times out / Docker API down | Probe error is not the same as instance-down: mark `unknown`, retry with backoff, do **not** restart on probe infrastructure failure alone. |
| Update in progress | Auto-restart suppressed; Health reports readiness to the Updater, which owns rollback. |
| Health monitor crashes | Data plane unaffected (Router keeps last-known poolable set until TTL); on restart, probes re-establish state from Inventory. |
| False-positive readiness | Application probe (real tool-call latency + error rate), not just container-up, reduces "up but broken" routing. |
## Security notes
Honors Architecture §9. Probes are least-privilege: readiness uses the same sandboxed MCP client path as the Router, never a privileged shell into the container. Probe credentials come from [Secrets](07-secrets-manager.md) by ref and are never logged or written into `HealthReport`. Failure `Reason` strings are sanitized before they reach the dashboard or notifications so upstream error payloads cannot leak secrets or internal topology. Every restart/quarantine/release is a mutating action and is **audited** with principal (or `system:health`) + target + outcome. Restart requests to the Runtime are the only privileged action and flow through the same audited Installer path.
## Open questions
- Should backoff/quarantine thresholds be per-`type` (a flaky IoT MCP vs. a rack Postgres MCP differ), overridable per Recipe?
- Do we expose raw probe latency histograms to the dashboard, or only rolled-up percentiles (topology-leak risk, per Discovery's parallel question)?
- Quarantine auto-recovery: should Nexus periodically re-probe quarantined instances with a long backoff, or strictly require operator/update release?
- How do readiness budgets interact with legitimately slow-starting MCP servers (model loads, large indexes)?
## Milestone
Delivered in **Phase 4** (Operate / day-2). Thin slice: container + MCP-readiness probes, the five-state machine, exponential-backoff auto-restart, and quarantine-after-N, wired to Router drain and Notifications. Exit proof (shared with the phase goal): kill a managed MCP container → it is auto-restarted, a Discord alert fires, and the dashboard reflects `running→offline→restarting→running` in real time.