docs: initial architecture and design for MCP Nexus

MCP Nexus is a self-hosted control plane for MCP servers — "Kubernetes for
MCP." Agents connect to one endpoint; Nexus discovers services, installs the
right MCP servers, aggregates them behind a namespaced router, secures access
with auth/RBAC, and heals/updates them via a reconcile loop.

This first commit is design-phase only (no runnable code yet):
- README.md            project front door + module map
- docs/ARCHITECTURE.md target design: tenets, two-plane split, reconcile
                       loop, domain types, storage, security, deployment
- docs/ROADMAP.md      phased delivery (foundations -> walking skeleton ->
                       discover+install -> secure -> operate -> extend/scale)
- docs/modules/01-20   one design doc per module, all cross-linked

Backend stack decision: Go (single static binary, embedded SQLite + embedded
React dashboard). Repo initialized in /root with a whitelist .gitignore so
only project files are tracked.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
drjones
2026-07-07 04:38:27 +00:00
commit 8d3ffef920
24 changed files with 3000 additions and 0 deletions

106
docs/modules/18-metrics.md Normal file
View File

@@ -0,0 +1,106 @@
# Metrics
> Module 18 · Plane: Cross-cutting · Roadmap phase: 4
> Part of [MCP Nexus architecture](../ARCHITECTURE.md).
## Purpose
Make Nexus observable. Metrics provides **Prometheus exposition** at `/metrics` and ships a set of **Grafana dashboards** (as JSON) so an operator can see request volume, latency, error rates, token usage, per-agent activity, installed-MCP and discovery counts, container/resource usage, and instance states at a glance. It is strictly **observational**: it records what other modules do and never changes behavior. Per tenet #4, exposition runs in-process in the single binary — no sidecar required — and Prometheus/Grafana remain optional external tools you point at Nexus.
## Responsibilities
- Own the process-wide metrics **registry** and the `GET /metrics` Prometheus text-exposition endpoint (OpenMetrics-compatible).
- Define and enforce the **metric taxonomy**: stable names, unit suffixes, and a controlled label set (bounded cardinality) that every module emits against.
- Provide thin emit helpers (counters, gauges, histograms) so modules record metrics without depending on the Prometheus client directly (keeps the client swappable and cardinality reviewable).
- Track the standard series: agent requests, tool-call latency, error rates, token usage, per-agent activity, installed MCP count, discovery counts, container/resource usage (CPU/mem per instance), and `MCPInstance` state gauges.
- Ship a curated **Grafana dashboard set** as JSON under `deploy/grafana/`, embedded and downloadable from the [Dashboard](12-web-dashboard.md).
- Expose Go runtime + build-info metrics (`go_*`, `nexus_build_info`) for baseline health.
## Non-goals
- **Not alerting.** Thresholds/alert routing are [Notifications](13-notifications.md) (event-driven) and Prometheus Alertmanager (metric-driven, external). Metrics only exposes series.
- **Not the audit log.** Audit is the immutable per-action record (§8); metrics are aggregate, lossy, and sampled. They answer different questions.
- **Not tracing.** Distributed tracing (OpenTelemetry spans) is a possible later addition, out of scope here.
- **Not health decisions.** [Health Monitoring](09-health-monitoring.md) owns state transitions and restart policy; it *emits* state gauges here.
- **Not long-term storage.** Nexus exposes; Prometheus scrapes and stores.
## Interfaces
```go
// Recorder is the narrow surface every module uses to emit metrics.
// Backed by the Prometheus client; kept minimal so the backend is swappable.
type Recorder interface {
Counter(name string, labels Labels) Counter
Gauge(name string, labels Labels) Gauge
Histogram(name string, labels Labels) Observer
}
type Labels map[string]string
type Observer interface{ Observe(v float64) }
type Counter interface{ Inc(); Add(float64) }
type Gauge interface{ Set(float64); Inc(); Dec() }
// Timer is sugar for latency histograms: defer t.ObserveDuration().
func StartTimer(h Observer) Timer
```
HTTP endpoint:
- `GET /metrics` — Prometheus/OpenMetrics exposition (behind Auth in hardened deployments; scrape token supported).
### Metric taxonomy (naming: `nexus_<subsystem>_<name>_<unit>`)
| Metric | Type | Key labels | Source module |
|---|---|---|---|
| `nexus_gateway_requests_total` | counter | `method`, `outcome` | [Gateway](04-gateway.md) |
| `nexus_toolcall_latency_seconds` | histogram | `namespace`, `tool`, `outcome` | [Router](05-dynamic-tool-registry.md) |
| `nexus_toolcall_errors_total` | counter | `namespace`, `tool`, `code` | Router |
| `nexus_tokens_total` | counter | `agent`, `direction` (in/out) | [Agent Profiles](14-agent-profiles.md) |
| `nexus_agent_activity_total` | counter | `agent`, `action` | Agent Profiles |
| `nexus_instances` | gauge | `state`, `package` | [Health](09-health-monitoring.md) |
| `nexus_mcp_installed` | gauge | `package` | [Installer](03-auto-installer.md) |
| `nexus_discovered_resources` | gauge | `type`, `confidence_bucket` | [Discovery](01-discovery-engine.md) |
| `nexus_container_cpu_seconds_total` | counter | `instance`, `package` | Runtime |
| `nexus_container_memory_bytes` | gauge | `instance`, `package` | Runtime |
| `nexus_updates_available` | gauge | `package` | [Update Manager](10-update-manager.md) |
| `nexus_notifications_sent_total` | counter | `channel`, `severity`, `outcome` | [Notifications](13-notifications.md) |
| `nexus_build_info` | gauge (=1) | `version`, `commit`, `go_version` | core |
**Cardinality rule:** high-cardinality identities (`agent`, `instance`) are allowed only on low-frequency series or via a bounded allowlist; free-form user input is never a label value. `confidence_bucket` discretizes the 0–1 score to keep Discovery cardinality flat.
**Method:** the request-path series follow the **RED** method (Rate, Errors, Duration) so the Overview dashboard reads as one story per subsystem; resource series follow **USE** (Utilization, Saturation, Errors) for containers/hosts. Histograms use fixed, documented buckets tuned to MCP tool-call latencies (sub-ms probes to multi-second model calls) so percentiles are comparable across instances and Grafana panels don't need per-panel re-bucketing.
## Data
- **Owns no store.** State lives in the in-memory Prometheus registry, reset on process restart (Prometheus persists the scraped history).
- **Reads** nothing persistent of its own; other modules push values through `Recorder`.
- Grafana dashboard JSON is a static, versioned asset shipped with the binary (`deploy/grafana/*.json`), not runtime state.
## Dependencies
- Emitted into by nearly every module (Gateway, Router, Discovery, Installer, Health, Update Manager, Notifications, Runtime, Agent Profiles) — it is cross-cutting by design.
- [Web Dashboard](12-web-dashboard.md) — the Metrics page renders selected series and links to the shipped Grafana dashboards.
- [Security](19-security.md) — `/metrics` auth + scrape-token policy.
- [Plugin System](11-plugin-system.md) — plugins receive a scoped `Recorder` so their metrics land in the same taxonomy.
## Grafana dashboard set
Shipped as JSON, importable or auto-provisioned:
1. **Overview** — request rate, p50/p95/p99 tool-call latency, error ratio, live instance count by state, updates available.
2. **Gateway & Router** — throughput, latency heatmap per namespace, error codes, connection-pool saturation.
3. **Agents** — per-agent activity, token usage (in/out), top tools per agent.
4. **Control plane** — discovery counts by type/confidence, installs over time, reconcile outcomes.
5. **Resources** — per-instance CPU/memory, container restarts (from Health), host headroom.
## Failure modes & handling
| Failure | Behavior |
|---|---|
| A module emits an unregistered/misnamed metric | Central taxonomy + emit helpers reject at registration; CI lints metric names. |
| Label cardinality explosion | Bounded label allowlist; free-form values are hashed/bucketed or dropped; cardinality budget checked in tests. |
| `/metrics` scrape is slow/large | Exposition is O(series); registry size is bounded by the taxonomy, so scrape cost stays flat. |
| Metrics registry contention under load | Client uses lock-free counters; emit is non-blocking and never on the request critical path's error path. |
| Prometheus not deployed | No effect — exposition is passive; Nexus functions fully without a scraper. |
## Security notes
Honors Architecture §9. `/metrics` can leak operational shape (instance names, package versions, agent identities), so in hardened deployments it is placed **behind Auth** and/or restricted to a scrape token / internal listener, never exposed alongside the public agent edge. No secret values, tokens, or raw request payloads are ever used as metric names or label values. Per-agent series use stable opaque IDs, not human PII. Access to `/metrics` and dashboard JSON is subject to [RBAC](08-rbac.md) when served through the dashboard.
## Open questions
- Do we adopt OpenTelemetry as the emit API now (metrics + future traces) or stay on the Prometheus client and bridge later?
- Auto-provision Grafana via its API on first run, or ship JSON for manual import only?
- Per-agent token accounting: exact per-call, or sampled to bound cardinality at high agent counts?
- Should `/metrics` default to authenticated even in the single-binary homelab topology, or open-on-loopback?
## Milestone
Delivered in **Phase 4** (Operate / day-2), though `/metrics` exists from **Phase 0** (exit criteria: `nexus serve` exposes `/healthz` and `/metrics`). Phase 4 fills the full taxonomy and ships the Grafana dashboard set. Exit proof (with Health + Notifications): killing a managed container is visible as an `nexus_instances{state="offline"}` blip and a container-restart counter increment on the Overview dashboard.