docs: initial architecture and design for MCP Nexus
MCP Nexus is a self-hosted control plane for MCP servers — "Kubernetes for
MCP." Agents connect to one endpoint; Nexus discovers services, installs the
right MCP servers, aggregates them behind a namespaced router, secures access
with auth/RBAC, and heals/updates them via a reconcile loop.
This first commit is design-phase only (no runnable code yet):
- README.md project front door + module map
- docs/ARCHITECTURE.md target design: tenets, two-plane split, reconcile
loop, domain types, storage, security, deployment
- docs/ROADMAP.md phased delivery (foundations -> walking skeleton ->
discover+install -> secure -> operate -> extend/scale)
- docs/modules/01-20 one design doc per module, all cross-linked
Backend stack decision: Go (single static binary, embedded SQLite + embedded
React dashboard). Repo initialized in /root with a whitelist .gitignore so
only project files are tracked.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
106
docs/modules/18-metrics.md
Normal file
106
docs/modules/18-metrics.md
Normal file
@@ -0,0 +1,106 @@
|
||||
# Metrics
|
||||
|
||||
> Module 18 · Plane: Cross-cutting · Roadmap phase: 4
|
||||
> Part of [MCP Nexus architecture](../ARCHITECTURE.md).
|
||||
|
||||
## Purpose
|
||||
Make Nexus observable. Metrics provides **Prometheus exposition** at `/metrics` and ships a set of **Grafana dashboards** (as JSON) so an operator can see request volume, latency, error rates, token usage, per-agent activity, installed-MCP and discovery counts, container/resource usage, and instance states at a glance. It is strictly **observational**: it records what other modules do and never changes behavior. Per tenet #4, exposition runs in-process in the single binary — no sidecar required — and Prometheus/Grafana remain optional external tools you point at Nexus.
|
||||
|
||||
## Responsibilities
|
||||
- Own the process-wide metrics **registry** and the `GET /metrics` Prometheus text-exposition endpoint (OpenMetrics-compatible).
|
||||
- Define and enforce the **metric taxonomy**: stable names, unit suffixes, and a controlled label set (bounded cardinality) that every module emits against.
|
||||
- Provide thin emit helpers (counters, gauges, histograms) so modules record metrics without depending on the Prometheus client directly (keeps the client swappable and cardinality reviewable).
|
||||
- Track the standard series: agent requests, tool-call latency, error rates, token usage, per-agent activity, installed MCP count, discovery counts, container/resource usage (CPU/mem per instance), and `MCPInstance` state gauges.
|
||||
- Ship a curated **Grafana dashboard set** as JSON under `deploy/grafana/`, embedded and downloadable from the [Dashboard](12-web-dashboard.md).
|
||||
- Expose Go runtime + build-info metrics (`go_*`, `nexus_build_info`) for baseline health.
|
||||
|
||||
## Non-goals
|
||||
- **Not alerting.** Thresholds/alert routing are [Notifications](13-notifications.md) (event-driven) and Prometheus Alertmanager (metric-driven, external). Metrics only exposes series.
|
||||
- **Not the audit log.** Audit is the immutable per-action record (§8); metrics are aggregate, lossy, and sampled. They answer different questions.
|
||||
- **Not tracing.** Distributed tracing (OpenTelemetry spans) is a possible later addition, out of scope here.
|
||||
- **Not health decisions.** [Health Monitoring](09-health-monitoring.md) owns state transitions and restart policy; it *emits* state gauges here.
|
||||
- **Not long-term storage.** Nexus exposes; Prometheus scrapes and stores.
|
||||
|
||||
## Interfaces
|
||||
```go
|
||||
// Recorder is the narrow surface every module uses to emit metrics.
|
||||
// Backed by the Prometheus client; kept minimal so the backend is swappable.
|
||||
type Recorder interface {
|
||||
Counter(name string, labels Labels) Counter
|
||||
Gauge(name string, labels Labels) Gauge
|
||||
Histogram(name string, labels Labels) Observer
|
||||
}
|
||||
|
||||
type Labels map[string]string
|
||||
|
||||
type Observer interface{ Observe(v float64) }
|
||||
type Counter interface{ Inc(); Add(float64) }
|
||||
type Gauge interface{ Set(float64); Inc(); Dec() }
|
||||
|
||||
// Timer is sugar for latency histograms: defer t.ObserveDuration().
|
||||
func StartTimer(h Observer) Timer
|
||||
```
|
||||
|
||||
HTTP endpoint:
|
||||
- `GET /metrics` — Prometheus/OpenMetrics exposition (behind Auth in hardened deployments; scrape token supported).
|
||||
|
||||
### Metric taxonomy (naming: `nexus_<subsystem>_<name>_<unit>`)
|
||||
| Metric | Type | Key labels | Source module |
|
||||
|---|---|---|---|
|
||||
| `nexus_gateway_requests_total` | counter | `method`, `outcome` | [Gateway](04-gateway.md) |
|
||||
| `nexus_toolcall_latency_seconds` | histogram | `namespace`, `tool`, `outcome` | [Router](05-dynamic-tool-registry.md) |
|
||||
| `nexus_toolcall_errors_total` | counter | `namespace`, `tool`, `code` | Router |
|
||||
| `nexus_tokens_total` | counter | `agent`, `direction` (in/out) | [Agent Profiles](14-agent-profiles.md) |
|
||||
| `nexus_agent_activity_total` | counter | `agent`, `action` | Agent Profiles |
|
||||
| `nexus_instances` | gauge | `state`, `package` | [Health](09-health-monitoring.md) |
|
||||
| `nexus_mcp_installed` | gauge | `package` | [Installer](03-auto-installer.md) |
|
||||
| `nexus_discovered_resources` | gauge | `type`, `confidence_bucket` | [Discovery](01-discovery-engine.md) |
|
||||
| `nexus_container_cpu_seconds_total` | counter | `instance`, `package` | Runtime |
|
||||
| `nexus_container_memory_bytes` | gauge | `instance`, `package` | Runtime |
|
||||
| `nexus_updates_available` | gauge | `package` | [Update Manager](10-update-manager.md) |
|
||||
| `nexus_notifications_sent_total` | counter | `channel`, `severity`, `outcome` | [Notifications](13-notifications.md) |
|
||||
| `nexus_build_info` | gauge (=1) | `version`, `commit`, `go_version` | core |
|
||||
|
||||
**Cardinality rule:** high-cardinality identities (`agent`, `instance`) are allowed only on low-frequency series or via a bounded allowlist; free-form user input is never a label value. `confidence_bucket` discretizes the 0–1 score to keep Discovery cardinality flat.
|
||||
|
||||
**Method:** the request-path series follow the **RED** method (Rate, Errors, Duration) so the Overview dashboard reads as one story per subsystem; resource series follow **USE** (Utilization, Saturation, Errors) for containers/hosts. Histograms use fixed, documented buckets tuned to MCP tool-call latencies (sub-ms probes to multi-second model calls) so percentiles are comparable across instances and Grafana panels don't need per-panel re-bucketing.
|
||||
|
||||
## Data
|
||||
- **Owns no store.** State lives in the in-memory Prometheus registry, reset on process restart (Prometheus persists the scraped history).
|
||||
- **Reads** nothing persistent of its own; other modules push values through `Recorder`.
|
||||
- Grafana dashboard JSON is a static, versioned asset shipped with the binary (`deploy/grafana/*.json`), not runtime state.
|
||||
|
||||
## Dependencies
|
||||
- Emitted into by nearly every module (Gateway, Router, Discovery, Installer, Health, Update Manager, Notifications, Runtime, Agent Profiles) — it is cross-cutting by design.
|
||||
- [Web Dashboard](12-web-dashboard.md) — the Metrics page renders selected series and links to the shipped Grafana dashboards.
|
||||
- [Security](19-security.md) — `/metrics` auth + scrape-token policy.
|
||||
- [Plugin System](11-plugin-system.md) — plugins receive a scoped `Recorder` so their metrics land in the same taxonomy.
|
||||
|
||||
## Grafana dashboard set
|
||||
Shipped as JSON, importable or auto-provisioned:
|
||||
1. **Overview** — request rate, p50/p95/p99 tool-call latency, error ratio, live instance count by state, updates available.
|
||||
2. **Gateway & Router** — throughput, latency heatmap per namespace, error codes, connection-pool saturation.
|
||||
3. **Agents** — per-agent activity, token usage (in/out), top tools per agent.
|
||||
4. **Control plane** — discovery counts by type/confidence, installs over time, reconcile outcomes.
|
||||
5. **Resources** — per-instance CPU/memory, container restarts (from Health), host headroom.
|
||||
|
||||
## Failure modes & handling
|
||||
| Failure | Behavior |
|
||||
|---|---|
|
||||
| A module emits an unregistered/misnamed metric | Central taxonomy + emit helpers reject at registration; CI lints metric names. |
|
||||
| Label cardinality explosion | Bounded label allowlist; free-form values are hashed/bucketed or dropped; cardinality budget checked in tests. |
|
||||
| `/metrics` scrape is slow/large | Exposition is O(series); registry size is bounded by the taxonomy, so scrape cost stays flat. |
|
||||
| Metrics registry contention under load | Client uses lock-free counters; emit is non-blocking and never on the request critical path's error path. |
|
||||
| Prometheus not deployed | No effect — exposition is passive; Nexus functions fully without a scraper. |
|
||||
|
||||
## Security notes
|
||||
Honors Architecture §9. `/metrics` can leak operational shape (instance names, package versions, agent identities), so in hardened deployments it is placed **behind Auth** and/or restricted to a scrape token / internal listener, never exposed alongside the public agent edge. No secret values, tokens, or raw request payloads are ever used as metric names or label values. Per-agent series use stable opaque IDs, not human PII. Access to `/metrics` and dashboard JSON is subject to [RBAC](08-rbac.md) when served through the dashboard.
|
||||
|
||||
## Open questions
|
||||
- Do we adopt OpenTelemetry as the emit API now (metrics + future traces) or stay on the Prometheus client and bridge later?
|
||||
- Auto-provision Grafana via its API on first run, or ship JSON for manual import only?
|
||||
- Per-agent token accounting: exact per-call, or sampled to bound cardinality at high agent counts?
|
||||
- Should `/metrics` default to authenticated even in the single-binary homelab topology, or open-on-loopback?
|
||||
|
||||
## Milestone
|
||||
Delivered in **Phase 4** (Operate / day-2), though `/metrics` exists from **Phase 0** (exit criteria: `nexus serve` exposes `/healthz` and `/metrics`). Phase 4 fills the full taxonomy and ships the Grafana dashboard set. Exit proof (with Health + Notifications): killing a managed container is visible as an `nexus_instances{state="offline"}` blip and a container-restart counter increment on the Overview dashboard.
|
||||
Reference in New Issue
Block a user