Files
mcp-gateway-nexus/docs/modules/18-metrics.md
drjones 8d3ffef920 docs: initial architecture and design for MCP Nexus
MCP Nexus is a self-hosted control plane for MCP servers — "Kubernetes for
MCP." Agents connect to one endpoint; Nexus discovers services, installs the
right MCP servers, aggregates them behind a namespaced router, secures access
with auth/RBAC, and heals/updates them via a reconcile loop.

This first commit is design-phase only (no runnable code yet):
- README.md            project front door + module map
- docs/ARCHITECTURE.md target design: tenets, two-plane split, reconcile
                       loop, domain types, storage, security, deployment
- docs/ROADMAP.md      phased delivery (foundations -> walking skeleton ->
                       discover+install -> secure -> operate -> extend/scale)
- docs/modules/01-20   one design doc per module, all cross-linked

Backend stack decision: Go (single static binary, embedded SQLite + embedded
React dashboard). Repo initialized in /root with a whitelist .gitignore so
only project files are tracked.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-07 04:38:27 +00:00

107 lines
8.6 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Metrics
> Module 18 · Plane: Cross-cutting · Roadmap phase: 4
> Part of [MCP Nexus architecture](../ARCHITECTURE.md).
## Purpose
Make Nexus observable. Metrics provides **Prometheus exposition** at `/metrics` and ships a set of **Grafana dashboards** (as JSON) so an operator can see request volume, latency, error rates, token usage, per-agent activity, installed-MCP and discovery counts, container/resource usage, and instance states at a glance. It is strictly **observational**: it records what other modules do and never changes behavior. Per tenet #4, exposition runs in-process in the single binary — no sidecar required — and Prometheus/Grafana remain optional external tools you point at Nexus.
## Responsibilities
- Own the process-wide metrics **registry** and the `GET /metrics` Prometheus text-exposition endpoint (OpenMetrics-compatible).
- Define and enforce the **metric taxonomy**: stable names, unit suffixes, and a controlled label set (bounded cardinality) that every module emits against.
- Provide thin emit helpers (counters, gauges, histograms) so modules record metrics without depending on the Prometheus client directly (keeps the client swappable and cardinality reviewable).
- Track the standard series: agent requests, tool-call latency, error rates, token usage, per-agent activity, installed MCP count, discovery counts, container/resource usage (CPU/mem per instance), and `MCPInstance` state gauges.
- Ship a curated **Grafana dashboard set** as JSON under `deploy/grafana/`, embedded and downloadable from the [Dashboard](12-web-dashboard.md).
- Expose Go runtime + build-info metrics (`go_*`, `nexus_build_info`) for baseline health.
## Non-goals
- **Not alerting.** Thresholds/alert routing are [Notifications](13-notifications.md) (event-driven) and Prometheus Alertmanager (metric-driven, external). Metrics only exposes series.
- **Not the audit log.** Audit is the immutable per-action record (§8); metrics are aggregate, lossy, and sampled. They answer different questions.
- **Not tracing.** Distributed tracing (OpenTelemetry spans) is a possible later addition, out of scope here.
- **Not health decisions.** [Health Monitoring](09-health-monitoring.md) owns state transitions and restart policy; it *emits* state gauges here.
- **Not long-term storage.** Nexus exposes; Prometheus scrapes and stores.
## Interfaces
```go
// Recorder is the narrow surface every module uses to emit metrics.
// Backed by the Prometheus client; kept minimal so the backend is swappable.
type Recorder interface {
Counter(name string, labels Labels) Counter
Gauge(name string, labels Labels) Gauge
Histogram(name string, labels Labels) Observer
}
type Labels map[string]string
type Observer interface{ Observe(v float64) }
type Counter interface{ Inc(); Add(float64) }
type Gauge interface{ Set(float64); Inc(); Dec() }
// Timer is sugar for latency histograms: defer t.ObserveDuration().
func StartTimer(h Observer) Timer
```
HTTP endpoint:
- `GET /metrics` — Prometheus/OpenMetrics exposition (behind Auth in hardened deployments; scrape token supported).
### Metric taxonomy (naming: `nexus_<subsystem>_<name>_<unit>`)
| Metric | Type | Key labels | Source module |
|---|---|---|---|
| `nexus_gateway_requests_total` | counter | `method`, `outcome` | [Gateway](04-gateway.md) |
| `nexus_toolcall_latency_seconds` | histogram | `namespace`, `tool`, `outcome` | [Router](05-dynamic-tool-registry.md) |
| `nexus_toolcall_errors_total` | counter | `namespace`, `tool`, `code` | Router |
| `nexus_tokens_total` | counter | `agent`, `direction` (in/out) | [Agent Profiles](14-agent-profiles.md) |
| `nexus_agent_activity_total` | counter | `agent`, `action` | Agent Profiles |
| `nexus_instances` | gauge | `state`, `package` | [Health](09-health-monitoring.md) |
| `nexus_mcp_installed` | gauge | `package` | [Installer](03-auto-installer.md) |
| `nexus_discovered_resources` | gauge | `type`, `confidence_bucket` | [Discovery](01-discovery-engine.md) |
| `nexus_container_cpu_seconds_total` | counter | `instance`, `package` | Runtime |
| `nexus_container_memory_bytes` | gauge | `instance`, `package` | Runtime |
| `nexus_updates_available` | gauge | `package` | [Update Manager](10-update-manager.md) |
| `nexus_notifications_sent_total` | counter | `channel`, `severity`, `outcome` | [Notifications](13-notifications.md) |
| `nexus_build_info` | gauge (=1) | `version`, `commit`, `go_version` | core |
**Cardinality rule:** high-cardinality identities (`agent`, `instance`) are allowed only on low-frequency series or via a bounded allowlist; free-form user input is never a label value. `confidence_bucket` discretizes the 0–1 score to keep Discovery cardinality flat.
**Method:** the request-path series follow the **RED** method (Rate, Errors, Duration) so the Overview dashboard reads as one story per subsystem; resource series follow **USE** (Utilization, Saturation, Errors) for containers/hosts. Histograms use fixed, documented buckets tuned to MCP tool-call latencies (sub-ms probes to multi-second model calls) so percentiles are comparable across instances and Grafana panels don't need per-panel re-bucketing.
## Data
- **Owns no store.** State lives in the in-memory Prometheus registry, reset on process restart (Prometheus persists the scraped history).
- **Reads** nothing persistent of its own; other modules push values through `Recorder`.
- Grafana dashboard JSON is a static, versioned asset shipped with the binary (`deploy/grafana/*.json`), not runtime state.
## Dependencies
- Emitted into by nearly every module (Gateway, Router, Discovery, Installer, Health, Update Manager, Notifications, Runtime, Agent Profiles) — it is cross-cutting by design.
- [Web Dashboard](12-web-dashboard.md) — the Metrics page renders selected series and links to the shipped Grafana dashboards.
- [Security](19-security.md) — `/metrics` auth + scrape-token policy.
- [Plugin System](11-plugin-system.md) — plugins receive a scoped `Recorder` so their metrics land in the same taxonomy.
## Grafana dashboard set
Shipped as JSON, importable or auto-provisioned:
1. **Overview** — request rate, p50/p95/p99 tool-call latency, error ratio, live instance count by state, updates available.
2. **Gateway & Router** — throughput, latency heatmap per namespace, error codes, connection-pool saturation.
3. **Agents** — per-agent activity, token usage (in/out), top tools per agent.
4. **Control plane** — discovery counts by type/confidence, installs over time, reconcile outcomes.
5. **Resources** — per-instance CPU/memory, container restarts (from Health), host headroom.
## Failure modes & handling
| Failure | Behavior |
|---|---|
| A module emits an unregistered/misnamed metric | Central taxonomy + emit helpers reject at registration; CI lints metric names. |
| Label cardinality explosion | Bounded label allowlist; free-form values are hashed/bucketed or dropped; cardinality budget checked in tests. |
| `/metrics` scrape is slow/large | Exposition is O(series); registry size is bounded by the taxonomy, so scrape cost stays flat. |
| Metrics registry contention under load | Client uses lock-free counters; emit is non-blocking and never on the request critical path's error path. |
| Prometheus not deployed | No effect — exposition is passive; Nexus functions fully without a scraper. |
## Security notes
Honors Architecture §9. `/metrics` can leak operational shape (instance names, package versions, agent identities), so in hardened deployments it is placed **behind Auth** and/or restricted to a scrape token / internal listener, never exposed alongside the public agent edge. No secret values, tokens, or raw request payloads are ever used as metric names or label values. Per-agent series use stable opaque IDs, not human PII. Access to `/metrics` and dashboard JSON is subject to [RBAC](08-rbac.md) when served through the dashboard.
## Open questions
- Do we adopt OpenTelemetry as the emit API now (metrics + future traces) or stay on the Prometheus client and bridge later?
- Auto-provision Grafana via its API on first run, or ship JSON for manual import only?
- Per-agent token accounting: exact per-call, or sampled to bound cardinality at high agent counts?
- Should `/metrics` default to authenticated even in the single-binary homelab topology, or open-on-loopback?
## Milestone
Delivered in **Phase 4** (Operate / day-2), though `/metrics` exists from **Phase 0** (exit criteria: `nexus serve` exposes `/healthz` and `/metrics`). Phase 4 fills the full taxonomy and ships the Grafana dashboard set. Exit proof (with Health + Notifications): killing a managed container is visible as an `nexus_instances{state="offline"}` blip and a container-restart counter increment on the Overview dashboard.