Files
mcp-gateway-nexus/docs/modules/18-metrics.md
drjones 8d3ffef920 docs: initial architecture and design for MCP Nexus
MCP Nexus is a self-hosted control plane for MCP servers — "Kubernetes for
MCP." Agents connect to one endpoint; Nexus discovers services, installs the
right MCP servers, aggregates them behind a namespaced router, secures access
with auth/RBAC, and heals/updates them via a reconcile loop.

This first commit is design-phase only (no runnable code yet):
- README.md            project front door + module map
- docs/ARCHITECTURE.md target design: tenets, two-plane split, reconcile
                       loop, domain types, storage, security, deployment
- docs/ROADMAP.md      phased delivery (foundations -> walking skeleton ->
                       discover+install -> secure -> operate -> extend/scale)
- docs/modules/01-20   one design doc per module, all cross-linked

Backend stack decision: Go (single static binary, embedded SQLite + embedded
React dashboard). Repo initialized in /root with a whitelist .gitignore so
only project files are tracked.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-07 04:38:27 +00:00

8.6 KiB
Raw Blame History

Metrics

Module 18 · Plane: Cross-cutting · Roadmap phase: 4 Part of MCP Nexus architecture.

Purpose

Make Nexus observable. Metrics provides Prometheus exposition at /metrics and ships a set of Grafana dashboards (as JSON) so an operator can see request volume, latency, error rates, token usage, per-agent activity, installed-MCP and discovery counts, container/resource usage, and instance states at a glance. It is strictly observational: it records what other modules do and never changes behavior. Per tenet #4, exposition runs in-process in the single binary — no sidecar required — and Prometheus/Grafana remain optional external tools you point at Nexus.

Responsibilities

  • Own the process-wide metrics registry and the GET /metrics Prometheus text-exposition endpoint (OpenMetrics-compatible).
  • Define and enforce the metric taxonomy: stable names, unit suffixes, and a controlled label set (bounded cardinality) that every module emits against.
  • Provide thin emit helpers (counters, gauges, histograms) so modules record metrics without depending on the Prometheus client directly (keeps the client swappable and cardinality reviewable).
  • Track the standard series: agent requests, tool-call latency, error rates, token usage, per-agent activity, installed MCP count, discovery counts, container/resource usage (CPU/mem per instance), and MCPInstance state gauges.
  • Ship a curated Grafana dashboard set as JSON under deploy/grafana/, embedded and downloadable from the Dashboard.
  • Expose Go runtime + build-info metrics (go_*, nexus_build_info) for baseline health.

Non-goals

  • Not alerting. Thresholds/alert routing are Notifications (event-driven) and Prometheus Alertmanager (metric-driven, external). Metrics only exposes series.
  • Not the audit log. Audit is the immutable per-action record (§8); metrics are aggregate, lossy, and sampled. They answer different questions.
  • Not tracing. Distributed tracing (OpenTelemetry spans) is a possible later addition, out of scope here.
  • Not health decisions. Health Monitoring owns state transitions and restart policy; it emits state gauges here.
  • Not long-term storage. Nexus exposes; Prometheus scrapes and stores.

Interfaces

// Recorder is the narrow surface every module uses to emit metrics.
// Backed by the Prometheus client; kept minimal so the backend is swappable.
type Recorder interface {
    Counter(name string, labels Labels) Counter
    Gauge(name string, labels Labels) Gauge
    Histogram(name string, labels Labels) Observer
}

type Labels map[string]string

type Observer interface{ Observe(v float64) }
type Counter  interface{ Inc(); Add(float64) }
type Gauge    interface{ Set(float64); Inc(); Dec() }

// Timer is sugar for latency histograms: defer t.ObserveDuration().
func StartTimer(h Observer) Timer

HTTP endpoint:

  • GET /metrics — Prometheus/OpenMetrics exposition (behind Auth in hardened deployments; scrape token supported).

Metric taxonomy (naming: nexus_<subsystem>_<name>_<unit>)

Metric Type Key labels Source module
nexus_gateway_requests_total counter method, outcome Gateway
nexus_toolcall_latency_seconds histogram namespace, tool, outcome Router
nexus_toolcall_errors_total counter namespace, tool, code Router
nexus_tokens_total counter agent, direction (in/out) Agent Profiles
nexus_agent_activity_total counter agent, action Agent Profiles
nexus_instances gauge state, package Health
nexus_mcp_installed gauge package Installer
nexus_discovered_resources gauge type, confidence_bucket Discovery
nexus_container_cpu_seconds_total counter instance, package Runtime
nexus_container_memory_bytes gauge instance, package Runtime
nexus_updates_available gauge package Update Manager
nexus_notifications_sent_total counter channel, severity, outcome Notifications
nexus_build_info gauge (=1) version, commit, go_version core

Cardinality rule: high-cardinality identities (agent, instance) are allowed only on low-frequency series or via a bounded allowlist; free-form user input is never a label value. confidence_bucket discretizes the 0–1 score to keep Discovery cardinality flat.

Method: the request-path series follow the RED method (Rate, Errors, Duration) so the Overview dashboard reads as one story per subsystem; resource series follow USE (Utilization, Saturation, Errors) for containers/hosts. Histograms use fixed, documented buckets tuned to MCP tool-call latencies (sub-ms probes to multi-second model calls) so percentiles are comparable across instances and Grafana panels don't need per-panel re-bucketing.

Data

  • Owns no store. State lives in the in-memory Prometheus registry, reset on process restart (Prometheus persists the scraped history).
  • Reads nothing persistent of its own; other modules push values through Recorder.
  • Grafana dashboard JSON is a static, versioned asset shipped with the binary (deploy/grafana/*.json), not runtime state.

Dependencies

  • Emitted into by nearly every module (Gateway, Router, Discovery, Installer, Health, Update Manager, Notifications, Runtime, Agent Profiles) — it is cross-cutting by design.
  • Web Dashboard — the Metrics page renders selected series and links to the shipped Grafana dashboards.
  • Security — /metrics auth + scrape-token policy.
  • Plugin System — plugins receive a scoped Recorder so their metrics land in the same taxonomy.

Grafana dashboard set

Shipped as JSON, importable or auto-provisioned:

  1. Overview — request rate, p50/p95/p99 tool-call latency, error ratio, live instance count by state, updates available.
  2. Gateway & Router — throughput, latency heatmap per namespace, error codes, connection-pool saturation.
  3. Agents — per-agent activity, token usage (in/out), top tools per agent.
  4. Control plane — discovery counts by type/confidence, installs over time, reconcile outcomes.
  5. Resources — per-instance CPU/memory, container restarts (from Health), host headroom.

Failure modes & handling

Failure Behavior
A module emits an unregistered/misnamed metric Central taxonomy + emit helpers reject at registration; CI lints metric names.
Label cardinality explosion Bounded label allowlist; free-form values are hashed/bucketed or dropped; cardinality budget checked in tests.
/metrics scrape is slow/large Exposition is O(series); registry size is bounded by the taxonomy, so scrape cost stays flat.
Metrics registry contention under load Client uses lock-free counters; emit is non-blocking and never on the request critical path's error path.
Prometheus not deployed No effect — exposition is passive; Nexus functions fully without a scraper.

Security notes

Honors Architecture §9. /metrics can leak operational shape (instance names, package versions, agent identities), so in hardened deployments it is placed behind Auth and/or restricted to a scrape token / internal listener, never exposed alongside the public agent edge. No secret values, tokens, or raw request payloads are ever used as metric names or label values. Per-agent series use stable opaque IDs, not human PII. Access to /metrics and dashboard JSON is subject to RBAC when served through the dashboard.

Open questions

  • Do we adopt OpenTelemetry as the emit API now (metrics + future traces) or stay on the Prometheus client and bridge later?
  • Auto-provision Grafana via its API on first run, or ship JSON for manual import only?
  • Per-agent token accounting: exact per-call, or sampled to bound cardinality at high agent counts?
  • Should /metrics default to authenticated even in the single-binary homelab topology, or open-on-loopback?

Milestone

Delivered in Phase 4 (Operate / day-2), though /metrics exists from Phase 0 (exit criteria: nexus serve exposes /healthz and /metrics). Phase 4 fills the full taxonomy and ships the Grafana dashboard set. Exit proof (with Health + Notifications): killing a managed container is visible as an nexus_instances{state="offline"} blip and a container-restart counter increment on the Overview dashboard.