# Metrics > Module 18 · Plane: Cross-cutting · Roadmap phase: 4 > Part of [MCP Nexus architecture](../ARCHITECTURE.md). ## Purpose Make Nexus observable. Metrics provides **Prometheus exposition** at `/metrics` and ships a set of **Grafana dashboards** (as JSON) so an operator can see request volume, latency, error rates, token usage, per-agent activity, installed-MCP and discovery counts, container/resource usage, and instance states at a glance. It is strictly **observational**: it records what other modules do and never changes behavior. Per tenet #4, exposition runs in-process in the single binary — no sidecar required — and Prometheus/Grafana remain optional external tools you point at Nexus. ## Responsibilities - Own the process-wide metrics **registry** and the `GET /metrics` Prometheus text-exposition endpoint (OpenMetrics-compatible). - Define and enforce the **metric taxonomy**: stable names, unit suffixes, and a controlled label set (bounded cardinality) that every module emits against. - Provide thin emit helpers (counters, gauges, histograms) so modules record metrics without depending on the Prometheus client directly (keeps the client swappable and cardinality reviewable). - Track the standard series: agent requests, tool-call latency, error rates, token usage, per-agent activity, installed MCP count, discovery counts, container/resource usage (CPU/mem per instance), and `MCPInstance` state gauges. - Ship a curated **Grafana dashboard set** as JSON under `deploy/grafana/`, embedded and downloadable from the [Dashboard](12-web-dashboard.md). - Expose Go runtime + build-info metrics (`go_*`, `nexus_build_info`) for baseline health. ## Non-goals - **Not alerting.** Thresholds/alert routing are [Notifications](13-notifications.md) (event-driven) and Prometheus Alertmanager (metric-driven, external). Metrics only exposes series. - **Not the audit log.** Audit is the immutable per-action record (§8); metrics are aggregate, lossy, and sampled. They answer different questions. - **Not tracing.** Distributed tracing (OpenTelemetry spans) is a possible later addition, out of scope here. - **Not health decisions.** [Health Monitoring](09-health-monitoring.md) owns state transitions and restart policy; it *emits* state gauges here. - **Not long-term storage.** Nexus exposes; Prometheus scrapes and stores. ## Interfaces ```go // Recorder is the narrow surface every module uses to emit metrics. // Backed by the Prometheus client; kept minimal so the backend is swappable. type Recorder interface { Counter(name string, labels Labels) Counter Gauge(name string, labels Labels) Gauge Histogram(name string, labels Labels) Observer } type Labels map[string]string type Observer interface{ Observe(v float64) } type Counter interface{ Inc(); Add(float64) } type Gauge interface{ Set(float64); Inc(); Dec() } // Timer is sugar for latency histograms: defer t.ObserveDuration(). func StartTimer(h Observer) Timer ``` HTTP endpoint: - `GET /metrics` — Prometheus/OpenMetrics exposition (behind Auth in hardened deployments; scrape token supported). ### Metric taxonomy (naming: `nexus___`) | Metric | Type | Key labels | Source module | |---|---|---|---| | `nexus_gateway_requests_total` | counter | `method`, `outcome` | [Gateway](04-gateway.md) | | `nexus_toolcall_latency_seconds` | histogram | `namespace`, `tool`, `outcome` | [Router](05-dynamic-tool-registry.md) | | `nexus_toolcall_errors_total` | counter | `namespace`, `tool`, `code` | Router | | `nexus_tokens_total` | counter | `agent`, `direction` (in/out) | [Agent Profiles](14-agent-profiles.md) | | `nexus_agent_activity_total` | counter | `agent`, `action` | Agent Profiles | | `nexus_instances` | gauge | `state`, `package` | [Health](09-health-monitoring.md) | | `nexus_mcp_installed` | gauge | `package` | [Installer](03-auto-installer.md) | | `nexus_discovered_resources` | gauge | `type`, `confidence_bucket` | [Discovery](01-discovery-engine.md) | | `nexus_container_cpu_seconds_total` | counter | `instance`, `package` | Runtime | | `nexus_container_memory_bytes` | gauge | `instance`, `package` | Runtime | | `nexus_updates_available` | gauge | `package` | [Update Manager](10-update-manager.md) | | `nexus_notifications_sent_total` | counter | `channel`, `severity`, `outcome` | [Notifications](13-notifications.md) | | `nexus_build_info` | gauge (=1) | `version`, `commit`, `go_version` | core | **Cardinality rule:** high-cardinality identities (`agent`, `instance`) are allowed only on low-frequency series or via a bounded allowlist; free-form user input is never a label value. `confidence_bucket` discretizes the 0–1 score to keep Discovery cardinality flat. **Method:** the request-path series follow the **RED** method (Rate, Errors, Duration) so the Overview dashboard reads as one story per subsystem; resource series follow **USE** (Utilization, Saturation, Errors) for containers/hosts. Histograms use fixed, documented buckets tuned to MCP tool-call latencies (sub-ms probes to multi-second model calls) so percentiles are comparable across instances and Grafana panels don't need per-panel re-bucketing. ## Data - **Owns no store.** State lives in the in-memory Prometheus registry, reset on process restart (Prometheus persists the scraped history). - **Reads** nothing persistent of its own; other modules push values through `Recorder`. - Grafana dashboard JSON is a static, versioned asset shipped with the binary (`deploy/grafana/*.json`), not runtime state. ## Dependencies - Emitted into by nearly every module (Gateway, Router, Discovery, Installer, Health, Update Manager, Notifications, Runtime, Agent Profiles) — it is cross-cutting by design. - [Web Dashboard](12-web-dashboard.md) — the Metrics page renders selected series and links to the shipped Grafana dashboards. - [Security](19-security.md) — `/metrics` auth + scrape-token policy. - [Plugin System](11-plugin-system.md) — plugins receive a scoped `Recorder` so their metrics land in the same taxonomy. ## Grafana dashboard set Shipped as JSON, importable or auto-provisioned: 1. **Overview** — request rate, p50/p95/p99 tool-call latency, error ratio, live instance count by state, updates available. 2. **Gateway & Router** — throughput, latency heatmap per namespace, error codes, connection-pool saturation. 3. **Agents** — per-agent activity, token usage (in/out), top tools per agent. 4. **Control plane** — discovery counts by type/confidence, installs over time, reconcile outcomes. 5. **Resources** — per-instance CPU/memory, container restarts (from Health), host headroom. ## Failure modes & handling | Failure | Behavior | |---|---| | A module emits an unregistered/misnamed metric | Central taxonomy + emit helpers reject at registration; CI lints metric names. | | Label cardinality explosion | Bounded label allowlist; free-form values are hashed/bucketed or dropped; cardinality budget checked in tests. | | `/metrics` scrape is slow/large | Exposition is O(series); registry size is bounded by the taxonomy, so scrape cost stays flat. | | Metrics registry contention under load | Client uses lock-free counters; emit is non-blocking and never on the request critical path's error path. | | Prometheus not deployed | No effect — exposition is passive; Nexus functions fully without a scraper. | ## Security notes Honors Architecture §9. `/metrics` can leak operational shape (instance names, package versions, agent identities), so in hardened deployments it is placed **behind Auth** and/or restricted to a scrape token / internal listener, never exposed alongside the public agent edge. No secret values, tokens, or raw request payloads are ever used as metric names or label values. Per-agent series use stable opaque IDs, not human PII. Access to `/metrics` and dashboard JSON is subject to [RBAC](08-rbac.md) when served through the dashboard. ## Open questions - Do we adopt OpenTelemetry as the emit API now (metrics + future traces) or stay on the Prometheus client and bridge later? - Auto-provision Grafana via its API on first run, or ship JSON for manual import only? - Per-agent token accounting: exact per-call, or sampled to bound cardinality at high agent counts? - Should `/metrics` default to authenticated even in the single-binary homelab topology, or open-on-loopback? ## Milestone Delivered in **Phase 4** (Operate / day-2), though `/metrics` exists from **Phase 0** (exit criteria: `nexus serve` exposes `/healthz` and `/metrics`). Phase 4 fills the full taxonomy and ships the Grafana dashboard set. Exit proof (with Health + Notifications): killing a managed container is visible as an `nexus_instances{state="offline"}` blip and a container-restart counter increment on the Overview dashboard.