MCP Nexus is a self-hosted control plane for MCP servers — "Kubernetes for
MCP." Agents connect to one endpoint; Nexus discovers services, installs the
right MCP servers, aggregates them behind a namespaced router, secures access
with auth/RBAC, and heals/updates them via a reconcile loop.
This first commit is design-phase only (no runnable code yet):
- README.md project front door + module map
- docs/ARCHITECTURE.md target design: tenets, two-plane split, reconcile
loop, domain types, storage, security, deployment
- docs/ROADMAP.md phased delivery (foundations -> walking skeleton ->
discover+install -> secure -> operate -> extend/scale)
- docs/modules/01-20 one design doc per module, all cross-linked
Backend stack decision: Go (single static binary, embedded SQLite + embedded
React dashboard). Repo initialized in /root with a whitelist .gitignore so
only project files are tracked.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
8.6 KiB
Metrics
Module 18 · Plane: Cross-cutting · Roadmap phase: 4 Part of MCP Nexus architecture.
Purpose
Make Nexus observable. Metrics provides Prometheus exposition at /metrics and ships a set of Grafana dashboards (as JSON) so an operator can see request volume, latency, error rates, token usage, per-agent activity, installed-MCP and discovery counts, container/resource usage, and instance states at a glance. It is strictly observational: it records what other modules do and never changes behavior. Per tenet #4, exposition runs in-process in the single binary — no sidecar required — and Prometheus/Grafana remain optional external tools you point at Nexus.
Responsibilities
- Own the process-wide metrics registry and the
GET /metricsPrometheus text-exposition endpoint (OpenMetrics-compatible). - Define and enforce the metric taxonomy: stable names, unit suffixes, and a controlled label set (bounded cardinality) that every module emits against.
- Provide thin emit helpers (counters, gauges, histograms) so modules record metrics without depending on the Prometheus client directly (keeps the client swappable and cardinality reviewable).
- Track the standard series: agent requests, tool-call latency, error rates, token usage, per-agent activity, installed MCP count, discovery counts, container/resource usage (CPU/mem per instance), and
MCPInstancestate gauges. - Ship a curated Grafana dashboard set as JSON under
deploy/grafana/, embedded and downloadable from the Dashboard. - Expose Go runtime + build-info metrics (
go_*,nexus_build_info) for baseline health.
Non-goals
- Not alerting. Thresholds/alert routing are Notifications (event-driven) and Prometheus Alertmanager (metric-driven, external). Metrics only exposes series.
- Not the audit log. Audit is the immutable per-action record (§8); metrics are aggregate, lossy, and sampled. They answer different questions.
- Not tracing. Distributed tracing (OpenTelemetry spans) is a possible later addition, out of scope here.
- Not health decisions. Health Monitoring owns state transitions and restart policy; it emits state gauges here.
- Not long-term storage. Nexus exposes; Prometheus scrapes and stores.
Interfaces
// Recorder is the narrow surface every module uses to emit metrics.
// Backed by the Prometheus client; kept minimal so the backend is swappable.
type Recorder interface {
Counter(name string, labels Labels) Counter
Gauge(name string, labels Labels) Gauge
Histogram(name string, labels Labels) Observer
}
type Labels map[string]string
type Observer interface{ Observe(v float64) }
type Counter interface{ Inc(); Add(float64) }
type Gauge interface{ Set(float64); Inc(); Dec() }
// Timer is sugar for latency histograms: defer t.ObserveDuration().
func StartTimer(h Observer) Timer
HTTP endpoint:
GET /metrics— Prometheus/OpenMetrics exposition (behind Auth in hardened deployments; scrape token supported).
Metric taxonomy (naming: nexus_<subsystem>_<name>_<unit>)
| Metric | Type | Key labels | Source module |
|---|---|---|---|
nexus_gateway_requests_total |
counter | method, outcome |
Gateway |
nexus_toolcall_latency_seconds |
histogram | namespace, tool, outcome |
Router |
nexus_toolcall_errors_total |
counter | namespace, tool, code |
Router |
nexus_tokens_total |
counter | agent, direction (in/out) |
Agent Profiles |
nexus_agent_activity_total |
counter | agent, action |
Agent Profiles |
nexus_instances |
gauge | state, package |
Health |
nexus_mcp_installed |
gauge | package |
Installer |
nexus_discovered_resources |
gauge | type, confidence_bucket |
Discovery |
nexus_container_cpu_seconds_total |
counter | instance, package |
Runtime |
nexus_container_memory_bytes |
gauge | instance, package |
Runtime |
nexus_updates_available |
gauge | package |
Update Manager |
nexus_notifications_sent_total |
counter | channel, severity, outcome |
Notifications |
nexus_build_info |
gauge (=1) | version, commit, go_version |
core |
Cardinality rule: high-cardinality identities (agent, instance) are allowed only on low-frequency series or via a bounded allowlist; free-form user input is never a label value. confidence_bucket discretizes the 0–1 score to keep Discovery cardinality flat.
Method: the request-path series follow the RED method (Rate, Errors, Duration) so the Overview dashboard reads as one story per subsystem; resource series follow USE (Utilization, Saturation, Errors) for containers/hosts. Histograms use fixed, documented buckets tuned to MCP tool-call latencies (sub-ms probes to multi-second model calls) so percentiles are comparable across instances and Grafana panels don't need per-panel re-bucketing.
Data
- Owns no store. State lives in the in-memory Prometheus registry, reset on process restart (Prometheus persists the scraped history).
- Reads nothing persistent of its own; other modules push values through
Recorder. - Grafana dashboard JSON is a static, versioned asset shipped with the binary (
deploy/grafana/*.json), not runtime state.
Dependencies
- Emitted into by nearly every module (Gateway, Router, Discovery, Installer, Health, Update Manager, Notifications, Runtime, Agent Profiles) — it is cross-cutting by design.
- Web Dashboard — the Metrics page renders selected series and links to the shipped Grafana dashboards.
- Security —
/metricsauth + scrape-token policy. - Plugin System — plugins receive a scoped
Recorderso their metrics land in the same taxonomy.
Grafana dashboard set
Shipped as JSON, importable or auto-provisioned:
- Overview — request rate, p50/p95/p99 tool-call latency, error ratio, live instance count by state, updates available.
- Gateway & Router — throughput, latency heatmap per namespace, error codes, connection-pool saturation.
- Agents — per-agent activity, token usage (in/out), top tools per agent.
- Control plane — discovery counts by type/confidence, installs over time, reconcile outcomes.
- Resources — per-instance CPU/memory, container restarts (from Health), host headroom.
Failure modes & handling
| Failure | Behavior |
|---|---|
| A module emits an unregistered/misnamed metric | Central taxonomy + emit helpers reject at registration; CI lints metric names. |
| Label cardinality explosion | Bounded label allowlist; free-form values are hashed/bucketed or dropped; cardinality budget checked in tests. |
/metrics scrape is slow/large |
Exposition is O(series); registry size is bounded by the taxonomy, so scrape cost stays flat. |
| Metrics registry contention under load | Client uses lock-free counters; emit is non-blocking and never on the request critical path's error path. |
| Prometheus not deployed | No effect — exposition is passive; Nexus functions fully without a scraper. |
Security notes
Honors Architecture §9. /metrics can leak operational shape (instance names, package versions, agent identities), so in hardened deployments it is placed behind Auth and/or restricted to a scrape token / internal listener, never exposed alongside the public agent edge. No secret values, tokens, or raw request payloads are ever used as metric names or label values. Per-agent series use stable opaque IDs, not human PII. Access to /metrics and dashboard JSON is subject to RBAC when served through the dashboard.
Open questions
- Do we adopt OpenTelemetry as the emit API now (metrics + future traces) or stay on the Prometheus client and bridge later?
- Auto-provision Grafana via its API on first run, or ship JSON for manual import only?
- Per-agent token accounting: exact per-call, or sampled to bound cardinality at high agent counts?
- Should
/metricsdefault to authenticated even in the single-binary homelab topology, or open-on-loopback?
Milestone
Delivered in Phase 4 (Operate / day-2), though /metrics exists from Phase 0 (exit criteria: nexus serve exposes /healthz and /metrics). Phase 4 fills the full taxonomy and ships the Grafana dashboard set. Exit proof (with Health + Notifications): killing a managed container is visible as an nexus_instances{state="offline"} blip and a container-restart counter increment on the Overview dashboard.