MCP Nexus is a self-hosted control plane for MCP servers — "Kubernetes for
MCP." Agents connect to one endpoint; Nexus discovers services, installs the
right MCP servers, aggregates them behind a namespaced router, secures access
with auth/RBAC, and heals/updates them via a reconcile loop.
This first commit is design-phase only (no runnable code yet):
- README.md project front door + module map
- docs/ARCHITECTURE.md target design: tenets, two-plane split, reconcile
loop, domain types, storage, security, deployment
- docs/ROADMAP.md phased delivery (foundations -> walking skeleton ->
discover+install -> secure -> operate -> extend/scale)
- docs/modules/01-20 one design doc per module, all cross-linked
Backend stack decision: Go (single static binary, embedded SQLite + embedded
React dashboard). Repo initialized in /root with a whitelist .gitignore so
only project files are tracked.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
9.4 KiB
Security
Module 19 · Plane: Cross-cutting · Roadmap phase: 3 Part of MCP Nexus architecture.
Purpose
Security is the umbrella model that makes tenet #5 ("secure by default") real and gives Architecture §9 its implementation. It is not one component but the defense-in-depth composition of the security-relevant modules — Authentication, RBAC, Secrets, AI Agent Profiles — plus the cross-cutting controls no single module owns: container sandboxing, append-only audit, signed-package/supply-chain verification, rate limiting, and TLS everywhere. This doc defines those cross-cutting controls and how the layers stack so a failure of any one layer is not a breach. Reference: Architecture §9.
Responsibilities
- Sandbox every MCP instance. Each
MCPInstanceruns in its own container with least privilege: dropped Linux capabilities,no-new-privileges, non-root user, read-only rootfs where the image allows, seccomp/AppArmor profiles, constrained networking (no host network unless the recipe demands it), and CPU/memory limits. - Container isolation defaults. Define the baseline security context the Auto Installer/Runtime applies to every container so sandboxing is the default, not an opt-in.
- Audit every action. Provide the append-only audit sink and schema for both control-plane mutations and every agent tool call, recorded with principal, target, and outcome (Architecture §9).
- Supply-chain verification. Own the trust model for signed packages: verify signatures before install and pin by digest (never by mutable tag), in concert with the Package Registry.
- Rate limiting. Provide per-principal / per-profile / per-tool rate limits at the Gateway so one agent cannot starve others or abuse an upstream (tenet #1 fairness).
- TLS everywhere. Define TLS termination at the Gateway and localhost/socket-only internal component calls.
- Compose the layers. Specify how AuthN → Profiles → RBAC → Secrets → sandbox → audit stack into defense-in-depth, and how a bypass of one is contained by the next.
Non-goals
- Implementing the identity/policy engines. AuthN (06), RBAC (08), Secrets (07), Profiles (14) own their mechanisms; this module composes them and owns the cross-cutting controls.
- Running containers. The Auto Installer/Runtime creates containers; this module defines the security context they must apply.
- Being a SIEM. Nexus emits an append-only audit log and metrics; long-term aggregation/alerting integrates via Notifications and external tooling.
- Guaranteeing upstream service security. Nexus sandboxes the MCP server; the real service behind it (Postgres, Home Assistant) has its own posture.
Interfaces
// The security context every managed container must be created with (Runtime applies it).
type SandboxSpec struct {
ReadOnlyRootfs bool
RunAsNonRoot bool
DropCaps []string // e.g. ["ALL"]; AddCaps only if a recipe justifies it
AddCaps []string
NoNewPrivs bool
Seccomp string // profile name/path
AppArmor string
Network NetMode // "none" | "bridge-scoped" | "host" (host requires justification)
CPU/*limits*/ Resource
Memory Resource
}
// Append-only audit sink (Architecture §8 Audit store). Records are immutable.
type AuditSink interface {
// Record is called for every mutating control-plane action AND every agent tool call.
Record(ctx context.Context, e AuditEvent) error
Query(ctx context.Context, f AuditFilter) ([]AuditEvent, error) // read-only
}
type AuditEvent struct {
Time time.Time
Principal string // who (user/agent/service)
Action string // "tools/call", "role.update", "secret.resolve", ...
Target string // "postgres.query", "role:developer", "secret://..."
Outcome string // "allow" | "deny" | "success" | "error"
Meta map[string]string // never contains secret plaintext
}
// Supply-chain verification (with the Package Registry).
type Verifier interface {
VerifySignature(ctx context.Context, ref string, sig Signature) error
ResolveDigest(ctx context.Context, ref string) (digest string, err error) // pin, never tag
}
// Rate limiting at the Gateway.
type RateLimiter interface {
Allow(ctx context.Context, key string) (bool, RetryAfter) // key = principal|profile|tool
}
HTTP/API surface:
GET /api/v1/audit?principal=&action=&target=&from=&to=— query the append-only log (RBAC-gated; read-only, no delete/edit endpoint by design).GET /api/v1/security/posture— summary: TLS status, sandbox defaults, unsigned-package count, recent denials.GET /metrics— security-relevant counters (auth failures, denials, rate-limit drops) for Metrics.
Data
- Writes/reads the Audit store (Architecture §8) — append-only; no update/delete path exists in code, and integrity may be reinforced by a hash chain over events.
- Reads signature/trust anchors and digests from the Package Registry before install.
- Reads the
SandboxSpecdefaults from the Config store; recipes may request (justified) deviations. - Emits security counters to Metrics and security transitions to Notifications.
Dependencies
- Composes Authentication, RBAC, Secrets, and AI Agent Profiles into defense-in-depth.
- Sandbox spec applied by the Auto Installer/Runtime.
- Signature/digest verification with the Package Registry.
- Rate limiting + TLS at the Gateway.
- Security events flow to Notifications and counters to Metrics.
Failure modes & handling
| Failure | Behavior |
|---|---|
| Unsigned or signature-invalid package | Blocked before install; never becomes an MCPInstance; audited + alerted (tenet #5). |
| Image only offers a mutable tag | Resolved to a digest and pinned; if no digest can be resolved, install fails closed. |
| Container cannot run read-only rootfs | Recipe must declare the writable paths (tmpfs/volume) explicitly; unjustified full-write rootfs is refused. |
| Recipe requests host network / extra caps | Requires explicit justification in the recipe; flagged in security/posture; audited on install. |
| Audit sink unavailable | Mutations/tool calls fail closed (no silent unaudited actions) or buffer to a durable local queue — never proceed unlogged. |
| Rate limit exceeded | Request rejected with 429/retry-after; drop counted; repeated abuse alerts. |
| TLS misconfigured / plaintext | Gateway refuses to serve the data plane in the clear; fails closed at boot. |
| A single layer bypassed (e.g. a tool leaks a name) | Next layer contains it: RBAC still denies the call, Secrets still redacts, audit still records — no single failure is a breach. |
Security notes
This module is the security notes for the system; the layering is the point (Architecture §9):
- Transport — TLS terminates at the Gateway; internal calls are localhost/socket only.
- Identity — Authentication resolves every request to a
Principal; fail closed if unresolved. - Binding — AI Agent Profiles map an agent identity to roles.
- Authorization — RBAC dual-gates
tools/listandtools/call; default deny. - Secrets — Secrets inject at runtime, never to agents/logs; envelope-encrypted at rest.
- Isolation — every MCP runs least-privilege, sandboxed, network-constrained.
- Supply chain — signature-verified, digest-pinned packages only.
- Observability — every action append-only audited; rate-limited; metered.
Each layer assumes the ones above it can fail. Secure defaults are non-optional: unauthenticated ⇒ no access, ungranted ⇒ no tool, unsigned ⇒ no install, unlogged ⇒ no action.
Open questions
- Audit integrity: is a hash-chained/tamper-evident log worth the write cost for v1, or is append-only-by-code enough?
- Do we ship default seccomp/AppArmor profiles per MCP archetype, or a single conservative baseline?
- Rate-limit granularity default: per-principal, per-profile, per-tool, or a composite — and where are limits configured?
- How do we attest the sandbox actually applied (verify the running container's security context vs. the spec)?
- Signing/trust roots for the community recipe & package ecosystem (mirrors Architecture §12).
Milestone
Delivered in Phase 3 (Secure). Thin slice first: container sandboxing defaults (non-root, dropped caps, read-only rootfs where possible, scoped network), append-only audit of every tool call + control-plane mutation, signed-package verification with digest pinning, rate limiting, and TLS termination at the Gateway — composing AuthN/RBAC/Secrets/Profiles into defense-in-depth. Underwrites the whole Phase 3 exit proof: differentiated multi-agent access with a secret-backed MCP that never leaks the secret to any agent-visible payload or log.