Files
mcp-gateway-nexus/docs/modules/19-security.md
drjones 8d3ffef920 docs: initial architecture and design for MCP Nexus
MCP Nexus is a self-hosted control plane for MCP servers — "Kubernetes for
MCP." Agents connect to one endpoint; Nexus discovers services, installs the
right MCP servers, aggregates them behind a namespaced router, secures access
with auth/RBAC, and heals/updates them via a reconcile loop.

This first commit is design-phase only (no runnable code yet):
- README.md            project front door + module map
- docs/ARCHITECTURE.md target design: tenets, two-plane split, reconcile
                       loop, domain types, storage, security, deployment
- docs/ROADMAP.md      phased delivery (foundations -> walking skeleton ->
                       discover+install -> secure -> operate -> extend/scale)
- docs/modules/01-20   one design doc per module, all cross-linked

Backend stack decision: Go (single static binary, embedded SQLite + embedded
React dashboard). Repo initialized in /root with a whitelist .gitignore so
only project files are tracked.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-07 04:38:27 +00:00

9.4 KiB

Security

Module 19 · Plane: Cross-cutting · Roadmap phase: 3 Part of MCP Nexus architecture.

Purpose

Security is the umbrella model that makes tenet #5 ("secure by default") real and gives Architecture §9 its implementation. It is not one component but the defense-in-depth composition of the security-relevant modules — Authentication, RBAC, Secrets, AI Agent Profiles — plus the cross-cutting controls no single module owns: container sandboxing, append-only audit, signed-package/supply-chain verification, rate limiting, and TLS everywhere. This doc defines those cross-cutting controls and how the layers stack so a failure of any one layer is not a breach. Reference: Architecture §9.

Responsibilities

  • Sandbox every MCP instance. Each MCPInstance runs in its own container with least privilege: dropped Linux capabilities, no-new-privileges, non-root user, read-only rootfs where the image allows, seccomp/AppArmor profiles, constrained networking (no host network unless the recipe demands it), and CPU/memory limits.
  • Container isolation defaults. Define the baseline security context the Auto Installer/Runtime applies to every container so sandboxing is the default, not an opt-in.
  • Audit every action. Provide the append-only audit sink and schema for both control-plane mutations and every agent tool call, recorded with principal, target, and outcome (Architecture §9).
  • Supply-chain verification. Own the trust model for signed packages: verify signatures before install and pin by digest (never by mutable tag), in concert with the Package Registry.
  • Rate limiting. Provide per-principal / per-profile / per-tool rate limits at the Gateway so one agent cannot starve others or abuse an upstream (tenet #1 fairness).
  • TLS everywhere. Define TLS termination at the Gateway and localhost/socket-only internal component calls.
  • Compose the layers. Specify how AuthN → Profiles → RBAC → Secrets → sandbox → audit stack into defense-in-depth, and how a bypass of one is contained by the next.

Non-goals

  • Implementing the identity/policy engines. AuthN (06), RBAC (08), Secrets (07), Profiles (14) own their mechanisms; this module composes them and owns the cross-cutting controls.
  • Running containers. The Auto Installer/Runtime creates containers; this module defines the security context they must apply.
  • Being a SIEM. Nexus emits an append-only audit log and metrics; long-term aggregation/alerting integrates via Notifications and external tooling.
  • Guaranteeing upstream service security. Nexus sandboxes the MCP server; the real service behind it (Postgres, Home Assistant) has its own posture.

Interfaces

// The security context every managed container must be created with (Runtime applies it).
type SandboxSpec struct {
    ReadOnlyRootfs bool
    RunAsNonRoot   bool
    DropCaps       []string // e.g. ["ALL"]; AddCaps only if a recipe justifies it
    AddCaps        []string
    NoNewPrivs     bool
    Seccomp        string   // profile name/path
    AppArmor       string
    Network        NetMode  // "none" | "bridge-scoped" | "host" (host requires justification)
    CPU/*limits*/  Resource
    Memory         Resource
}

// Append-only audit sink (Architecture §8 Audit store). Records are immutable.
type AuditSink interface {
    // Record is called for every mutating control-plane action AND every agent tool call.
    Record(ctx context.Context, e AuditEvent) error
    Query(ctx context.Context, f AuditFilter) ([]AuditEvent, error) // read-only
}

type AuditEvent struct {
    Time      time.Time
    Principal string // who (user/agent/service)
    Action    string // "tools/call", "role.update", "secret.resolve", ...
    Target    string // "postgres.query", "role:developer", "secret://..."
    Outcome   string // "allow" | "deny" | "success" | "error"
    Meta      map[string]string // never contains secret plaintext
}

// Supply-chain verification (with the Package Registry).
type Verifier interface {
    VerifySignature(ctx context.Context, ref string, sig Signature) error
    ResolveDigest(ctx context.Context, ref string) (digest string, err error) // pin, never tag
}

// Rate limiting at the Gateway.
type RateLimiter interface {
    Allow(ctx context.Context, key string) (bool, RetryAfter) // key = principal|profile|tool
}

HTTP/API surface:

  • GET /api/v1/audit?principal=&action=&target=&from=&to= — query the append-only log (RBAC-gated; read-only, no delete/edit endpoint by design).
  • GET /api/v1/security/posture — summary: TLS status, sandbox defaults, unsigned-package count, recent denials.
  • GET /metrics — security-relevant counters (auth failures, denials, rate-limit drops) for Metrics.

Data

  • Writes/reads the Audit store (Architecture §8) — append-only; no update/delete path exists in code, and integrity may be reinforced by a hash chain over events.
  • Reads signature/trust anchors and digests from the Package Registry before install.
  • Reads the SandboxSpec defaults from the Config store; recipes may request (justified) deviations.
  • Emits security counters to Metrics and security transitions to Notifications.

Dependencies

Failure modes & handling

Failure Behavior
Unsigned or signature-invalid package Blocked before install; never becomes an MCPInstance; audited + alerted (tenet #5).
Image only offers a mutable tag Resolved to a digest and pinned; if no digest can be resolved, install fails closed.
Container cannot run read-only rootfs Recipe must declare the writable paths (tmpfs/volume) explicitly; unjustified full-write rootfs is refused.
Recipe requests host network / extra caps Requires explicit justification in the recipe; flagged in security/posture; audited on install.
Audit sink unavailable Mutations/tool calls fail closed (no silent unaudited actions) or buffer to a durable local queue — never proceed unlogged.
Rate limit exceeded Request rejected with 429/retry-after; drop counted; repeated abuse alerts.
TLS misconfigured / plaintext Gateway refuses to serve the data plane in the clear; fails closed at boot.
A single layer bypassed (e.g. a tool leaks a name) Next layer contains it: RBAC still denies the call, Secrets still redacts, audit still records — no single failure is a breach.

Security notes

This module is the security notes for the system; the layering is the point (Architecture §9):

  1. Transport — TLS terminates at the Gateway; internal calls are localhost/socket only.
  2. Identity — Authentication resolves every request to a Principal; fail closed if unresolved.
  3. Binding — AI Agent Profiles map an agent identity to roles.
  4. Authorization — RBAC dual-gates tools/list and tools/call; default deny.
  5. Secrets — Secrets inject at runtime, never to agents/logs; envelope-encrypted at rest.
  6. Isolation — every MCP runs least-privilege, sandboxed, network-constrained.
  7. Supply chain — signature-verified, digest-pinned packages only.
  8. Observability — every action append-only audited; rate-limited; metered.

Each layer assumes the ones above it can fail. Secure defaults are non-optional: unauthenticated ⇒ no access, ungranted ⇒ no tool, unsigned ⇒ no install, unlogged ⇒ no action.

Open questions

  • Audit integrity: is a hash-chained/tamper-evident log worth the write cost for v1, or is append-only-by-code enough?
  • Do we ship default seccomp/AppArmor profiles per MCP archetype, or a single conservative baseline?
  • Rate-limit granularity default: per-principal, per-profile, per-tool, or a composite — and where are limits configured?
  • How do we attest the sandbox actually applied (verify the running container's security context vs. the spec)?
  • Signing/trust roots for the community recipe & package ecosystem (mirrors Architecture §12).

Milestone

Delivered in Phase 3 (Secure). Thin slice first: container sandboxing defaults (non-root, dropped caps, read-only rootfs where possible, scoped network), append-only audit of every tool call + control-plane mutation, signed-package verification with digest pinning, rate limiting, and TLS termination at the Gateway — composing AuthN/RBAC/Secrets/Profiles into defense-in-depth. Underwrites the whole Phase 3 exit proof: differentiated multi-agent access with a secret-backed MCP that never leaks the secret to any agent-visible payload or log.