MCP Nexus is a self-hosted control plane for MCP servers — "Kubernetes for
MCP." Agents connect to one endpoint; Nexus discovers services, installs the
right MCP servers, aggregates them behind a namespaced router, secures access
with auth/RBAC, and heals/updates them via a reconcile loop.
This first commit is design-phase only (no runnable code yet):
- README.md project front door + module map
- docs/ARCHITECTURE.md target design: tenets, two-plane split, reconcile
loop, domain types, storage, security, deployment
- docs/ROADMAP.md phased delivery (foundations -> walking skeleton ->
discover+install -> secure -> operate -> extend/scale)
- docs/modules/01-20 one design doc per module, all cross-linked
Backend stack decision: Go (single static binary, embedded SQLite + embedded
React dashboard). Repo initialized in /root with a whitelist .gitignore so
only project files are tracked.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
121 lines
9.4 KiB
Markdown
121 lines
9.4 KiB
Markdown
# Security
|
|
|
|
> Module 19 · Plane: Cross-cutting · Roadmap phase: 3
|
|
> Part of [MCP Nexus architecture](../ARCHITECTURE.md).
|
|
|
|
## Purpose
|
|
Security is the umbrella model that makes tenet #5 ("secure by default") real and gives Architecture §9 its implementation. It is not one component but the **defense-in-depth composition** of the security-relevant modules — [Authentication](06-authentication.md), [RBAC](08-rbac.md), [Secrets](07-secrets-manager.md), [AI Agent Profiles](14-agent-profiles.md) — plus the cross-cutting controls no single module owns: **container sandboxing**, **append-only audit**, **signed-package/supply-chain verification**, **rate limiting**, and **TLS everywhere**. This doc defines those cross-cutting controls and how the layers stack so a failure of any one layer is not a breach. Reference: Architecture §9.
|
|
|
|
## Responsibilities
|
|
- **Sandbox every MCP instance.** Each `MCPInstance` runs in its own container with least privilege: dropped Linux capabilities, `no-new-privileges`, non-root user, **read-only rootfs** where the image allows, seccomp/AppArmor profiles, constrained networking (no host network unless the recipe demands it), and CPU/memory limits.
|
|
- **Container isolation defaults.** Define the baseline security context the [Auto Installer](03-auto-installer.md)/Runtime applies to every container so sandboxing is the default, not an opt-in.
|
|
- **Audit every action.** Provide the append-only audit sink and schema for *both* control-plane mutations and *every* agent tool call, recorded with principal, target, and outcome (Architecture §9).
|
|
- **Supply-chain verification.** Own the trust model for signed packages: verify signatures before install and **pin by digest** (never by mutable tag), in concert with the [Package Registry](02-package-registry.md).
|
|
- **Rate limiting.** Provide per-principal / per-profile / per-tool rate limits at the [Gateway](04-gateway.md) so one agent cannot starve others or abuse an upstream (tenet #1 fairness).
|
|
- **TLS everywhere.** Define TLS termination at the Gateway and localhost/socket-only internal component calls.
|
|
- **Compose the layers.** Specify how AuthN → Profiles → RBAC → Secrets → sandbox → audit stack into defense-in-depth, and how a bypass of one is contained by the next.
|
|
|
|
## Non-goals
|
|
- **Implementing the identity/policy engines.** AuthN ([06](06-authentication.md)), RBAC ([08](08-rbac.md)), Secrets ([07](07-secrets-manager.md)), Profiles ([14](14-agent-profiles.md)) own their mechanisms; this module composes them and owns the *cross-cutting* controls.
|
|
- **Running containers.** The [Auto Installer](03-auto-installer.md)/Runtime creates containers; this module defines the security context they must apply.
|
|
- **Being a SIEM.** Nexus emits an append-only audit log and metrics; long-term aggregation/alerting integrates via [Notifications](13-notifications.md) and external tooling.
|
|
- **Guaranteeing upstream service security.** Nexus sandboxes the MCP server; the real service behind it (Postgres, Home Assistant) has its own posture.
|
|
|
|
## Interfaces
|
|
```go
|
|
// The security context every managed container must be created with (Runtime applies it).
|
|
type SandboxSpec struct {
|
|
ReadOnlyRootfs bool
|
|
RunAsNonRoot bool
|
|
DropCaps []string // e.g. ["ALL"]; AddCaps only if a recipe justifies it
|
|
AddCaps []string
|
|
NoNewPrivs bool
|
|
Seccomp string // profile name/path
|
|
AppArmor string
|
|
Network NetMode // "none" | "bridge-scoped" | "host" (host requires justification)
|
|
CPU/*limits*/ Resource
|
|
Memory Resource
|
|
}
|
|
|
|
// Append-only audit sink (Architecture §8 Audit store). Records are immutable.
|
|
type AuditSink interface {
|
|
// Record is called for every mutating control-plane action AND every agent tool call.
|
|
Record(ctx context.Context, e AuditEvent) error
|
|
Query(ctx context.Context, f AuditFilter) ([]AuditEvent, error) // read-only
|
|
}
|
|
|
|
type AuditEvent struct {
|
|
Time time.Time
|
|
Principal string // who (user/agent/service)
|
|
Action string // "tools/call", "role.update", "secret.resolve", ...
|
|
Target string // "postgres.query", "role:developer", "secret://..."
|
|
Outcome string // "allow" | "deny" | "success" | "error"
|
|
Meta map[string]string // never contains secret plaintext
|
|
}
|
|
|
|
// Supply-chain verification (with the Package Registry).
|
|
type Verifier interface {
|
|
VerifySignature(ctx context.Context, ref string, sig Signature) error
|
|
ResolveDigest(ctx context.Context, ref string) (digest string, err error) // pin, never tag
|
|
}
|
|
|
|
// Rate limiting at the Gateway.
|
|
type RateLimiter interface {
|
|
Allow(ctx context.Context, key string) (bool, RetryAfter) // key = principal|profile|tool
|
|
}
|
|
```
|
|
|
|
HTTP/API surface:
|
|
- `GET /api/v1/audit?principal=&action=&target=&from=&to=` — query the append-only log (RBAC-gated; read-only, no delete/edit endpoint by design).
|
|
- `GET /api/v1/security/posture` — summary: TLS status, sandbox defaults, unsigned-package count, recent denials.
|
|
- `GET /metrics` — security-relevant counters (auth failures, denials, rate-limit drops) for [Metrics](18-metrics.md).
|
|
|
|
## Data
|
|
- **Writes/reads** the **Audit** store (Architecture §8) — append-only; no update/delete path exists in code, and integrity may be reinforced by a hash chain over events.
|
|
- **Reads** signature/trust anchors and digests from the [Package Registry](02-package-registry.md) before install.
|
|
- **Reads** the `SandboxSpec` defaults from the **Config** store; recipes may request (justified) deviations.
|
|
- Emits security counters to [Metrics](18-metrics.md) and security transitions to [Notifications](13-notifications.md).
|
|
|
|
## Dependencies
|
|
- Composes [Authentication](06-authentication.md), [RBAC](08-rbac.md), [Secrets](07-secrets-manager.md), and [AI Agent Profiles](14-agent-profiles.md) into defense-in-depth.
|
|
- Sandbox spec applied by the [Auto Installer](03-auto-installer.md)/Runtime.
|
|
- Signature/digest verification with the [Package Registry](02-package-registry.md).
|
|
- Rate limiting + TLS at the [Gateway](04-gateway.md).
|
|
- Security events flow to [Notifications](13-notifications.md) and counters to [Metrics](18-metrics.md).
|
|
|
|
## Failure modes & handling
|
|
| Failure | Behavior |
|
|
|---|---|
|
|
| Unsigned or signature-invalid package | Blocked before install; never becomes an `MCPInstance`; audited + alerted (tenet #5). |
|
|
| Image only offers a mutable tag | Resolved to a digest and pinned; if no digest can be resolved, install fails closed. |
|
|
| Container cannot run read-only rootfs | Recipe must declare the writable paths (tmpfs/volume) explicitly; unjustified full-write rootfs is refused. |
|
|
| Recipe requests host network / extra caps | Requires explicit justification in the recipe; flagged in `security/posture`; audited on install. |
|
|
| Audit sink unavailable | Mutations/tool calls fail closed (no silent unaudited actions) or buffer to a durable local queue — never proceed unlogged. |
|
|
| Rate limit exceeded | Request rejected with `429`/retry-after; drop counted; repeated abuse alerts. |
|
|
| TLS misconfigured / plaintext | Gateway refuses to serve the data plane in the clear; fails closed at boot. |
|
|
| A single layer bypassed (e.g. a tool leaks a name) | Next layer contains it: RBAC still denies the call, Secrets still redacts, audit still records — no single failure is a breach. |
|
|
|
|
## Security notes
|
|
This module *is* the security notes for the system; the layering is the point (Architecture §9):
|
|
|
|
1. **Transport** — TLS terminates at the Gateway; internal calls are localhost/socket only.
|
|
2. **Identity** — [Authentication](06-authentication.md) resolves every request to a `Principal`; fail closed if unresolved.
|
|
3. **Binding** — [AI Agent Profiles](14-agent-profiles.md) map an agent identity to roles.
|
|
4. **Authorization** — [RBAC](08-rbac.md) dual-gates `tools/list` and `tools/call`; default deny.
|
|
5. **Secrets** — [Secrets](07-secrets-manager.md) inject at runtime, never to agents/logs; envelope-encrypted at rest.
|
|
6. **Isolation** — every MCP runs least-privilege, sandboxed, network-constrained.
|
|
7. **Supply chain** — signature-verified, digest-pinned packages only.
|
|
8. **Observability** — every action append-only audited; rate-limited; metered.
|
|
|
|
Each layer assumes the ones above it can fail. Secure defaults are non-optional: unauthenticated ⇒ no access, ungranted ⇒ no tool, unsigned ⇒ no install, unlogged ⇒ no action.
|
|
|
|
## Open questions
|
|
- Audit integrity: is a hash-chained/tamper-evident log worth the write cost for v1, or is append-only-by-code enough?
|
|
- Do we ship default seccomp/AppArmor profiles per MCP archetype, or a single conservative baseline?
|
|
- Rate-limit granularity default: per-principal, per-profile, per-tool, or a composite — and where are limits configured?
|
|
- How do we attest the sandbox actually applied (verify the running container's security context vs. the spec)?
|
|
- Signing/trust roots for the community recipe & package ecosystem (mirrors Architecture §12).
|
|
|
|
## Milestone
|
|
Delivered in **Phase 3** (Secure). Thin slice first: **container sandboxing defaults** (non-root, dropped caps, read-only rootfs where possible, scoped network), **append-only audit** of every tool call + control-plane mutation, **signed-package verification with digest pinning**, **rate limiting**, and **TLS termination** at the Gateway — composing AuthN/RBAC/Secrets/Profiles into defense-in-depth. Underwrites the whole Phase 3 exit proof: differentiated multi-agent access with a secret-backed MCP that never leaks the secret to any agent-visible payload or log.
|