Files
mcp-gateway-nexus/docs/modules/08-rbac.md
drjones 8d3ffef920 docs: initial architecture and design for MCP Nexus
MCP Nexus is a self-hosted control plane for MCP servers — "Kubernetes for
MCP." Agents connect to one endpoint; Nexus discovers services, installs the
right MCP servers, aggregates them behind a namespaced router, secures access
with auth/RBAC, and heals/updates them via a reconcile loop.

This first commit is design-phase only (no runnable code yet):
- README.md            project front door + module map
- docs/ARCHITECTURE.md target design: tenets, two-plane split, reconcile
                       loop, domain types, storage, security, deployment
- docs/ROADMAP.md      phased delivery (foundations -> walking skeleton ->
                       discover+install -> secure -> operate -> extend/scale)
- docs/modules/01-20   one design doc per module, all cross-linked

Backend stack decision: Go (single static binary, embedded SQLite + embedded
React dashboard). Repo initialized in /root with a whitelist .gitignore so
only project files are tracked.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-07 04:38:27 +00:00

8.9 KiB

RBAC

Module 08 · Plane: Data · Roadmap phase: 3 Part of MCP Nexus architecture.

Purpose

RBAC (role-based access control) answers the question Authentication does not: "Now that we know who is calling, what may they see and do?" It maps a verified Principal to a set of roles, and each role to a set of allowed MCP namespaces and tools. The resolution collapses to one artifact the hot path cares about: the principal's allowed tool set. This is the concrete realization of Architecture §9 "request scoping" and tenet #5 — an agent may not call, or even see, a tool outside its role.

Example roles ship as a starter pack, each exposing a different slice of the tool catalog:

Role Sees (illustrative namespaces/tools)
Developer github.*, postgres.*, filesystem.*, docker.*
Home Automation homeassistant.*, frigate.*, mqtt.*
Networking unifi.*, pihole.*, traefik.*
Infrastructure proxmox.*, docker.*, prometheus.*, grafana.*
Finance finance.*, curated read-only reporting tools
Admin * (all namespaces) + control-plane mutations

Responsibilities

  • Own the Role and grant model (Architecture §7): a role is a named set of allow rules over namespaces/tools; principals hold one or more roles (directly, or via IdP group claims).
  • Resolve a Principal → roles → the union of allowed {namespace}.{tool} patterns → a concrete allowed tool set, evaluated against the current Dynamic Tool Registry.
  • Provide two enforcement primitives the Gateway/Router calls on the hot path:
    • Filter — at tools/list, return only the tools the principal may see.
    • Can — at tools/call, decide whether this exact tool invocation is permitted.
  • Support wildcard and namespace-level grants (homeassistant.*) as well as tool-level grants (postgres.query) and explicit denies (deny overrides allow).
  • Cache resolved tool sets per principal with invalidation on role change, tool-catalog change, or session end.
  • Reconcile with AI Agent Profiles: a profile is a binding of an agent identity to roles; RBAC is the engine that turns those roles into the visible set.
  • Audit every enforcement decision, especially denials, with principal, tool, and outcome (Architecture §9).

Non-goals

  • Authentication. Establishing the Principal is Authentication; RBAC trusts the resolved identity.
  • The agent-facing policy object. Per-agent identity binding is AI Agent Profiles; RBAC is the underlying role engine both users and profiles share.
  • Secret access policy. Whether a secret may be revealed is enforced by the Secrets Manager (though it may consult roles); RBAC governs tool visibility/invocation.
  • Multi-tenancy above roles. Projects/tenants are deferred (Architecture §12); v1 is RBAC only.
  • Row/field-level data authz inside a tool's result. RBAC gates the tool, not the upstream service's internal permissions.

Interfaces

// A single allow/deny rule over the namespaced tool space.
type Rule struct {
    Effect  Effect // Allow | Deny (Deny wins on conflict)
    Pattern string // "homeassistant.*", "postgres.query", "*"
}

type Effect string

const (
    Allow Effect = "allow"
    Deny  Effect = "deny"
)

// A Role is a named set of rules (Architecture §7).
type Role struct {
    Name  string
    Rules []Rule
}

// The allowed set resolved for a principal at a point in time.
type AllowedSet struct {
    Principal string
    Tools     map[string]bool // fully-qualified "ns.tool" -> allowed
    Roles     []string
    ResolvedAt time.Time
}

type Enforcer interface {
    // Resolve principal -> roles -> concrete allowed tool set (against the live catalog).
    Resolve(ctx context.Context, p Principal) (AllowedSet, error)
    // Filter is called at tools/list: keep only visible tools.
    Filter(ctx context.Context, p Principal, tools []Tool) ([]Tool, error)
    // Can is called at tools/call: may this principal invoke this tool?
    Can(ctx context.Context, p Principal, tool string) (bool, error)
}

type RoleStore interface {
    Upsert(ctx context.Context, r Role) error
    Get(ctx context.Context, name string) (Role, error)
    List(ctx context.Context) ([]Role, error)
    AssignRoles(ctx context.Context, principal string, roles []string) error
}

HTTP/API surface (control-plane API, behind Auth, RBAC-gated — Admin only):

  • GET /api/v1/roles · POST /api/v1/roles · PUT /api/v1/roles/{name} · DELETE /api/v1/roles/{name}.
  • POST /api/v1/principals/{id}/roles — assign/unassign roles to a user or agent profile.
  • GET /api/v1/principals/{id}/allowed — debug view: the resolved allowed tool set.

MCP data-plane integration (not new endpoints — hooks inside the Router):

  • On tools/list → Enforcer.Filter removes disallowed tools before the response is serialized.
  • On tools/call → Enforcer.Can gates invocation; a disallowed call returns a JSON-RPC error (as if the tool did not exist), never reaching the upstream client.

Data

  • Reads/writes Role definitions and principal→role assignments in the Identity store (Architecture §8), alongside Users and AgentProfiles.
  • Reads the live tool catalog from the Dynamic Tool Registry to expand wildcard grants into concrete ns.tool entries.
  • Writes authorization decisions (esp. denials) to the Audit store.
  • Module-local: an in-memory, TTL'd cache of AllowedSet per principal, invalidated on role edits and catalog changes via the event bus.

Dependencies

Failure modes & handling

Failure Behavior
Principal has no roles Empty allowed set → sees zero tools, can call nothing. Fail closed (default deny).
Role references a namespace/tool that no longer exists Wildcard resolves against the live catalog; stale explicit grants are simply inert (no error, no phantom access).
Allow and Deny both match a tool Deny wins — explicit denies always override allows.
Tool catalog changes mid-session Cached AllowedSet invalidated on the catalog-change event; next tools/list/tools/call re-resolves. New tools are not auto-visible until re-resolution.
Role edited while an agent is connected Cache invalidated; the change takes effect on the next request without dropping the connection.
RBAC store unavailable Fail closed — deny rather than fall back to allow; surfaced via Notifications.
A denied tools/call Returned as a JSON-RPC method-not-found / not-permitted error; never forwarded upstream; audited.

Security notes

Honors Architecture §9 and tenet #5. Enforcement is dual-gate and mandatory: tools/list filtering means a disallowed tool is invisible (no information leak about what exists), and tools/call gating means even a guessed tool name cannot be invoked — the two together defeat both enumeration and direct-call bypass. Default is deny: absence of a grant is denial, and explicit Deny overrides any allow. Wildcards expand against the live catalog so a newly installed tool is not silently exposed to an over-broad role without re-resolution. All role mutations and all denials are audited with principal + tool. RBAC is enforced in the Router (server side), never the client — the agent's view is a consequence of enforcement, not the enforcement itself.

Open questions

  • Role composition: flat roles only, or hierarchical/inheritable roles for large orgs?
  • Do we need per-tool parameter constraints (e.g. postgres.query read-only), or is tool-level granularity enough for v1? (Leaning: tool-level v1.)
  • How do IdP group claims map onto Nexus roles — static mapping table, or claim-expression rules?
  • Should denies be expressible per-principal (an override) as well as per-role?
  • Time-bounded / just-in-time role grants for elevated access windows?

Milestone

Delivered in Phase 3 (Secure). Thin slice first: the starter role pack (Developer, Home Automation, Admin, …), principal→role assignment, and dual-gate enforcement in the Router at tools/list and tools/call with default-deny. Directly powers the Phase 3 exit proof: two agents with different roles connect and each sees a different tool set, with disallowed tools neither visible nor callable.