MCP Nexus is a self-hosted control plane for MCP servers — "Kubernetes for
MCP." Agents connect to one endpoint; Nexus discovers services, installs the
right MCP servers, aggregates them behind a namespaced router, secures access
with auth/RBAC, and heals/updates them via a reconcile loop.
This first commit is design-phase only (no runnable code yet):
- README.md project front door + module map
- docs/ARCHITECTURE.md target design: tenets, two-plane split, reconcile
loop, domain types, storage, security, deployment
- docs/ROADMAP.md phased delivery (foundations -> walking skeleton ->
discover+install -> secure -> operate -> extend/scale)
- docs/modules/01-20 one design doc per module, all cross-linked
Backend stack decision: Go (single static binary, embedded SQLite + embedded
React dashboard). Repo initialized in /root with a whitelist .gitignore so
only project files are tracked.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
8.9 KiB
RBAC
Module 08 · Plane: Data · Roadmap phase: 3 Part of MCP Nexus architecture.
Purpose
RBAC (role-based access control) answers the question Authentication does not: "Now that we know who is calling, what may they see and do?" It maps a verified Principal to a set of roles, and each role to a set of allowed MCP namespaces and tools. The resolution collapses to one artifact the hot path cares about: the principal's allowed tool set. This is the concrete realization of Architecture §9 "request scoping" and tenet #5 — an agent may not call, or even see, a tool outside its role.
Example roles ship as a starter pack, each exposing a different slice of the tool catalog:
| Role | Sees (illustrative namespaces/tools) |
|---|---|
| Developer | github.*, postgres.*, filesystem.*, docker.* |
| Home Automation | homeassistant.*, frigate.*, mqtt.* |
| Networking | unifi.*, pihole.*, traefik.* |
| Infrastructure | proxmox.*, docker.*, prometheus.*, grafana.* |
| Finance | finance.*, curated read-only reporting tools |
| Admin | * (all namespaces) + control-plane mutations |
Responsibilities
- Own the Role and grant model (Architecture §7): a role is a named set of allow rules over namespaces/tools; principals hold one or more roles (directly, or via IdP group claims).
- Resolve a
Principal→ roles → the union of allowed{namespace}.{tool}patterns → a concrete allowed tool set, evaluated against the current Dynamic Tool Registry. - Provide two enforcement primitives the Gateway/Router calls on the hot path:
Filter— attools/list, return only the tools the principal may see.Can— attools/call, decide whether this exact tool invocation is permitted.
- Support wildcard and namespace-level grants (
homeassistant.*) as well as tool-level grants (postgres.query) and explicit denies (deny overrides allow). - Cache resolved tool sets per principal with invalidation on role change, tool-catalog change, or session end.
- Reconcile with AI Agent Profiles: a profile is a binding of an agent identity to roles; RBAC is the engine that turns those roles into the visible set.
- Audit every enforcement decision, especially denials, with principal, tool, and outcome (Architecture §9).
Non-goals
- Authentication. Establishing the
Principalis Authentication; RBAC trusts the resolved identity. - The agent-facing policy object. Per-agent identity binding is AI Agent Profiles; RBAC is the underlying role engine both users and profiles share.
- Secret access policy. Whether a secret may be revealed is enforced by the Secrets Manager (though it may consult roles); RBAC governs tool visibility/invocation.
- Multi-tenancy above roles. Projects/tenants are deferred (Architecture §12); v1 is RBAC only.
- Row/field-level data authz inside a tool's result. RBAC gates the tool, not the upstream service's internal permissions.
Interfaces
// A single allow/deny rule over the namespaced tool space.
type Rule struct {
Effect Effect // Allow | Deny (Deny wins on conflict)
Pattern string // "homeassistant.*", "postgres.query", "*"
}
type Effect string
const (
Allow Effect = "allow"
Deny Effect = "deny"
)
// A Role is a named set of rules (Architecture §7).
type Role struct {
Name string
Rules []Rule
}
// The allowed set resolved for a principal at a point in time.
type AllowedSet struct {
Principal string
Tools map[string]bool // fully-qualified "ns.tool" -> allowed
Roles []string
ResolvedAt time.Time
}
type Enforcer interface {
// Resolve principal -> roles -> concrete allowed tool set (against the live catalog).
Resolve(ctx context.Context, p Principal) (AllowedSet, error)
// Filter is called at tools/list: keep only visible tools.
Filter(ctx context.Context, p Principal, tools []Tool) ([]Tool, error)
// Can is called at tools/call: may this principal invoke this tool?
Can(ctx context.Context, p Principal, tool string) (bool, error)
}
type RoleStore interface {
Upsert(ctx context.Context, r Role) error
Get(ctx context.Context, name string) (Role, error)
List(ctx context.Context) ([]Role, error)
AssignRoles(ctx context.Context, principal string, roles []string) error
}
HTTP/API surface (control-plane API, behind Auth, RBAC-gated — Admin only):
GET /api/v1/roles·POST /api/v1/roles·PUT /api/v1/roles/{name}·DELETE /api/v1/roles/{name}.POST /api/v1/principals/{id}/roles— assign/unassign roles to a user or agent profile.GET /api/v1/principals/{id}/allowed— debug view: the resolved allowed tool set.
MCP data-plane integration (not new endpoints — hooks inside the Router):
- On
tools/list→Enforcer.Filterremoves disallowed tools before the response is serialized. - On
tools/call→Enforcer.Cangates invocation; a disallowed call returns a JSON-RPC error (as if the tool did not exist), never reaching the upstream client.
Data
- Reads/writes
Roledefinitions and principal→role assignments in the Identity store (Architecture §8), alongsideUsersandAgentProfiles. - Reads the live tool catalog from the Dynamic Tool Registry to expand wildcard grants into concrete
ns.toolentries. - Writes authorization decisions (esp. denials) to the Audit store.
- Module-local: an in-memory, TTL'd cache of
AllowedSetper principal, invalidated on role edits and catalog changes via the event bus.
Dependencies
- Consumes the
Principalfrom Authentication. - Enforced inside the Gateway/Router at both
tools/listandtools/call. - Expands grants against the Dynamic Tool Registry.
- Bound to agents by AI Agent Profiles (identity → roles).
- Managed through the Web Dashboard (roles, assignments).
- Decisions audited per Security.
Failure modes & handling
| Failure | Behavior |
|---|---|
| Principal has no roles | Empty allowed set → sees zero tools, can call nothing. Fail closed (default deny). |
| Role references a namespace/tool that no longer exists | Wildcard resolves against the live catalog; stale explicit grants are simply inert (no error, no phantom access). |
| Allow and Deny both match a tool | Deny wins — explicit denies always override allows. |
| Tool catalog changes mid-session | Cached AllowedSet invalidated on the catalog-change event; next tools/list/tools/call re-resolves. New tools are not auto-visible until re-resolution. |
| Role edited while an agent is connected | Cache invalidated; the change takes effect on the next request without dropping the connection. |
| RBAC store unavailable | Fail closed — deny rather than fall back to allow; surfaced via Notifications. |
A denied tools/call |
Returned as a JSON-RPC method-not-found / not-permitted error; never forwarded upstream; audited. |
Security notes
Honors Architecture §9 and tenet #5. Enforcement is dual-gate and mandatory: tools/list filtering means a disallowed tool is invisible (no information leak about what exists), and tools/call gating means even a guessed tool name cannot be invoked — the two together defeat both enumeration and direct-call bypass. Default is deny: absence of a grant is denial, and explicit Deny overrides any allow. Wildcards expand against the live catalog so a newly installed tool is not silently exposed to an over-broad role without re-resolution. All role mutations and all denials are audited with principal + tool. RBAC is enforced in the Router (server side), never the client — the agent's view is a consequence of enforcement, not the enforcement itself.
Open questions
- Role composition: flat roles only, or hierarchical/inheritable roles for large orgs?
- Do we need per-tool parameter constraints (e.g.
postgres.queryread-only), or is tool-level granularity enough for v1? (Leaning: tool-level v1.) - How do IdP group claims map onto Nexus roles — static mapping table, or claim-expression rules?
- Should denies be expressible per-principal (an override) as well as per-role?
- Time-bounded / just-in-time role grants for elevated access windows?
Milestone
Delivered in Phase 3 (Secure). Thin slice first: the starter role pack (Developer, Home Automation, Admin, …), principal→role assignment, and dual-gate enforcement in the Router at tools/list and tools/call with default-deny. Directly powers the Phase 3 exit proof: two agents with different roles connect and each sees a different tool set, with disallowed tools neither visible nor callable.