Files
mcp-gateway-nexus/docs/modules/08-rbac.md
drjones 8d3ffef920 docs: initial architecture and design for MCP Nexus
MCP Nexus is a self-hosted control plane for MCP servers — "Kubernetes for
MCP." Agents connect to one endpoint; Nexus discovers services, installs the
right MCP servers, aggregates them behind a namespaced router, secures access
with auth/RBAC, and heals/updates them via a reconcile loop.

This first commit is design-phase only (no runnable code yet):
- README.md            project front door + module map
- docs/ARCHITECTURE.md target design: tenets, two-plane split, reconcile
                       loop, domain types, storage, security, deployment
- docs/ROADMAP.md      phased delivery (foundations -> walking skeleton ->
                       discover+install -> secure -> operate -> extend/scale)
- docs/modules/01-20   one design doc per module, all cross-linked

Backend stack decision: Go (single static binary, embedded SQLite + embedded
React dashboard). Repo initialized in /root with a whitelist .gitignore so
only project files are tracked.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-07 04:38:27 +00:00

130 lines
8.9 KiB
Markdown

# RBAC
> Module 08 · Plane: Data · Roadmap phase: 3
> Part of [MCP Nexus architecture](../ARCHITECTURE.md).
## Purpose
RBAC (role-based access control) answers the question [Authentication](06-authentication.md) does not: *"Now that we know who is calling, what may they see and do?"* It maps a verified `Principal` to a set of **roles**, and each role to a set of allowed MCP **namespaces and tools**. The resolution collapses to one artifact the hot path cares about: the principal's **allowed tool set**. This is the concrete realization of Architecture §9 "request scoping" and tenet #5 — an agent may not call, *or even see*, a tool outside its role.
Example roles ship as a starter pack, each exposing a **different** slice of the tool catalog:
| Role | Sees (illustrative namespaces/tools) |
|---|---|
| Developer | `github.*`, `postgres.*`, `filesystem.*`, `docker.*` |
| Home Automation | `homeassistant.*`, `frigate.*`, `mqtt.*` |
| Networking | `unifi.*`, `pihole.*`, `traefik.*` |
| Infrastructure | `proxmox.*`, `docker.*`, `prometheus.*`, `grafana.*` |
| Finance | `finance.*`, curated read-only reporting tools |
| Admin | `*` (all namespaces) + control-plane mutations |
## Responsibilities
- Own the **Role** and **grant** model (Architecture §7): a role is a named set of allow rules over namespaces/tools; principals hold one or more roles (directly, or via IdP group claims).
- **Resolve** a `Principal` → roles → the union of allowed `{namespace}.{tool}` patterns → a concrete **allowed tool set**, evaluated against the current [Dynamic Tool Registry](05-dynamic-tool-registry.md).
- Provide two enforcement primitives the [Gateway](04-gateway.md)/Router calls on the hot path:
- **`Filter`** — at `tools/list`, return only the tools the principal may see.
- **`Can`** — at `tools/call`, decide whether this exact tool invocation is permitted.
- Support **wildcard and namespace-level grants** (`homeassistant.*`) as well as tool-level grants (`postgres.query`) and explicit denies (deny overrides allow).
- Cache resolved tool sets per principal with invalidation on role change, tool-catalog change, or session end.
- Reconcile with [AI Agent Profiles](14-agent-profiles.md): a profile is a binding of an agent identity to roles; RBAC is the engine that turns those roles into the visible set.
- Audit every enforcement decision, especially denials, with principal, tool, and outcome (Architecture §9).
## Non-goals
- **Authentication.** Establishing the `Principal` is [Authentication](06-authentication.md); RBAC trusts the resolved identity.
- **The agent-facing policy object.** Per-agent identity binding is [AI Agent Profiles](14-agent-profiles.md); RBAC is the underlying role engine both users and profiles share.
- **Secret access policy.** Whether a secret may be revealed is enforced by the [Secrets Manager](07-secrets-manager.md) (though it may consult roles); RBAC governs tool visibility/invocation.
- **Multi-tenancy above roles.** Projects/tenants are deferred (Architecture §12); v1 is RBAC only.
- **Row/field-level data authz inside a tool's result.** RBAC gates the *tool*, not the upstream service's internal permissions.
## Interfaces
```go
// A single allow/deny rule over the namespaced tool space.
type Rule struct {
Effect Effect // Allow | Deny (Deny wins on conflict)
Pattern string // "homeassistant.*", "postgres.query", "*"
}
type Effect string
const (
Allow Effect = "allow"
Deny Effect = "deny"
)
// A Role is a named set of rules (Architecture §7).
type Role struct {
Name string
Rules []Rule
}
// The allowed set resolved for a principal at a point in time.
type AllowedSet struct {
Principal string
Tools map[string]bool // fully-qualified "ns.tool" -> allowed
Roles []string
ResolvedAt time.Time
}
type Enforcer interface {
// Resolve principal -> roles -> concrete allowed tool set (against the live catalog).
Resolve(ctx context.Context, p Principal) (AllowedSet, error)
// Filter is called at tools/list: keep only visible tools.
Filter(ctx context.Context, p Principal, tools []Tool) ([]Tool, error)
// Can is called at tools/call: may this principal invoke this tool?
Can(ctx context.Context, p Principal, tool string) (bool, error)
}
type RoleStore interface {
Upsert(ctx context.Context, r Role) error
Get(ctx context.Context, name string) (Role, error)
List(ctx context.Context) ([]Role, error)
AssignRoles(ctx context.Context, principal string, roles []string) error
}
```
HTTP/API surface (control-plane API, behind Auth, RBAC-gated — Admin only):
- `GET /api/v1/roles` · `POST /api/v1/roles` · `PUT /api/v1/roles/{name}` · `DELETE /api/v1/roles/{name}`.
- `POST /api/v1/principals/{id}/roles` — assign/unassign roles to a user or agent profile.
- `GET /api/v1/principals/{id}/allowed` — debug view: the resolved allowed tool set.
MCP data-plane integration (not new endpoints — hooks inside the Router):
- On `tools/list` → `Enforcer.Filter` removes disallowed tools **before the response is serialized**.
- On `tools/call` → `Enforcer.Can` gates invocation; a disallowed call returns a JSON-RPC error (as if the tool did not exist), never reaching the upstream client.
## Data
- **Reads/writes** `Role` definitions and principal→role assignments in the **Identity** store (Architecture §8), alongside `Users` and `AgentProfiles`.
- **Reads** the live tool catalog from the [Dynamic Tool Registry](05-dynamic-tool-registry.md) to expand wildcard grants into concrete `ns.tool` entries.
- **Writes** authorization decisions (esp. denials) to the **Audit** store.
- Module-local: an in-memory, TTL'd cache of `AllowedSet` per principal, invalidated on role edits and catalog changes via the event bus.
## Dependencies
- Consumes the `Principal` from [Authentication](06-authentication.md).
- Enforced inside the [Gateway](04-gateway.md)/Router at both `tools/list` and `tools/call`.
- Expands grants against the [Dynamic Tool Registry](05-dynamic-tool-registry.md).
- Bound to agents by [AI Agent Profiles](14-agent-profiles.md) (identity → roles).
- Managed through the [Web Dashboard](12-web-dashboard.md) (roles, assignments).
- Decisions audited per [Security](19-security.md).
## Failure modes & handling
| Failure | Behavior |
|---|---|
| Principal has no roles | Empty allowed set → sees zero tools, can call nothing. **Fail closed** (default deny). |
| Role references a namespace/tool that no longer exists | Wildcard resolves against the live catalog; stale explicit grants are simply inert (no error, no phantom access). |
| Allow and Deny both match a tool | **Deny wins** — explicit denies always override allows. |
| Tool catalog changes mid-session | Cached `AllowedSet` invalidated on the catalog-change event; next `tools/list`/`tools/call` re-resolves. New tools are not auto-visible until re-resolution. |
| Role edited while an agent is connected | Cache invalidated; the change takes effect on the next request without dropping the connection. |
| RBAC store unavailable | Fail closed — deny rather than fall back to allow; surfaced via [Notifications](13-notifications.md). |
| A denied `tools/call` | Returned as a JSON-RPC method-not-found / not-permitted error; never forwarded upstream; audited. |
## Security notes
Honors Architecture §9 and tenet #5. Enforcement is **dual-gate and mandatory**: `tools/list` filtering means a disallowed tool is invisible (no information leak about what exists), and `tools/call` gating means even a guessed tool name cannot be invoked — the two together defeat both enumeration and direct-call bypass. Default is **deny**: absence of a grant is denial, and explicit `Deny` overrides any allow. Wildcards expand against the live catalog so a newly installed tool is not silently exposed to an over-broad role without re-resolution. All role mutations and all denials are audited with principal + tool. RBAC is enforced in the Router (server side), never the client — the agent's view is a *consequence* of enforcement, not the enforcement itself.
## Open questions
- Role composition: flat roles only, or hierarchical/inheritable roles for large orgs?
- Do we need per-tool **parameter** constraints (e.g. `postgres.query` read-only), or is tool-level granularity enough for v1? (Leaning: tool-level v1.)
- How do IdP group claims map onto Nexus roles — static mapping table, or claim-expression rules?
- Should denies be expressible per-principal (an override) as well as per-role?
- Time-bounded / just-in-time role grants for elevated access windows?
## Milestone
Delivered in **Phase 3** (Secure). Thin slice first: the starter role pack (Developer, Home Automation, Admin, …), principal→role assignment, and **dual-gate enforcement in the Router** at `tools/list` and `tools/call` with default-deny. Directly powers the Phase 3 exit proof: two agents with different roles connect and each sees a *different* tool set, with disallowed tools neither visible nor callable.