MCP Nexus is a self-hosted control plane for MCP servers — "Kubernetes for
MCP." Agents connect to one endpoint; Nexus discovers services, installs the
right MCP servers, aggregates them behind a namespaced router, secures access
with auth/RBAC, and heals/updates them via a reconcile loop.
This first commit is design-phase only (no runnable code yet):
- README.md project front door + module map
- docs/ARCHITECTURE.md target design: tenets, two-plane split, reconcile
loop, domain types, storage, security, deployment
- docs/ROADMAP.md phased delivery (foundations -> walking skeleton ->
discover+install -> secure -> operate -> extend/scale)
- docs/modules/01-20 one design doc per module, all cross-linked
Backend stack decision: Go (single static binary, embedded SQLite + embedded
React dashboard). Repo initialized in /root with a whitelist .gitignore so
only project files are tracked.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
130 lines
8.9 KiB
Markdown
130 lines
8.9 KiB
Markdown
# RBAC
|
|
|
|
> Module 08 · Plane: Data · Roadmap phase: 3
|
|
> Part of [MCP Nexus architecture](../ARCHITECTURE.md).
|
|
|
|
## Purpose
|
|
RBAC (role-based access control) answers the question [Authentication](06-authentication.md) does not: *"Now that we know who is calling, what may they see and do?"* It maps a verified `Principal` to a set of **roles**, and each role to a set of allowed MCP **namespaces and tools**. The resolution collapses to one artifact the hot path cares about: the principal's **allowed tool set**. This is the concrete realization of Architecture §9 "request scoping" and tenet #5 — an agent may not call, *or even see*, a tool outside its role.
|
|
|
|
Example roles ship as a starter pack, each exposing a **different** slice of the tool catalog:
|
|
|
|
| Role | Sees (illustrative namespaces/tools) |
|
|
|---|---|
|
|
| Developer | `github.*`, `postgres.*`, `filesystem.*`, `docker.*` |
|
|
| Home Automation | `homeassistant.*`, `frigate.*`, `mqtt.*` |
|
|
| Networking | `unifi.*`, `pihole.*`, `traefik.*` |
|
|
| Infrastructure | `proxmox.*`, `docker.*`, `prometheus.*`, `grafana.*` |
|
|
| Finance | `finance.*`, curated read-only reporting tools |
|
|
| Admin | `*` (all namespaces) + control-plane mutations |
|
|
|
|
## Responsibilities
|
|
- Own the **Role** and **grant** model (Architecture §7): a role is a named set of allow rules over namespaces/tools; principals hold one or more roles (directly, or via IdP group claims).
|
|
- **Resolve** a `Principal` → roles → the union of allowed `{namespace}.{tool}` patterns → a concrete **allowed tool set**, evaluated against the current [Dynamic Tool Registry](05-dynamic-tool-registry.md).
|
|
- Provide two enforcement primitives the [Gateway](04-gateway.md)/Router calls on the hot path:
|
|
- **`Filter`** — at `tools/list`, return only the tools the principal may see.
|
|
- **`Can`** — at `tools/call`, decide whether this exact tool invocation is permitted.
|
|
- Support **wildcard and namespace-level grants** (`homeassistant.*`) as well as tool-level grants (`postgres.query`) and explicit denies (deny overrides allow).
|
|
- Cache resolved tool sets per principal with invalidation on role change, tool-catalog change, or session end.
|
|
- Reconcile with [AI Agent Profiles](14-agent-profiles.md): a profile is a binding of an agent identity to roles; RBAC is the engine that turns those roles into the visible set.
|
|
- Audit every enforcement decision, especially denials, with principal, tool, and outcome (Architecture §9).
|
|
|
|
## Non-goals
|
|
- **Authentication.** Establishing the `Principal` is [Authentication](06-authentication.md); RBAC trusts the resolved identity.
|
|
- **The agent-facing policy object.** Per-agent identity binding is [AI Agent Profiles](14-agent-profiles.md); RBAC is the underlying role engine both users and profiles share.
|
|
- **Secret access policy.** Whether a secret may be revealed is enforced by the [Secrets Manager](07-secrets-manager.md) (though it may consult roles); RBAC governs tool visibility/invocation.
|
|
- **Multi-tenancy above roles.** Projects/tenants are deferred (Architecture §12); v1 is RBAC only.
|
|
- **Row/field-level data authz inside a tool's result.** RBAC gates the *tool*, not the upstream service's internal permissions.
|
|
|
|
## Interfaces
|
|
```go
|
|
// A single allow/deny rule over the namespaced tool space.
|
|
type Rule struct {
|
|
Effect Effect // Allow | Deny (Deny wins on conflict)
|
|
Pattern string // "homeassistant.*", "postgres.query", "*"
|
|
}
|
|
|
|
type Effect string
|
|
|
|
const (
|
|
Allow Effect = "allow"
|
|
Deny Effect = "deny"
|
|
)
|
|
|
|
// A Role is a named set of rules (Architecture §7).
|
|
type Role struct {
|
|
Name string
|
|
Rules []Rule
|
|
}
|
|
|
|
// The allowed set resolved for a principal at a point in time.
|
|
type AllowedSet struct {
|
|
Principal string
|
|
Tools map[string]bool // fully-qualified "ns.tool" -> allowed
|
|
Roles []string
|
|
ResolvedAt time.Time
|
|
}
|
|
|
|
type Enforcer interface {
|
|
// Resolve principal -> roles -> concrete allowed tool set (against the live catalog).
|
|
Resolve(ctx context.Context, p Principal) (AllowedSet, error)
|
|
// Filter is called at tools/list: keep only visible tools.
|
|
Filter(ctx context.Context, p Principal, tools []Tool) ([]Tool, error)
|
|
// Can is called at tools/call: may this principal invoke this tool?
|
|
Can(ctx context.Context, p Principal, tool string) (bool, error)
|
|
}
|
|
|
|
type RoleStore interface {
|
|
Upsert(ctx context.Context, r Role) error
|
|
Get(ctx context.Context, name string) (Role, error)
|
|
List(ctx context.Context) ([]Role, error)
|
|
AssignRoles(ctx context.Context, principal string, roles []string) error
|
|
}
|
|
```
|
|
|
|
HTTP/API surface (control-plane API, behind Auth, RBAC-gated — Admin only):
|
|
- `GET /api/v1/roles` · `POST /api/v1/roles` · `PUT /api/v1/roles/{name}` · `DELETE /api/v1/roles/{name}`.
|
|
- `POST /api/v1/principals/{id}/roles` — assign/unassign roles to a user or agent profile.
|
|
- `GET /api/v1/principals/{id}/allowed` — debug view: the resolved allowed tool set.
|
|
|
|
MCP data-plane integration (not new endpoints — hooks inside the Router):
|
|
- On `tools/list` → `Enforcer.Filter` removes disallowed tools **before the response is serialized**.
|
|
- On `tools/call` → `Enforcer.Can` gates invocation; a disallowed call returns a JSON-RPC error (as if the tool did not exist), never reaching the upstream client.
|
|
|
|
## Data
|
|
- **Reads/writes** `Role` definitions and principal→role assignments in the **Identity** store (Architecture §8), alongside `Users` and `AgentProfiles`.
|
|
- **Reads** the live tool catalog from the [Dynamic Tool Registry](05-dynamic-tool-registry.md) to expand wildcard grants into concrete `ns.tool` entries.
|
|
- **Writes** authorization decisions (esp. denials) to the **Audit** store.
|
|
- Module-local: an in-memory, TTL'd cache of `AllowedSet` per principal, invalidated on role edits and catalog changes via the event bus.
|
|
|
|
## Dependencies
|
|
- Consumes the `Principal` from [Authentication](06-authentication.md).
|
|
- Enforced inside the [Gateway](04-gateway.md)/Router at both `tools/list` and `tools/call`.
|
|
- Expands grants against the [Dynamic Tool Registry](05-dynamic-tool-registry.md).
|
|
- Bound to agents by [AI Agent Profiles](14-agent-profiles.md) (identity → roles).
|
|
- Managed through the [Web Dashboard](12-web-dashboard.md) (roles, assignments).
|
|
- Decisions audited per [Security](19-security.md).
|
|
|
|
## Failure modes & handling
|
|
| Failure | Behavior |
|
|
|---|---|
|
|
| Principal has no roles | Empty allowed set → sees zero tools, can call nothing. **Fail closed** (default deny). |
|
|
| Role references a namespace/tool that no longer exists | Wildcard resolves against the live catalog; stale explicit grants are simply inert (no error, no phantom access). |
|
|
| Allow and Deny both match a tool | **Deny wins** — explicit denies always override allows. |
|
|
| Tool catalog changes mid-session | Cached `AllowedSet` invalidated on the catalog-change event; next `tools/list`/`tools/call` re-resolves. New tools are not auto-visible until re-resolution. |
|
|
| Role edited while an agent is connected | Cache invalidated; the change takes effect on the next request without dropping the connection. |
|
|
| RBAC store unavailable | Fail closed — deny rather than fall back to allow; surfaced via [Notifications](13-notifications.md). |
|
|
| A denied `tools/call` | Returned as a JSON-RPC method-not-found / not-permitted error; never forwarded upstream; audited. |
|
|
|
|
## Security notes
|
|
Honors Architecture §9 and tenet #5. Enforcement is **dual-gate and mandatory**: `tools/list` filtering means a disallowed tool is invisible (no information leak about what exists), and `tools/call` gating means even a guessed tool name cannot be invoked — the two together defeat both enumeration and direct-call bypass. Default is **deny**: absence of a grant is denial, and explicit `Deny` overrides any allow. Wildcards expand against the live catalog so a newly installed tool is not silently exposed to an over-broad role without re-resolution. All role mutations and all denials are audited with principal + tool. RBAC is enforced in the Router (server side), never the client — the agent's view is a *consequence* of enforcement, not the enforcement itself.
|
|
|
|
## Open questions
|
|
- Role composition: flat roles only, or hierarchical/inheritable roles for large orgs?
|
|
- Do we need per-tool **parameter** constraints (e.g. `postgres.query` read-only), or is tool-level granularity enough for v1? (Leaning: tool-level v1.)
|
|
- How do IdP group claims map onto Nexus roles — static mapping table, or claim-expression rules?
|
|
- Should denies be expressible per-principal (an override) as well as per-role?
|
|
- Time-bounded / just-in-time role grants for elevated access windows?
|
|
|
|
## Milestone
|
|
Delivered in **Phase 3** (Secure). Thin slice first: the starter role pack (Developer, Home Automation, Admin, …), principal→role assignment, and **dual-gate enforcement in the Router** at `tools/list` and `tools/call` with default-deny. Directly powers the Phase 3 exit proof: two agents with different roles connect and each sees a *different* tool set, with disallowed tools neither visible nor callable.
|