docs: initial architecture and design for MCP Nexus

MCP Nexus is a self-hosted control plane for MCP servers — "Kubernetes for
MCP." Agents connect to one endpoint; Nexus discovers services, installs the
right MCP servers, aggregates them behind a namespaced router, secures access
with auth/RBAC, and heals/updates them via a reconcile loop.

This first commit is design-phase only (no runnable code yet):
- README.md            project front door + module map
- docs/ARCHITECTURE.md target design: tenets, two-plane split, reconcile
                       loop, domain types, storage, security, deployment
- docs/ROADMAP.md      phased delivery (foundations -> walking skeleton ->
                       discover+install -> secure -> operate -> extend/scale)
- docs/modules/01-20   one design doc per module, all cross-linked

Backend stack decision: Go (single static binary, embedded SQLite + embedded
React dashboard). Repo initialized in /root with a whitelist .gitignore so
only project files are tracked.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
drjones
2026-07-07 04:38:27 +00:00
commit 8d3ffef920
24 changed files with 3000 additions and 0 deletions

View File

@@ -0,0 +1,167 @@
# Update Manager
> Module 10 · Plane: Control · Roadmap phase: 4
> Part of [MCP Nexus architecture](../ARCHITECTURE.md).
## Purpose
The Update Manager keeps managed `MCPInstance`s current and safe. It watches
upstream sources for newer versions of the packages an instance was installed
from — Docker image tags/digests, GitHub releases, and
[Registry](02-package-registry.md) index updates — decides whether to act based
on **policy** (automatic, manual-approval-gated, or pinned), and applies updates
as **re-installs** through the [Auto Installer](03-auto-installer.md). Crucially,
it keeps the previous known-good spec so any update can be **rolled back** safely.
It is a day-2 control-plane reconciler (Architecture §5): a failure here degrades
*management currency*, never the data plane. Per tenet #2 it proposes and applies
state transitions through the loop, it does not run ad-hoc upgrade scripts.
## Responsibilities
- **Watch** upstream sources per instance and detect available updates:
- Docker image **tags/digests** (a new digest behind a tracked tag).
- **GitHub releases** for `source=github` packages.
- **Registry index updates** (a package's recommended version advanced).
- Evaluate each candidate against the instance's **update policy**:
`automatic`, `manual` (approval-gated), or `pinned` (never auto-update).
- Enforce **version pinning**: a pinned instance is skipped entirely (still
reported as "update available").
- Apply an approved update by driving a **re-install** at the new digest via the
[Installer](03-auto-installer.md), preserving the same `bound_resource_uuid`.
- Retain the **previous known-good `MCPInstance` spec** and perform **rollback**
when the new version fails to become/stay healthy.
- Coordinate with [Health Monitoring](09-health-monitoring.md) to gate and
validate updates, and with [Notifications](13-notifications.md) to alert on
available / applied / failed / rolled-back updates.
- Record all decisions and transitions to Audit (Architecture §9).
## Non-goals
- Performing the actual pull/create/start — delegated to the
[Installer](03-auto-installer.md) and Runtime.
- Choosing/verifying package signatures or digests — owned by the
[Registry](02-package-registry.md).
- Restarting crashed-but-same-version instances — that is
[Health Monitoring](09-health-monitoring.md) (self-healing), not an update.
- Deciding which package a service uses ([Recipes](15-smart-recipes.md)).
## Interfaces
```go
type UpdatePolicy string
const (
PolicyAuto UpdatePolicy = "automatic" // apply without human action
PolicyManual UpdatePolicy = "manual" // detect, then wait for approval
PolicyPinned UpdatePolicy = "pinned" // never auto-update
)
// Everything-is-a-plugin (tenet #3): each upstream is a watcher.
type UpdateWatcher interface {
Name() string // "docker-digest", "github-release", "registry-index"
// Return a candidate if a newer version exists for this instance.
Check(ctx context.Context, inst MCPInstance) (Candidate, bool, error)
}
type Candidate struct {
InstanceID string
From PackageVersion // current (digest-pinned)
To PackageVersion // proposed (digest-pinned)
Source string // watcher name
Notes string // e.g. GitHub release notes
}
type UpdateManager interface {
Scan(ctx context.Context) ([]Candidate, error) // periodic + event-driven
Approve(ctx context.Context, candidateID string) error
Apply(ctx context.Context, c Candidate) (MCPInstance, error) // re-install
Rollback(ctx context.Context, instanceID string) error // to known-good
SetPolicy(ctx context.Context, instanceID string, p UpdatePolicy) error
Pin(ctx context.Context, instanceID string, version string) error
}
```
HTTP/API surface (behind Auth + [RBAC](08-rbac.md)):
- `GET /api/v1/updates` — available/pending candidates across instances.
- `POST /api/v1/updates/{id}/approve` — approve a manual-gated candidate.
- `POST /api/v1/updates/{id}/apply` — force-apply now (audited).
- `POST /api/v1/instances/{id}/rollback` — roll back to previous known-good.
- `PUT /api/v1/instances/{id}/update-policy` — set policy / pin a version.
## Data
- **Reads:** `MCPInstance` (current `package`, `version`, digest) from
**Inventory**; `Package`/`PackageVersion` from the **Registry** store
(Architecture §8); instance `health` from Health Monitoring.
- **Writes:** on apply, the Installer updates the `MCPInstance` (new `version`,
`container_ref`); the Update Manager updates `state` around the transition
(`Updating`, `RollingBack`).
- **Module-local persistence:**
- `update_policy` per instance (`automatic|manual|pinned`, optional pinned
version/digest).
- `previous_good_spec` — the last healthy `MCPInstance` spec (package, version,
digest, rendered config-with-secret-refs) for rollback.
- `update_candidate` rows (from→to, source, status: available/approved/applied/
failed/rolled-back) and per-watcher cursors/ETags to avoid re-alerting.
## Dependencies
- [Auto Installer](03-auto-installer.md) — updates are re-installs of a new version.
- [MCP Package Registry](02-package-registry.md) — resolves/validates target digests.
- [Health Monitoring](09-health-monitoring.md) — gates the new version; signals rollback.
- [Notifications](13-notifications.md) — alerts on available/applied/failed/rolled-back.
- [Secrets Manager](07-secrets-manager.md) — re-injected at the re-install (unchanged).
- [Plugin System](11-plugin-system.md) — `UpdateWatcher` is a plugin point.
- [RBAC](08-rbac.md) — who may approve/apply/pin/rollback.
## Failure modes & handling
- **New version unhealthy:** the Installer's health gate fails, or Health
Monitoring reports the freshly updated instance unhealthy within a validation
window → automatic **rollback** to `previous_good_spec`; fire a failed-update
notification.
- **Watcher/upstream unreachable:** skip silently (best-effort), keep the last
cursor; repeated failures alert but never block other updates or the data plane.
- **Registry can't validate target digest:** candidate is rejected (not applied);
reported as `failed` with reason.
- **Pinned instance with newer upstream:** never applied; surfaced as
"update available (pinned)" only.
- **Rollback itself fails:** instance marked degraded, high-severity notification,
left for Health Monitoring quarantine + human intervention.
- **Concurrent apply + heal:** updates are serialized per `instance-id`
(shared lock with the Installer) so an update and a restart don't race.
- **Approval race / stale candidate:** applying a candidate whose `From` no longer
matches the live instance is rejected and re-scanned.
## Security notes
- Only **digest-pinned, signature-verified** target versions are applied; the
Registry re-verifies before the Installer pulls (Architecture §9).
- Secrets are re-injected at runtime by the Installer on the re-install — never
copied into `previous_good_spec` as plaintext; the retained spec holds secret
*references* only.
- Approve/apply/pin/rollback are audited mutating actions with principal, target,
from→to versions, and outcome.
- Auto-update policy is itself an RBAC-guarded setting; a compromised auto policy
is a supply-chain risk, so downgrades/policy changes are audited.
## Open questions
- Health **validation window** length after an update before declaring success —
fixed, or per-recipe?
- Should GitHub-release watching drive version *selection* directly, or only
notify the Registry to advance its index (single source of truth)?
- Batch/canary updates across many instances of the same package vs. one-at-a-time.
- Retention depth for `previous_good_spec` — just N−1, or a short history for
multi-step rollback?
- Maintenance windows / update scheduling to avoid disrupting active agent calls.
## Milestone
Phase 4 — *Operate (day-2)*. Update Manager ships with Docker-digest, GitHub-release,
and Registry-index watchers; automatic/manual/pinned policies; and health-gated
rollback. It contributes to the Phase 4 exit story alongside
[Health](09-health-monitoring.md), [Metrics](18-metrics.md),
[Notifications](13-notifications.md), and the [Dashboard](12-web-dashboard.md).