Files
mcp-gateway-nexus/docs/modules/10-update-manager.md
drjones 8d3ffef920 docs: initial architecture and design for MCP Nexus
MCP Nexus is a self-hosted control plane for MCP servers — "Kubernetes for
MCP." Agents connect to one endpoint; Nexus discovers services, installs the
right MCP servers, aggregates them behind a namespaced router, secures access
with auth/RBAC, and heals/updates them via a reconcile loop.

This first commit is design-phase only (no runnable code yet):
- README.md            project front door + module map
- docs/ARCHITECTURE.md target design: tenets, two-plane split, reconcile
                       loop, domain types, storage, security, deployment
- docs/ROADMAP.md      phased delivery (foundations -> walking skeleton ->
                       discover+install -> secure -> operate -> extend/scale)
- docs/modules/01-20   one design doc per module, all cross-linked

Backend stack decision: Go (single static binary, embedded SQLite + embedded
React dashboard). Repo initialized in /root with a whitelist .gitignore so
only project files are tracked.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-07 04:38:27 +00:00

8.2 KiB
Raw Permalink Blame History

Update Manager

Module 10 · Plane: Control · Roadmap phase: 4 Part of MCP Nexus architecture.

Purpose

The Update Manager keeps managed MCPInstances current and safe. It watches upstream sources for newer versions of the packages an instance was installed from — Docker image tags/digests, GitHub releases, and Registry index updates — decides whether to act based on policy (automatic, manual-approval-gated, or pinned), and applies updates as re-installs through the Auto Installer. Crucially, it keeps the previous known-good spec so any update can be rolled back safely.

It is a day-2 control-plane reconciler (Architecture §5): a failure here degrades management currency, never the data plane. Per tenet #2 it proposes and applies state transitions through the loop, it does not run ad-hoc upgrade scripts.

Responsibilities

  • Watch upstream sources per instance and detect available updates:
    • Docker image tags/digests (a new digest behind a tracked tag).
    • GitHub releases for source=github packages.
    • Registry index updates (a package's recommended version advanced).
  • Evaluate each candidate against the instance's update policy: automatic, manual (approval-gated), or pinned (never auto-update).
  • Enforce version pinning: a pinned instance is skipped entirely (still reported as "update available").
  • Apply an approved update by driving a re-install at the new digest via the Installer, preserving the same bound_resource_uuid.
  • Retain the previous known-good MCPInstance spec and perform rollback when the new version fails to become/stay healthy.
  • Coordinate with Health Monitoring to gate and validate updates, and with Notifications to alert on available / applied / failed / rolled-back updates.
  • Record all decisions and transitions to Audit (Architecture §9).

Non-goals

  • Performing the actual pull/create/start — delegated to the Installer and Runtime.
  • Choosing/verifying package signatures or digests — owned by the Registry.
  • Restarting crashed-but-same-version instances — that is Health Monitoring (self-healing), not an update.
  • Deciding which package a service uses (Recipes).

Interfaces

type UpdatePolicy string

const (
    PolicyAuto   UpdatePolicy = "automatic" // apply without human action
    PolicyManual UpdatePolicy = "manual"     // detect, then wait for approval
    PolicyPinned UpdatePolicy = "pinned"      // never auto-update
)

// Everything-is-a-plugin (tenet #3): each upstream is a watcher.
type UpdateWatcher interface {
    Name() string // "docker-digest", "github-release", "registry-index"
    // Return a candidate if a newer version exists for this instance.
    Check(ctx context.Context, inst MCPInstance) (Candidate, bool, error)
}

type Candidate struct {
    InstanceID string
    From       PackageVersion // current (digest-pinned)
    To         PackageVersion // proposed (digest-pinned)
    Source     string         // watcher name
    Notes      string         // e.g. GitHub release notes
}

type UpdateManager interface {
    Scan(ctx context.Context) ([]Candidate, error)      // periodic + event-driven
    Approve(ctx context.Context, candidateID string) error
    Apply(ctx context.Context, c Candidate) (MCPInstance, error) // re-install
    Rollback(ctx context.Context, instanceID string) error      // to known-good
    SetPolicy(ctx context.Context, instanceID string, p UpdatePolicy) error
    Pin(ctx context.Context, instanceID string, version string) error
}

HTTP/API surface (behind Auth + RBAC):

  • GET /api/v1/updates — available/pending candidates across instances.
  • POST /api/v1/updates/{id}/approve — approve a manual-gated candidate.
  • POST /api/v1/updates/{id}/apply — force-apply now (audited).
  • POST /api/v1/instances/{id}/rollback — roll back to previous known-good.
  • PUT /api/v1/instances/{id}/update-policy — set policy / pin a version.

Data

  • Reads: MCPInstance (current package, version, digest) from Inventory; Package/PackageVersion from the Registry store (Architecture §8); instance health from Health Monitoring.
  • Writes: on apply, the Installer updates the MCPInstance (new version, container_ref); the Update Manager updates state around the transition (Updating, RollingBack).
  • Module-local persistence:
    • update_policy per instance (automatic|manual|pinned, optional pinned version/digest).
    • previous_good_spec — the last healthy MCPInstance spec (package, version, digest, rendered config-with-secret-refs) for rollback.
    • update_candidate rows (from→to, source, status: available/approved/applied/ failed/rolled-back) and per-watcher cursors/ETags to avoid re-alerting.

Dependencies

Failure modes & handling

  • New version unhealthy: the Installer's health gate fails, or Health Monitoring reports the freshly updated instance unhealthy within a validation window → automatic rollback to previous_good_spec; fire a failed-update notification.
  • Watcher/upstream unreachable: skip silently (best-effort), keep the last cursor; repeated failures alert but never block other updates or the data plane.
  • Registry can't validate target digest: candidate is rejected (not applied); reported as failed with reason.
  • Pinned instance with newer upstream: never applied; surfaced as "update available (pinned)" only.
  • Rollback itself fails: instance marked degraded, high-severity notification, left for Health Monitoring quarantine + human intervention.
  • Concurrent apply + heal: updates are serialized per instance-id (shared lock with the Installer) so an update and a restart don't race.
  • Approval race / stale candidate: applying a candidate whose From no longer matches the live instance is rejected and re-scanned.

Security notes

  • Only digest-pinned, signature-verified target versions are applied; the Registry re-verifies before the Installer pulls (Architecture §9).
  • Secrets are re-injected at runtime by the Installer on the re-install — never copied into previous_good_spec as plaintext; the retained spec holds secret references only.
  • Approve/apply/pin/rollback are audited mutating actions with principal, target, from→to versions, and outcome.
  • Auto-update policy is itself an RBAC-guarded setting; a compromised auto policy is a supply-chain risk, so downgrades/policy changes are audited.

Open questions

  • Health validation window length after an update before declaring success — fixed, or per-recipe?
  • Should GitHub-release watching drive version selection directly, or only notify the Registry to advance its index (single source of truth)?
  • Batch/canary updates across many instances of the same package vs. one-at-a-time.
  • Retention depth for previous_good_spec — just N−1, or a short history for multi-step rollback?
  • Maintenance windows / update scheduling to avoid disrupting active agent calls.

Milestone

Phase 4 — Operate (day-2). Update Manager ships with Docker-digest, GitHub-release, and Registry-index watchers; automatic/manual/pinned policies; and health-gated rollback. It contributes to the Phase 4 exit story alongside Health, Metrics, Notifications, and the Dashboard.