MCP Nexus is a self-hosted control plane for MCP servers — "Kubernetes for
MCP." Agents connect to one endpoint; Nexus discovers services, installs the
right MCP servers, aggregates them behind a namespaced router, secures access
with auth/RBAC, and heals/updates them via a reconcile loop.
This first commit is design-phase only (no runnable code yet):
- README.md project front door + module map
- docs/ARCHITECTURE.md target design: tenets, two-plane split, reconcile
loop, domain types, storage, security, deployment
- docs/ROADMAP.md phased delivery (foundations -> walking skeleton ->
discover+install -> secure -> operate -> extend/scale)
- docs/modules/01-20 one design doc per module, all cross-linked
Backend stack decision: Go (single static binary, embedded SQLite + embedded
React dashboard). Repo initialized in /root with a whitelist .gitignore so
only project files are tracked.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
8.6 KiB
Notifications
Module 13 · Plane: Cross-cutting · Roadmap phase: 4 Part of MCP Nexus architecture.
Purpose
Tell a human when something meaningful happens. Notifications turns internal events — a new service discovered, an update available or applied, a failed install, a service gone offline, a security alert — into messages delivered over pluggable notifiers: Discord, Slack, Email, Pushover, ntfy, Telegram, and Signal. It owns the notify stage of the reconciliation loop (Architecture §6): the reconciler and other modules only emit typed events; this module decides what is worth telling whom, over which channel, how often, and in what words.
Responsibilities
- Define a typed event model and subscribe to the event bus for notification-worthy transitions.
- Implement a Notifier plugin interface with built-in notifiers for Discord, Slack, Email (SMTP), Pushover, ntfy, Telegram, and Signal.
- Own the event→channel routing/subscription model: rules that map (event type, severity, source) → one or more configured channels.
- Assign and honor severity levels (
info,warning,error,critical) with per-channel minimum-severity filters. - Provide throttling, dedup, and grouping: collapse repeated/flapping events, rate-limit per channel, and batch bursts into digests.
- Template messages per event type with a per-channel renderer (rich embeds for Discord/Slack, plain text for ntfy/SMS-style, subject+body for email).
- Record delivery outcome (sent/failed/suppressed) to Audit and emit
nexus_notifications_sent_totalto Metrics.
Non-goals
- Does not detect conditions. It never polls health or updates itself — it reacts to events emitted by Health, Update Manager, Discovery, Installer, and Security.
- Not metric alerting. Threshold-on-series alerting is Prometheus Alertmanager against Metrics; this module is event-driven.
- Not a message queue / inbox. Fire-and-forward with bounded retry; it is not durable messaging or a ticketing system.
- Not the plugin transport. The out-of-process notifier transport is generalized by the Plugin System in Phase 5; here notifiers are in-process interfaces.
- Does not store secrets — channel credentials (webhook URLs, SMTP creds, bot tokens) come by ref from Secrets.
Interfaces
// Notifier is a pluggable delivery channel.
type Notifier interface {
Name() string // "discord", "slack", "email", "ntfy", "telegram", "pushover", "signal"
Configure(ctx context.Context, cfg ChannelConfig) error // creds via Secret refs
Send(ctx context.Context, msg Message) error
Health(ctx context.Context) error // reachability check for the dashboard
}
// Event is what the rest of Nexus emits onto the bus.
type Event struct {
Type EventType // ServiceDiscovered, UpdateAvailable, UpdateApplied,
// InstallFailed, ServiceOffline, SecurityAlert, ...
Severity Severity // Info | Warning | Error | Critical
Source string // module + subject, e.g. "health/postgres-01"
Subject string // short title
Fields map[string]any // structured detail for templating
Time time.Time
DedupKey string // events sharing a key collapse within a window
}
// Router decides which channels an event reaches and applies throttling.
type Router interface {
Register(n Notifier) error
Route(ctx context.Context, e Event) error // subscription match → throttle → render → Send
Test(ctx context.Context, channel string) error // "send test notification"
}
// Template renders an Event into a channel-specific Message.
type Template interface {
Render(e Event, channel string) (Message, error)
}
Internal HTTP (control API, RBAC-guarded):
GET/POST /api/v1/notifications/channels— list/configure channels (creds by Secret ref).POST /api/v1/notifications/channels/{name}/test— send a test message.GET/POST /api/v1/notifications/rules— manage event→channel subscription rules.GET /api/v1/notifications/history— recent delivery log (from Audit).
Templating
Each event type has a base template rendered per channel: rich embeds (title, color-by-severity, fields, action link back to the Dashboard) for Discord/Slack; subject + body for email; a compact single line for ntfy/Pushover/Telegram/Signal. Templates are overridable in config and rendered from the Event.Fields map, so a new event type ships a default template without code changes. Rendering is isolated from delivery — a template error degrades to a safe plain-text fallback rather than dropping the alert.
Event → severity defaults
| Event | Default severity |
|---|---|
ServiceDiscovered |
info |
UpdateAvailable |
info |
UpdateApplied |
info / warning (if rollback) |
InstallFailed |
error |
ServiceOffline / quarantined |
error / critical |
SecurityAlert |
critical |
Data
- Reads channel + rule config from the Config store; channel credentials by ref from Secrets (never persisted here, never logged).
- Writes delivery outcomes to the Audit store (§8) — append-only: event, channels, sent/failed/suppressed, timestamp.
- Module-local in-memory state: dedup windows, per-channel rate-limit token buckets, and pending-digest buffers (rebuilt on restart; at-most-once semantics on crash).
Dependencies
- Consumes events from Health, Update Manager, Discovery, Auto Installer, and Security.
- Secrets Manager — channel credentials by ref.
- Plugin System — the
Notifierinterface is one of its first-class plugin types; third-party notifiers land Phase 5. - Metrics —
nexus_notifications_sent_total{channel,severity,outcome}. - Web Dashboard — channel config, rule editing, and history UI.
Failure modes & handling
| Failure | Behavior |
|---|---|
| A channel is unreachable (webhook 5xx, SMTP down) | Bounded retry with backoff; then mark delivery failed in Audit; surface channel health in dashboard. Never block the emitter. |
| Notification storm (flapping instance) | Dedup by DedupKey within a window + per-channel rate limiting; bursts collapse into a digest. |
| A notifier plugin hangs | Send is deadline-bounded; a slow/hung channel is isolated and does not delay others (per-channel workers). |
| Misconfigured channel (bad token) | Configure/Test validates on save; runtime failures degrade to failed + dashboard warning, not a crash. |
| Nexus restarts mid-burst | At-most-once: pending in-memory digests may be lost; durable audit records what was sent. Critical events are sent eagerly, not batched. |
| Secret rotation | Channel re-reads cred by ref on next send; no restart required. |
Security notes
Honors Architecture §9. Channel credentials (webhook URLs, bot tokens, SMTP passwords) are Secrets referenced by ref, never stored in config sent to the dashboard and never logged. Message bodies are sanitized before send: no secret material, no raw agent payloads, and security-alert messages avoid leaking exploit detail to low-trust channels. Configuring/testing channels and editing rules are mutating actions, gated by RBAC and audited. SecurityAlert routing should prefer channels with delivery guarantees (email/Pushover) over best-effort chat webhooks.
Open questions
- Do we support acknowledgement / two-way interactions (e.g. approve an update from a Slack button), or keep notifications strictly one-way in v1?
- Digest cadence: fixed windows, or adaptive based on event rate?
- Per-user vs. per-system channels — should individual users subscribe personal channels, or are channels a system-wide concern in v1?
- Signal delivery requires a linked device/
signal-clisidecar — bundle guidance or treat as advanced/optional?
Milestone
Delivered in Phase 4 (Operate / day-2). Thin slice: the Notifier interface with Discord + Email built in, event subscriptions for ServiceDiscovered/InstallFailed/ServiceOffline/UpdateApplied/SecurityAlert, severity filtering, and basic dedup/throttle. Exit proof (shared with the phase goal): kill a managed MCP container → Health emits ServiceOffline → a Discord alert fires within one cycle.