Two remaining pieces of the two-application coupling are gone. Overclock profiles were switched by naming 'comfy' and 'ollama' directly, so a third application could never get tuned clocks. A tenant declares overclock_profile and the arbitrator applies whichever the highest-priority *working* tenant asks for, falling back to the idle profile when nothing is running. The websocket listener parsed ComfyUI's message schema -- status, execution_start, executing, execution_success -- which tied the fast path to one application. An event source is now declarative and the messages are not parsed at all: any message means "look now", and the tenant's own busy probe decides what is true. That gives the same sub-second reaction to any application that emits anything on state change, with no knowledge of what it emits. Generalising this exposed a design error in the priority rule I had introduced. plan_release excluded candidates ranking above the demander, which broke both directions in turn. With the LLM at priority 60 and diffusion at 50, ComfyUI could never reclaim from Ollama -- the premise the whole service is built on, and preserved until now only by the ComfyUI-specific trigger that was about to be removed. Swapping the ranks then broke the reverse: a starved Ollama could no longer reclaim from an idle ComfyUI. Priority now orders rather than vetoes. Any idle reclaimable tenant is a candidate, because an idle tenant is not using its VRAM; priority decides who is asked first, and busy tenants are never interrupted whatever their rank. Diffusion outranks the LLM, whose weights reload from page cache in seconds. All three cases are pinned by tests, including that busy work is never interrupted even by a far higher-priority demander. Tests: 244 (was 242). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
HyperSwap // GPU Program Swapper & Memory Orchestrator
HyperSwap is an ultra-low-latency VRAM arbitrator, host RAM cache pre-warmer, dynamic hardware overclocker, and real-time telemetry dashboard designed specifically for Linux deployment machines that simultaneously host Ollama LLM workloads and ComfyUI Diffusion pipelines on a single NVIDIA GPU.
Real-Time Telemetry & Control Dashboard
The HyperSwap live dashboard running on :9090, demonstrating real-time VRAM allocation tracking (Ollama 14.14 GB, ComfyUI 0.38 GB, Desktop 0.6 GB), 33.89 GB of models resident in 64GB host RAM cache, live dual-axis memory charts, hardware fan control, and sub-25ms model VRAM purges and soft-yields.
1. Feature Matrix
⚡ Bidirectional VRAM Hot-Swapping & Arbitration
- A stale ComfyUI queue entry no longer disables arbitration. ComfyUI can leave a
dead job in
queue_runningindefinitely; one was found sitting there with the GPU idle and ComfyUI holding 0.56 GB. Trusting that flag made this service believe ComfyUI was permanently busy — so it evicted the LLM on every poll, never ran the idle purge, and never checked for CPU spill. Instrumenting the watchdog showedbusy=6, idle_check=0. A running entry is now corroborated against ComfyUI's own VRAM (a real job loads gigabytes; a dead one holds only its CUDA context) before it is believed, and a stale entry is reported by/api/health. Utilisation is deliberately not the signal — it is shared with Ollama and any third-party process. - Both directions are now automatic. Yielding Ollama for ComfyUI always was; the
reverse was not, despite "bidirectional" in this heading. Which way an LLM fails when
it cannot fit depends on configuration: with
n_gpu_layersleft to Ollama it spills layers to the CPU and reportssize_vram < size(roughly an order of magnitude slower, and silent). Withn_gpu_layerspinned — 99 on this box — it refuses outright withcudaMalloc failed: out of memory. Both are handled: the spill triggers a reclaim from an idle ComfyUI, and the hard failure is caught byswitch_ollama_model, which reclaims and retries once. Measured: a 12.87 GB model that returned HTTP 500 from Ollama directly now loads through HyperSwap after reclaiming 6.83 GB, at 3.85 GB/s. - A busy LLM is not a failed yield. A model mid-generation cannot unload; the
keep_alive: 0request queues behind it and applies when it finishes. That is reported asbusy(returning in ~610 ms) rather than blocking, with per-model backoff and a detached watcher that logs the eventual release. Only VRAM held while the GPU sits idle counts as a fault. - Confirmed Soft-Yield (barrier, not fire-and-forget): Releases Ollama VRAM allocations (
keep_alive: 0) down to 0 MB, then waits on NVML until the driver has actually freed the allocation before letting ComfyUI proceed. Postingkeep_alive: 0only asks Ollama to unload; on this box the HTTP call returns in ~63 ms while the driver takes a further ~77 ms to release 14.9 GB. Returning during that window is how diffusion ends up allocating into VRAM that is still occupied. - Idle-Aware ComfyUI Purge: Diffusion checkpoints are held for
COMFY_IDLE_PURGE_S(30 s) of genuinely empty queue rather than purged 1.5 s after every prompt — iterating on a workflow no longer pays a full checkpoint reload per run. An immediate purge still happens the moment Ollama actually asks for VRAM (POST /api/request-vram). - Real-Time ComfyUI WebSocket & Watchdog Listener: Subscribes directly to
ws://127.0.0.1:8188/ws. The WebSocket is the primary signal; a connection-pooled watchdog polls/queueat 1 Hz purely as a fallback, backing off to 3 s while the socket is healthy. - Process-Level VRAM Attribution: Live NVML process inspection attributes exact GPU memory usage across Ollama (
llama-server), ComfyUI (python), and Desktop display servers (gnome-shell,Xorg). - Bandwidth-Classified Transition History: Every switch is classified by the bandwidth it actually achieved (
model size ÷ load duration) rather than a fixed duration threshold:RAM Cache Hit ⚡(≥5 GB/s),Partial Cache 🌤(≥1.5 GB/s),Cold Disk Load 💾(below that). The previousload_duration < 2500 msrule called a 12.9 GB model read at 2.9 GB/s a "cold disk load" and a 0.5 GB model read from NVMe a "cache hit".
🧠 64GB Host RAM Cache & Page Pre-warmer
- Zero-Latency Model Discovery: Automatic cataloging of all local Ollama models (
/usr/share/ollama/.ollama/models,~/.ollama/models) and ComfyUI model directories (checkpoints,diffusion_models,unet,vae,clip,loras,controlnet). - POSIX
fadvise& Pinned Pre-warmer: Pre-faults multi-gigabyte GGUFs and Safetensors into the Linux OS Page Cache so that reloading models runs at page-cache speed rather than disk speed. - Measured Residency via
cachestat(2): Residency is measured, not assumed.cachestat(2)gives exact cached-page counts per file. Where the kernel refuses it — it only permits introspection of files you own, and Ollama's blobs are owned by uidollama— HyperSwap falls back to a randomised read-rate probe and labels the result as such. Files it cannot measure are reported as unmeasurable rather than guessed at. - Budgeted, Ranked Warming: This box has 64 GB of RAM and >270 GB of model files; reading everything simply evicts whatever was warmed first. Files are ranked by recency/frequency (from the persisted event log) and warmed until a byte budget is spent, skipping anything already resident.
GET /api/warm-planpreviews the decision without executing it. - Memory Telemetry: Real-time breakdown of Total Host RAM, Applications Memory, Active Model Page Cache, Free Memory, and measured Cache Residency Ratio.
🎛️ Measured Overclock Profiles & Thermal Management
Every profile setting in this repo is now backed by a measurement from autotune.py on
this specific card and driver. Several long-standing settings turned out to do nothing.
What this driver actually honours (NVIDIA 595.84, RTX 4080 SUPER):
| Lever | Mechanism | Works? |
|---|---|---|
| Power limit | nvidia-smi -pl |
✅ Yes — and it is the only lever that changes anything measurable |
| Core / memory clock lock | nvidia-smi -lgc / -lmc |
✅ Applies correctly, but made no measurable difference to either workload |
| Core / memory clock offsets | nvidia-settings -a ...Offset |
❌ Silently ignored. The driver reports assigned value 0 and the attribute still reads back 250. Detected automatically by offsets_supported(); apply_profile now skips them and says so rather than pretending. |
| Fan control | nvidia-settings GPUTargetFanSpeed |
✅ Yes |
Measured results (POST /api/autotune/sweep):
LLM decode is not power-bound. Throughput is flat across the card's entire power range — the GPU never drew more than 224 W no matter what the limit allowed:
| Power limit | 222 W | 259 W | 296 W | 320 W | 333 W | 370 W |
|---|---|---|---|---|---|---|
tok/s (qwen3.8long) |
73.17 | 73.13 | 73.43 | 73.51 | 73.10 | 73.04 |
Diffusion is power-bound. Here the watts genuinely buy throughput:
| Power limit | 222 W | 259 W | 296 W | 320 W | 333 W | 370 W |
|---|---|---|---|---|---|---|
| it/s (SDXL 1024, 20 steps) | 5.48 | 6.22 | 6.50 | 6.52 | 6.63 | 6.71 |
Clock locks changed nothing for either workload. Memory clock: 72.6 tok/s locked at 11251 MHz vs 72.7 unlocked. Core clock: 6.73 it/s unlocked vs 6.77 locked at 3105 MHz — and 6.78 at 2400 MHz, so diffusion here is not core-clock-bound at all.
Memory bandwidth is the decode bottleneck, confirming the profile's original premise — dropping the memory clock to 5001 MHz halves throughput (35.9 tok/s vs 72.6). The card simply reaches its top memory clock on its own; pinning it there adds nothing.
Resulting profiles:
ollama— 320 W (stock), no locks, automatic fans. Decode draws ~224 W and is bandwidth-bound, so the previous 370 W limit and 100% fan pinning bought nothing.comfy— 370 W, no locks, automatic fans. The extra power is worth a measured +2.8% over the 320 W stock default.balanced— stock power and boost, automatic fans.
On fans: all three profiles previously pinned the fans to manual 100%. Across 48,435
telemetry samples this card has never exceeded 81 °C and has logged zero thermal
throttle events; the ollama profile was holding 49.6 °C average by running the fans at
87%. Fans are now automatic in every profile, with the thermal governor escalating them
only if the card actually needs it.
🌡️ Thermal Governor (closed-loop de-escalation)
- Every overclock lever here is sticky: a profile locks clocks and pins the fans to a manual PWM, and nothing used to undo that. The governor watches the telemetry the sampler already collects (so it costs no extra NVML calls) and walks the overclock back through a four-step derate ladder when the card runs hot or reports a hardware throttle.
- Hysteresis by design: escalation needs 5 consecutive bad samples, recovery needs 30 consecutive good ones, with a 20 s cooldown between changes — a single spike during a diffusion step will not cause profile thrash.
- Guaranteed restore: stock clocks, default power limit and automatic fans are restored by the server's shutdown hook and by a systemd
ExecStopPost=, so aSIGKILLcannot leave the card with locked clocks and fans pinned at 100%.
🔬 Overclock Autotune (autotune.py)
- Sweeps a knob (
power_limit_w,lock_mem_mhz,lock_core_max, clock offsets) and reports the fastest stable value. - Both workloads are measurable.
workload=ollamabenchmarks decode throughput in tok/s;workload=comfyqueues a fixed SDXL 1024/20-step graph through ComfyUI's API and measures it/s. Without the second one there was no way to tell whether the compute-orientedcomfyprofile was doing anything at all — and it was not. - Refuses to sweep a knob the driver ignores. A preflight applies a probe value and confirms the hardware moved; the probe is chosen as the candidate furthest from the current reading, since probing with the maximum proves nothing when the card already sits there. This is what caught the silently-discarded clock offsets.
- Instability detection: kernel
Xid/NVRMmessages viajournalctl -k, benchmark failure, degenerate output, and a temperature ceiling. The sweep stops climbing the moment a step looks unstable. - Honest gain reporting: gain against the profile's current setting is reported separately from the spread across values tried. Conflating them turns a flat result into a headline "+102%".
- Safety: refuses to start while ComfyUI is executing, suspends the arbitrator's automatic profile switching for the duration (otherwise a diffusion benchmark trips
trigger_comfy_priority, which reapplies the whole profile and overwrites the clock being measured), and restores the original profile in afinallyblock — including on exception or cancellation.
🔧 Live Engine Configuration (engines.py)
GET /api/enginesreports the real, current configuration of both engines and what each setting implies for arbitration — because the settings that dictate this service's behaviour live outside its own codebase.OLLAMA_NUM_PARALLEL=1is why akeep_alive: 0unload queues behind a running generation and is reported as deferred rather than failed.OLLAMA_MAX_LOADED_MODELS=1is why every swap evicts the previous model. Working these out originally meant reading journald and the systemd unit by hand.- The dashboard's engine subtitles now come from this endpoint. They were previously hardcoded — and happened to be accurate, which is worse than being wrong, since they would have kept looking accurate after the configuration changed.
🩺 Dependency Self-Check (health.py)
GET /api/healthverifies everything this service depends on: NVML, passwordless sudo fornvidia-smi, fan control through the headless X server, overclock drift, the telemetry store, residency-measurement capability, model directories, the ComfyUI WebSocket, and both upstream HTTP services.- Each check reports what is broken, what that breaks, and how to fix it — not just a red light. Shown on the dashboard as a badge that expands only when something is wrong.
- It exists because fan control once failed for an entire session, recoverably and
silently: the unit started before the X server that owns the GPU was accepting
connections, the assignment failed with
Error resolving target specification 'gpu:0', nothing retried, and nothing ever asked whether fans worked. That failure now shows up in three places — a retry, a drift check, and this endpoint.
🧮 Honest VRAM Accounting
- Processes are bucketed ollama / comfy / desktop / unmanaged rather than into one
catch-all. On this machine a long-running
stt_relay.pyheld 842 MB for three days while the compositor held 3.9 MB; a single "system" number reported them as one figure. - That distinction matters because ComfyUI's memory can be reclaimed and a third
party's cannot.
unmanaged_gbis headroom the arbitrator can never give back, so it is reported explicitly, shown on the dashboard, and named in the error when a reclaim-and-retry still cannot fit a model.
🗄️ Persistent Telemetry Store (telemetry_store.py)
- Swap history used to be an in-memory
deque(maxlen=50)that evaporated on every restart. Telemetry and events now persist to SQLite (WAL, single writer thread, batched 1 Hz inserts, automatic retention pruning) at roughly 0.4 MB per hour. - This is what makes the app's central question answerable:
GET /api/analytics/profilescompares decode throughput per overclock profile, joined against the thermals recorded while that profile was active.
📊 Real-Time Web Telemetry Dashboard (:9090)
- Live Hardware Telemetry: GPU utilization %, GPU temperature (°C), power draw (W), fan speeds (%), and graphics/memory clock frequencies (MHz).
- Live Dual-Axis Time-Series Chart: Real-time graphical visualization of VRAM usage (GB) and Host RAM Cache (GB) with zero frontend polling overhead.
- Interactive Control Center: Trigger model hot-swaps, soft-yields, cache pre-warms, fan adjustments, and benchmarks directly from the web interface.
- Server-Sent Events (SSE): A single background sampler produces one 1 Hz snapshot and fans it out to every subscriber via
GET /api/stream. Frames are trimmed: the installed-model catalog was 81% of a 13.1 KB payload and changes only when a model is pulled, so it is sent on a subscriber's first frame and whenever it changes. Steady-state frames dropped 14041 → 3664 bytes (74% smaller; 135 → 38 MB/hour across three tabs), while/api/statsstill returns the complete snapshot. Previously each connected client independently re-ran the whole snapshot — NVML,/proc/meminfo, an HTTP round-trip each to Ollama and ComfyUI, and a recursive walk of the ComfyUI models tree with astat()per checkpoint — once per second, so opening the dashboard in three tabs tripled the load on the thing it was measuring. Slow clients drop stale frames instead of stalling the sampler.
🤖 Model Context Protocol (MCP 2.0) Server
- 23 Native Agentic Tools: Allows AI agents (Antigravity CLI, Claude Desktop, Cursor) to manage GPU resources, trigger model hot-swaps, measure page-cache residency, read persisted performance analytics, drive the thermal governor, and run overclock sweeps.
- 6 Live MCP Resources: Live metrics, model catalog, switch log, measured cache residency, per-profile analytics, and the overclock profiles with the evidence behind each setting.
- Dual Transport Support: Run via standard input/output (
--stdio) or network Server-Sent Events (--sse --port 8001).
⏱️ Automated Latency & Throughput Benchmark Engine
- Conducts automated round-trip model switching benchmarks to measure transition latency, model load time, tokens per second, and RAM cache effectiveness.
1b. Any Application, Not Just These Two
The purpose is fast handoff of one GPU between applications. It grew up around the two on this box, and their names ended up compiled into process matching, VRAM attribution, busy detection and release calls alike — about 385 references. That made it a script for Ollama and ComfyUI rather than a GPU arbitrator.
A tenant is now described as data in tenants.json:
{
"name": "trainer",
"kind": "other",
"priority": 80,
"match": { "cmdline": ["train.py"] },
"busy": { "type": "vram", "vram_busy_gb": 1.0 },
"release": { "type": "http_post", "url": "http://localhost:9999/release" }
}
| Field | What it answers |
|---|---|
match |
Which GPU processes belong to this application (name, cmdline substring, or suffix — ComfyUI is a bare python main.py) |
busy |
Whether it is genuinely working. http_count sums queue lists; vram needs no API at all. vram_floor_gb catches a queue that claims work while nothing is loaded |
release |
How to ask for VRAM back — http_post with a body, per_model for Ollama's per-model unload, or none |
priority |
Who is asked to yield first among idle tenants — it never protects idle memory, and never interrupts work |
overclock_profile |
GPU profile applied while this tenant is the active workload |
events |
Optional stream (e.g. a websocket) used purely as a wake-up, so reaction is sub-second rather than waiting for the next poll |
Two more fields drive the decision loop: needs_vram_gb (how much free memory the
application needs before it can work) and idle_release_after_s (how long it may sit
idle holding VRAM before being asked for it back — deliberately not immediate, so
iterating on a ComfyUI workflow does not reload the checkpoint between every run).
Priority orders, it does not veto. An idle tenant is not using its VRAM, so outranking the demander is no reason to keep it; busy tenants are never interrupted whatever their rank. Getting this wrong broke both directions in turn — with the LLM ranked above diffusion, ComfyUI could never preempt Ollama (the service's central behaviour), and once the ranks were swapped, a starved Ollama could no longer reclaim from an idle ComfyUI. Diffusion now outranks the LLM, whose weights reload from page cache in seconds.
plan_release() then arbitrates generically: a busy tenant that cannot reach
needs_vram_gb even counting what it already holds is starved, and the memory is taken
from idle reclaimable tenants below it in priority, lowest first, stopping as soon as
enough is freed. Tenants that cannot be released are named as blockers rather than
ignored, so possible: false comes with the reason. The plan is returned before it is
acted on, which makes the decision testable and loggable.
Ollama, ComfyUI and the desktop compositor ship as defaults, so behaviour is unchanged —
but nothing in the arbitration logic knows their names, and three applications can
contend for the card as easily as two. The dashboard's GPU Tenants panel lists all of
them ordered by the priority arbitration actually considers, with the last decision and
why it could or could not be satisfied. Endpoints are generic:
GET /api/tenants, GET /api/tenants/{name}, POST /api/tenants/{name}/release.
A tenant with "release": {"type": "none"} is still worth declaring. The 842 MB speech
relay on this box cannot be reclaimed, and naming it turns anonymous "unmanaged VRAM" into
"held by stt-relay, which exposes no release API" — and a release request returns 409
explaining that, rather than silently doing nothing.
1a. Tests
/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/ -q # 244 passed in ~3.8s
Hermetic: no GPU, no network, no sleeps. An autouse fixture stubs overclock_manager._sh
— the single choke point for every nvidia-smi/nvidia-settings write — so no test can
mutate the card, and HYPERSWAP_DB is redirected before telemetry_store imports.
The suite deliberately pins empirically measured constants, so that a future edit which contradicts the hardware fails loudly rather than silently:
| Pinned fact | Measured value | Why it is pinned |
|---|---|---|
| Warm model load | 12.87 GB in 4901 ms = 2.63 GB/s | The cache-hit threshold must stay below this, or no load can ever qualify |
| Cold model load | 12.87 GB in 34267 ms = 0.38 GB/s | Separates a genuine cold read from a partial hit |
| Busy yield | VRAM held at ≥50% GPU utilisation | A mid-generation model is finishing, not failing |
| Residency confidence | probe trusted only at 100% | A 12-window probe once cleared 90% on a mostly-cold file |
1b. End-to-End Verification
python verify_arbitration.py # full cycle, a few minutes
python verify_arbitration.py --quick # skip the diffusion stages
The unit suite covers logic in isolation. This exercises the promise the service exists to make — an LLM and a diffusion pipeline sharing one 16 GB card — against real hardware, and reports what actually happened at each stage. It restores what it changes and refuses to start if ComfyUI is busy.
A representative run on this machine:
| Stage | Result |
|---|---|
| LLM load, classified by achieved bandwidth | 1.96 GB in 1327 ms → 1.47 GB/s → Partial Cache |
| VRAM yield confirmed against NVML | released in 43 ms, 2.39 GB freed |
| Diffusion, cold (includes checkpoint load) | 17863 ms → 1.12 it/s |
| Diffusion, warm | 3645 ms → 5.49 it/s |
| ComfyUI retains its checkpoint | 7.03 GB held through the idle window |
| VRAM attribution adds up | 15.58 GB attributed vs 15.80 GB NVML — Ollama 7.71 + ComfyUI 7.03 coexisting |
| Reported GPU state matches hardware | profile asks 320 W, card reports 320 W |
The stages report warnings rather than passes when they did not actually prove anything — a reclaim that was never needed is not evidence that reclaiming works.
2. Architectural Overview
flowchart TD
subgraph HostRAM["64 GB Host System RAM (Page Cache & Staging Buffer)"]
OllamaGGUFs["Ollama GGUF Weights<br/>(Qwen, Gemma, Nemotron)"]
ComfySafetensors["ComfyUI Safetensors & VAEs<br/>(Wan2.1, Flux, SDXL)"]
end
subgraph GPU["NVIDIA GeForce RTX 4080 SUPER (16 GB VRAM)"]
direction LR
ActiveLLM["Active LLM<br/>(0–15 GB VRAM)"]
ActiveDiffusion["Active Diffusion Pipeline<br/>(0–15 GB VRAM)"]
end
subgraph Orchestrator["HyperSwap Control Plane (:9090)"]
REST["REST API & OpenAPI Docs"]
MCP["Model Context Protocol (MCP 2.0)"]
SSE["1Hz Real-Time SSE Stream"]
Arbitrator["VRAM Arbitrator (confirmed yield)"]
Overclock["Overclock & Fan Manager"]
Warmer["Page Cache Pre-Warmer"]
end
HostRAM <== "PCIe 4.0 x16 Bus (measured 2.6 GB/s warm model load)" ==> GPU
Orchestrator --> GPU
Orchestrator --> HostRAM
The Physics of Sub-Second Switching
-
Host RAM as Staging: Active LLMs and diffusion checkpoints remain resident in the 64GB Linux Page Cache.
-
Warm vs cold model loads, measured. The same 12.87 GB model, loaded through Ollama on this box:
Page-cache residency Load time Effective rate 3.1% (dropped with FADV_DONTNEED)34.3 s 0.38 GB/s 100% (force-warmed) 4.9 s 2.63 GB/s A 6.9× speedup, and the reason the page cache matters. Note the effective rate is well below the PCIe 4.0 x16 bus rate and below the 6.4 GB/s the page cache itself reads at: Ollama's
load_durationalso covers host-to-device transfer and model initialisation, not just the file read. Classification thresholds are calibrated against these measured numbers rather than the theoretical bus bandwidth — an earlier 5 GB/s cache-hit bar sat above what a fully warm load can even achieve, so every warm load was misreported as a partial hit. -
Soft-Yielding: Dropping Ollama's VRAM allocation via
keep_alive: 0preserves the weights in host RAM. Measured on this box: the HTTP request returns in ~63 ms, and the driver finishes releasing 14.9 GB ~77 ms after that. HyperSwap waits for the second number before handing VRAM to ComfyUI — the earlier "~15 ms" figure timed the request, not the release.
3. REST API Reference
The HyperSwap server runs on port 9090 by default. Interactive OpenAPI/Swagger docs are available at http://localhost:9090/docs.
Telemetry & Hardware Endpoints
| Endpoint | Method | Description |
|---|---|---|
/api/stats |
GET |
Complete unified JSON snapshot of hardware sensors, VRAM breakdown, host RAM, Ollama status, ComfyUI queue, and switch logs. |
/api/gpu |
GET |
NVIDIA GPU sensors (utilization %, temperature, power draw in Watts, fan speeds, clocks, and active PIDs). |
/api/memory |
GET |
Precise /proc/meminfo metrics (Total, Used, OS Page Cache containing models, Free memory). |
/api/gpu/fan |
GET |
Current GPU fan mode (auto vs manual), target speed %, and live fan RPM/PWM status. |
/api/gpu/fan |
POST |
Sets GPU fan speed mode (auto or manual) with target speed % (30–100%). |
/api/overclock |
GET |
Active overclock profile, configured profiles, GPU clock limits, and fan status. |
/api/overclock/apply |
POST |
Applies a named profile (ollama, comfy, balanced). |
/api/overclock/profile |
POST |
Creates or updates an overclock profile configuration. |
/api/stream |
GET |
Server-Sent Events (SSE) stream pushing full telemetry updates at 1Hz (text/event-stream). |
Model Orchestration & Hot-Swap Endpoints
| Endpoint | Method | Description |
|---|---|---|
/api/switch-model |
POST |
Hot-swaps the active Ollama LLM in VRAM and tracks transition timing. |
/api/free-vram |
POST |
Soft-yields Ollama VRAM to 0 MB and waits for NVML to confirm the release (?confirm=false to skip). Returns request_ms, confirm_ms and the GB actually freed. |
/api/comfy-free |
POST |
Instructs ComfyUI to purge loaded diffusion weights and VRAM cache. |
/api/request-vram |
POST |
Ollama-priority path: purges ComfyUI immediately if there is not enough free VRAM. |
/api/warm-all |
POST |
Warms the highest-value models into page cache within a byte budget (budget_gb). |
/api/warm-plan |
GET |
Previews what warming would read, in what order, and what it would skip — without doing it. |
/api/warm-model |
POST |
Pre-warms a specific model or file into RAM (blob_only warms weights without touching VRAM). |
/api/cache/report |
GET |
Measured page-cache residency per model file, with the measurement method used for each. |
/api/benchmark |
POST |
Runs an automated back-and-forth model swap benchmark and calculates average latency. |
Analytics Endpoints (persisted)
| Endpoint | Method | Description |
|---|---|---|
/api/analytics/profiles |
GET |
Decode throughput per overclock profile, joined with the thermals recorded under it. |
/api/analytics/swaps |
GET |
Aggregated swap/yield/purge latencies, cache-hit split, and per-model throughput. |
/api/analytics/timeseries |
GET |
Downsampled telemetry history for charts that outlive a page refresh. |
/api/analytics/models |
GET |
Recency/frequency model ranking used to prioritise the warm budget. |
/api/history?durable=true |
GET |
Swap history from the persistent store rather than the in-memory ring. |
/api/db |
GET |
Store location, row counts and how many hours of history are held. |
/api/engines |
GET |
Live Ollama and ComfyUI configuration, with what each setting implies for arbitration. |
/api/health |
GET |
Dependency self-check: NVML, sudo, fan control, drift, store, upstreams — each with impact and remediation. |
Governor & Autotune Endpoints
| Endpoint | Method | Description |
|---|---|---|
/api/governor |
GET / POST |
Current derate level and why; enable/disable, or clear an active derate. |
/api/overclock/restore |
POST |
Drop all clock locks and offsets, restore default power limit and automatic fans. |
/api/autotune |
GET |
Sweep progress, last result, and every recorded autotune step. |
/api/autotune/sweep |
POST |
Sweep a knob against a real workload (workload: ollama decode tok/s, comfy SDXL it/s), verifying the knob moves the hardware first. |
/api/autotune/cancel |
POST |
Stop the current sweep after the step in flight; the profile is restored either way. |
4. Model Context Protocol (MCP 2.0) Reference
HyperSwap includes a native MCP 2.0 server (mcp_server.py) exposing orchestration and telemetry tools to AI agents.
MCP Tools List
| Tool Name | Parameters | Description |
|---|---|---|
get_gpu_status |
None | Live NVIDIA GPU hardware telemetry, VRAM breakdown, temps, power, fan %, and PIDs. |
get_gpu_fan_status |
None | Current GPU fan mode (auto/manual) and target fan percentage. |
set_gpu_fan_speed |
mode (str), percent (optional int) |
Sets fan speed mode (auto|manual) and target PWM % (30–100%). |
get_host_memory_status |
None | 64GB host RAM breakdown, active page cache size, and cache ratio. |
switch_ollama_model |
model_name (str), keep_alive (str) |
Hot-swaps active LLM in VRAM, measures latency (ms) and tokens/sec. |
soft_yield_ollama_vram |
model_name (optional str) |
Yields Ollama VRAM to 0 MB and waits for NVML to confirm the driver actually released it. Returns the request/confirm split. |
purge_comfyui_vram |
None | Purges loaded diffusion models from ComfyUI pipeline VRAM. |
prewarm_all_models_to_ram |
None | Warms the highest-value models into page cache within a byte budget, skipping what is already resident. |
prewarm_single_model |
model_name (optional str), filepath (optional str) |
Pre-warms a single GGUF or Safetensors file into RAM. |
list_available_models |
None | Lists all installed Ollama models and discovered ComfyUI Safetensors on disk. |
get_switch_history |
limit (int, default 20) |
Retrieves recent switch events, millisecond latencies, and RAM hit status. |
run_model_switch_benchmark |
iterations (int, default 2) |
Automated round-trip latency benchmark between installed models. |
get_page_cache_residency |
include_files (bool) |
Measured page-cache residency per model file, with the measurement method used for each. |
get_warm_plan |
budget_gb (optional float) |
Previews what warming would read and skip, ranked by recency/frequency. Does not warm. |
request_vram_for_ollama |
needed_gb (float) |
Purges ComfyUI's checkpoints immediately if VRAM headroom is short, bypassing the idle timer. |
get_profile_performance |
days (float, default 7) |
Measured tok/s and thermals per overclock profile, from persisted history. |
get_thermal_governor_status |
None | Current derate level, the reason for it, and escalation history. |
set_thermal_governor |
enabled (optional bool), reset (bool) |
Enable/disable the governor, or clear an active derate. |
get_overclock_status |
None | Active profile, all profiles with their evidence, and which levers this driver honours. |
apply_overclock_profile |
profile (str) |
Apply ollama | comfy | balanced. |
restore_stock_gpu_state |
None | Drop clock locks and offsets, restore default power limit, return fans to automatic. |
run_overclock_sweep |
knob, profile, workload, start, stop, repeats, apply_best |
Sweep a knob against a real workload and report the fastest stable value. Verifies the knob moves the hardware first. Takes minutes. |
get_autotune_status |
None | Sweep progress, the last result table, and all recorded autotune steps. |
MCP Resources List
gpu://metrics/live: Real-time snapshot of GPU sensors and RAM page cache.gpu://models/catalog: Catalog of all discovered GGUF and Safetensors models.gpu://history/switches: Event log of recent model transitions and swap speeds.gpu://cache/residency: Measured page-cache residency across every model on disk.gpu://analytics/profiles: Measured throughput and thermals per overclock profile.gpu://overclock/profiles: Overclock profiles including the measurement behind each setting.
MCP Client Configurations
Antigravity Configuration (~/.gemini/antigravity-cli/mcp_config.json)
{
"mcpServers": {
"hyperswap": {
"command": "/home/drjones/comfy-mcp-venv/bin/python",
"args": ["/home/drjones/unified-model-manager/mcp_server.py", "--stdio"]
}
}
}
Claude Desktop Configuration (claude_desktop_config.json)
{
"mcpServers": {
"hyperswap": {
"command": "/home/drjones/comfy-mcp-venv/bin/python",
"args": ["/home/drjones/unified-model-manager/mcp_server.py", "--stdio"]
}
}
}
5. Linux Kernel & Host Tuning
To ensure that model weights remain permanently in RAM without kernel eviction:
# Set CPU scaling governor to performance
sudo cpupower frequency-set -g performance
# Configure sysctl optimizations in /etc/sysctl.d/99-hyperswap.conf
cat << 'EOF' | sudo tee /etc/sysctl.d/99-hyperswap.conf
# Retain model file cache aggressively in RAM
vm.vfs_cache_pressure = 50
# Prevent swapping cached models
vm.swappiness = 10
# Support large memory maps for high-parameter models
vm.max_map_count = 1048576
# Flush dirty pages quickly
vm.dirty_background_ratio = 5
vm.dirty_ratio = 10
EOF
# Apply sysctl settings immediately
sudo sysctl --system
6. Systemd Service Management
The HyperSwap server runs as a systemd service:
# Check service status
systemctl status hyperswap.service
# Restart service
sudo systemctl restart hyperswap.service
# View live telemetry and arbitration logs
journalctl -u hyperswap.service -f
7. License
MIT License. Developed for Google Antigravity & High-Throughput Linux AI Deployments.
