Nine changes, in rough order of how much they affect real behaviour: 1. VRAM yield is now a barrier. Posting keep_alive:0 only asks Ollama to unload; measured here, the HTTP call returns in 63ms while the driver takes a further 77ms to release 14.9GB. Returning inside that window is how ComfyUI ends up allocating into VRAM that is still occupied. instant_free_ollama_vram() polls NVML until the allocation is actually gone and reports request/confirm split. 2. ComfyUI VRAM is no longer purged 1.5s after every prompt, which forced a full checkpoint reload on each workflow iteration. It is held for 30s of genuinely empty queue, with an immediate purge when Ollama actually asks for the memory. 3. Cache-hit classification uses achieved bandwidth (size / load duration) rather than a fixed `load_duration < 2500ms`. That constant called a 12.9GB model read at 2.9GB/s a cold load, and a 0.5GB model read from NVMe a cache hit. 4. Page-cache residency is measured, not assumed. mincore(2) reported 128GB resident on a box with 46GB of page cache: the kernel only permits page-cache introspection on files you own, and the Ollama blobs are owned by uid ollama, for which mincore answers "all resident" instead of failing. Uses cachestat(2) where permitted and a randomised read-rate probe elsewhere, labelling which was used. Fixed-offset probing was self-fulfilling, so windows are random and cold ones are returned with FADV_DONTNEED. 5. Warming is budgeted and ranked by recency/frequency instead of reading every file top-to-bottom, which on 64GB of RAM just evicts whatever was warmed first. 6. Telemetry and events persist to SQLite (~0.38 MB/hour) instead of living in a 50-entry in-memory deque, so /api/analytics/profiles can finally answer whether an overclock profile actually delivers more tok/s. 7. Thermal governor walks the overclock back on sustained heat or hardware throttling, with hysteresis, fed from the existing sampler. 8. Autotune sweeps a clock offset, benchmarks decode at each step, watches for Xid errors and degenerate output, and restores the profile in a finally block. 9. Stock clocks/power/fans are restored on shutdown and via systemd ExecStopPost. Nothing previously undid a locked clock or a manually pinned fan. Also: one shared 1Hz telemetry sampler fanned out to SSE subscribers rather than every client re-running the whole snapshot; wall-clock timestamps in place of the event loop's monotonic clock; cached nvidia-smi shell-outs; quieter httpx logging. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
HyperSwap // GPU Program Swapper & Memory Orchestrator
HyperSwap is an ultra-low-latency VRAM arbitrator, host RAM cache pre-warmer, dynamic hardware overclocker, and real-time telemetry dashboard designed specifically for Linux deployment machines that simultaneously host Ollama LLM workloads and ComfyUI Diffusion pipelines on a single NVIDIA GPU.
Real-Time Telemetry & Control Dashboard
The HyperSwap live dashboard running on :9090, demonstrating real-time VRAM allocation tracking (Ollama 14.14 GB, ComfyUI 0.38 GB, Desktop 0.6 GB), 33.89 GB of models resident in 64GB host RAM cache, live dual-axis memory charts, hardware fan control, and sub-25ms model VRAM purges and soft-yields.
1. Feature Matrix
⚡ Bidirectional VRAM Hot-Swapping & Arbitration
- Confirmed Soft-Yield (barrier, not fire-and-forget): Releases Ollama VRAM allocations (
keep_alive: 0) down to 0 MB, then waits on NVML until the driver has actually freed the allocation before letting ComfyUI proceed. Postingkeep_alive: 0only asks Ollama to unload; on this box the HTTP call returns in ~63 ms while the driver takes a further ~77 ms to release 14.9 GB. Returning during that window is how diffusion ends up allocating into VRAM that is still occupied. - Idle-Aware ComfyUI Purge: Diffusion checkpoints are held for
COMFY_IDLE_PURGE_S(30 s) of genuinely empty queue rather than purged 1.5 s after every prompt — iterating on a workflow no longer pays a full checkpoint reload per run. An immediate purge still happens the moment Ollama actually asks for VRAM (POST /api/request-vram). - Real-Time ComfyUI WebSocket & Watchdog Listener: Subscribes directly to
ws://127.0.0.1:8188/ws. The WebSocket is the primary signal; a connection-pooled watchdog polls/queueat 1 Hz purely as a fallback, backing off to 3 s while the socket is healthy. - Process-Level VRAM Attribution: Live NVML process inspection attributes exact GPU memory usage across Ollama (
llama-server), ComfyUI (python), and Desktop display servers (gnome-shell,Xorg). - Bandwidth-Classified Transition History: Every switch is classified by the bandwidth it actually achieved (
model size ÷ load duration) rather than a fixed duration threshold:RAM Cache Hit ⚡(≥5 GB/s),Partial Cache 🌤(≥1.5 GB/s),Cold Disk Load 💾(below that). The previousload_duration < 2500 msrule called a 12.9 GB model read at 2.9 GB/s a "cold disk load" and a 0.5 GB model read from NVMe a "cache hit".
🧠 64GB Host RAM Cache & Page Pre-warmer
- Zero-Latency Model Discovery: Automatic cataloging of all local Ollama models (
/usr/share/ollama/.ollama/models,~/.ollama/models) and ComfyUI model directories (checkpoints,diffusion_models,unet,vae,clip,loras,controlnet). - POSIX
fadvise& Pinned Pre-warmer: Pre-faults multi-gigabyte GGUFs and Safetensors into the Linux OS Page Cache so that reloading models runs at page-cache speed rather than disk speed. - Measured Residency via
cachestat(2): Residency is measured, not assumed.cachestat(2)gives exact cached-page counts per file. Where the kernel refuses it — it only permits introspection of files you own, and Ollama's blobs are owned by uidollama— HyperSwap falls back to a randomised read-rate probe and labels the result as such. Files it cannot measure are reported as unmeasurable rather than guessed at. - Budgeted, Ranked Warming: This box has 64 GB of RAM and >270 GB of model files; reading everything simply evicts whatever was warmed first. Files are ranked by recency/frequency (from the persisted event log) and warmed until a byte budget is spent, skipping anything already resident.
GET /api/warm-planpreviews the decision without executing it. - Memory Telemetry: Real-time breakdown of Total Host RAM, Applications Memory, Active Model Page Cache, Free Memory, and measured Cache Residency Ratio.
🎛️ Dynamic Overclocking & Thermal Management
- Workload-Aware Overclock Profiles:
ollamaProfile (Memory-Bandwidth Bound): Max 370W power limit, +150 MHz Core Offset, +825 MHz Memory Offset, and 100% fan speed for maximum prompt eval / generation bandwidth.comfyProfile (Compute Bound): Max 370W power limit, +100 MHz Core Offset, +500 MHz Memory Offset, Core Clock locked to 2900–3105 MHz, and 75% fan speed for maximum diffusion compute.balancedProfile (Stock/General Purpose): Unlocked 370W power limit with stock dynamic boost curves and automatic fan control.
- Hardware Actuation Hierarchy:
- Level 1: Power Limit Control (
nvidia-smi -pl 370). - Level 2: Core & Memory Clock Locking (
nvidia-smi -lgc/-lmc). - Level 3: Clock Offsets via headless X display (
:8) with Coolbits support (nvidia-settings).
- Level 1: Power Limit Control (
- Hardware Fan Control: Switch between
autoandmanualPWM control (30%–100%) with synchronized dual-fan actuation ([fan:0]and[fan:1]). - Automated Lockstep Profile Switching: AutoArbitrator automatically switches hardware profiles in lockstep with the active workload (
comfyon generation start,ollamaon completion).
🌡️ Thermal Governor (closed-loop de-escalation)
- Every overclock lever here is sticky: a profile locks clocks and pins the fans to a manual PWM, and nothing used to undo that. The governor watches the telemetry the sampler already collects (so it costs no extra NVML calls) and walks the overclock back through a four-step derate ladder when the card runs hot or reports a hardware throttle.
- Hysteresis by design: escalation needs 5 consecutive bad samples, recovery needs 30 consecutive good ones, with a 20 s cooldown between changes — a single spike during a diffusion step will not cause profile thrash.
- Guaranteed restore: stock clocks, default power limit and automatic fans are restored by the server's shutdown hook and by a systemd
ExecStopPost=, so aSIGKILLcannot leave the card with locked clocks and fans pinned at 100%.
🔬 Overclock Autotune (autotune.py)
- Walks a clock offset upward, running a fixed decode benchmark at each step, and reports the fastest stable value with its measured gain over baseline.
- Instability detection: kernel
Xid/NVRMmessages viajournalctl -k, benchmark failure, degenerate output, and a temperature ceiling. The sweep stops climbing the moment a step looks unstable. - Safety: refuses to start while ComfyUI is executing, and restores the original profile in a
finallyblock — including on exception or cancellation.
🗄️ Persistent Telemetry Store (telemetry_store.py)
- Swap history used to be an in-memory
deque(maxlen=50)that evaporated on every restart. Telemetry and events now persist to SQLite (WAL, single writer thread, batched 1 Hz inserts, automatic retention pruning) at roughly 0.4 MB per hour. - This is what makes the app's central question answerable:
GET /api/analytics/profilescompares decode throughput per overclock profile, joined against the thermals recorded while that profile was active.
📊 Real-Time Web Telemetry Dashboard (:9090)
- Live Hardware Telemetry: GPU utilization %, GPU temperature (°C), power draw (W), fan speeds (%), and graphics/memory clock frequencies (MHz).
- Live Dual-Axis Time-Series Chart: Real-time graphical visualization of VRAM usage (GB) and Host RAM Cache (GB) with zero frontend polling overhead.
- Interactive Control Center: Trigger model hot-swaps, soft-yields, cache pre-warms, fan adjustments, and benchmarks directly from the web interface.
- Server-Sent Events (SSE): A single background sampler produces one 1 Hz snapshot and fans it out to every subscriber via
GET /api/stream. Previously each connected client independently re-ran the whole snapshot — NVML,/proc/meminfo, an HTTP round-trip each to Ollama and ComfyUI, and a recursive walk of the ComfyUI models tree with astat()per checkpoint — once per second, so opening the dashboard in three tabs tripled the load on the thing it was measuring. Slow clients drop stale frames instead of stalling the sampler.
🤖 Model Context Protocol (MCP 2.0) Server
- 12 Native Agentic Tools: Allows AI agents (Antigravity CLI, Claude Desktop, Cursor) to manage GPU resources, trigger model hot-swaps, tune fan curves, and inspect telemetry.
- 3 Live MCP Resources: Exposes live metrics, model catalogs, and switch logs as streamable resources (
gpu://metrics/live,gpu://models/catalog,gpu://history/switches). - Dual Transport Support: Run via standard input/output (
--stdio) or network Server-Sent Events (--sse --port 8001).
⏱️ Automated Latency & Throughput Benchmark Engine
- Conducts automated round-trip model switching benchmarks to measure transition latency, model load time, tokens per second, and RAM cache effectiveness.
2. Architectural Overview
flowchart TD
subgraph HostRAM["64 GB Host System RAM (Page Cache & Staging Buffer)"]
OllamaGGUFs["Ollama GGUF Weights<br/>(Qwen, Gemma, Nemotron)"]
ComfySafetensors["ComfyUI Safetensors & VAEs<br/>(Wan2.1, Flux, SDXL)"]
end
subgraph GPU["NVIDIA GeForce RTX 4080 SUPER (16 GB VRAM)"]
direction LR
ActiveLLM["Active LLM<br/>(0–15 GB VRAM)"]
ActiveDiffusion["Active Diffusion Pipeline<br/>(0–15 GB VRAM)"]
end
subgraph Orchestrator["HyperSwap Control Plane (:9090)"]
REST["REST API & OpenAPI Docs"]
MCP["Model Context Protocol (MCP 2.0)"]
SSE["1Hz Real-Time SSE Stream"]
Arbitrator["VRAM Arbitrator (15ms Soft-Yield)"]
Overclock["Overclock & Fan Manager"]
Warmer["Page Cache Pre-Warmer"]
end
HostRAM <== "PCIe 4.0 x16 Bus (~31.5 GB/s Hot-Swap)" ==> GPU
Orchestrator --> GPU
Orchestrator --> HostRAM
The Physics of Sub-Second Switching
- Host RAM as Staging: Active LLMs and diffusion checkpoints remain resident in the 64GB Linux Page Cache.
- PCIe 4.0 x16 Hot-Swapping: Transferring weights across PCIe 4.0 x16 achieves ~31.5 GB/s bandwidth, reducing model loads from 30+ seconds (disk) to under 1.5 seconds.
- Soft-Yielding: Dropping Ollama's VRAM allocation via
keep_alive: 0preserves the weights in host RAM. Measured on this box: the HTTP request returns in ~63 ms, and the driver finishes releasing 14.9 GB ~77 ms after that. HyperSwap waits for the second number before handing VRAM to ComfyUI — the earlier "~15 ms" figure timed the request, not the release.
3. REST API Reference
The HyperSwap server runs on port 9090 by default. Interactive OpenAPI/Swagger docs are available at http://localhost:9090/docs.
Telemetry & Hardware Endpoints
| Endpoint | Method | Description |
|---|---|---|
/api/stats |
GET |
Complete unified JSON snapshot of hardware sensors, VRAM breakdown, host RAM, Ollama status, ComfyUI queue, and switch logs. |
/api/gpu |
GET |
NVIDIA GPU sensors (utilization %, temperature, power draw in Watts, fan speeds, clocks, and active PIDs). |
/api/memory |
GET |
Precise /proc/meminfo metrics (Total, Used, OS Page Cache containing models, Free memory). |
/api/gpu/fan |
GET |
Current GPU fan mode (auto vs manual), target speed %, and live fan RPM/PWM status. |
/api/gpu/fan |
POST |
Sets GPU fan speed mode (auto or manual) with target speed % (30–100%). |
/api/overclock |
GET |
Active overclock profile, configured profiles, GPU clock limits, and fan status. |
/api/overclock/apply |
POST |
Applies a named profile (ollama, comfy, balanced). |
/api/overclock/profile |
POST |
Creates or updates an overclock profile configuration. |
/api/stream |
GET |
Server-Sent Events (SSE) stream pushing full telemetry updates at 1Hz (text/event-stream). |
Model Orchestration & Hot-Swap Endpoints
| Endpoint | Method | Description |
|---|---|---|
/api/switch-model |
POST |
Hot-swaps the active Ollama LLM in VRAM and tracks transition timing. |
/api/free-vram |
POST |
Soft-yields Ollama VRAM to 0 MB and waits for NVML to confirm the release (?confirm=false to skip). Returns request_ms, confirm_ms and the GB actually freed. |
/api/comfy-free |
POST |
Instructs ComfyUI to purge loaded diffusion weights and VRAM cache. |
/api/request-vram |
POST |
Ollama-priority path: purges ComfyUI immediately if there is not enough free VRAM. |
/api/warm-all |
POST |
Warms the highest-value models into page cache within a byte budget (budget_gb). |
/api/warm-plan |
GET |
Previews what warming would read, in what order, and what it would skip — without doing it. |
/api/warm-model |
POST |
Pre-warms a specific model or file into RAM (blob_only warms weights without touching VRAM). |
/api/cache/report |
GET |
Measured page-cache residency per model file, with the measurement method used for each. |
/api/benchmark |
POST |
Runs an automated back-and-forth model swap benchmark and calculates average latency. |
Analytics Endpoints (persisted)
| Endpoint | Method | Description |
|---|---|---|
/api/analytics/profiles |
GET |
Decode throughput per overclock profile, joined with the thermals recorded under it. |
/api/analytics/swaps |
GET |
Aggregated swap/yield/purge latencies, cache-hit split, and per-model throughput. |
/api/analytics/timeseries |
GET |
Downsampled telemetry history for charts that outlive a page refresh. |
/api/analytics/models |
GET |
Recency/frequency model ranking used to prioritise the warm budget. |
/api/history?durable=true |
GET |
Swap history from the persistent store rather than the in-memory ring. |
/api/db |
GET |
Store location, row counts and how many hours of history are held. |
Governor & Autotune Endpoints
| Endpoint | Method | Description |
|---|---|---|
/api/governor |
GET / POST |
Current derate level and why; enable/disable, or clear an active derate. |
/api/overclock/restore |
POST |
Drop all clock locks and offsets, restore default power limit and automatic fans. |
/api/autotune |
GET |
Sweep progress, last result, and every recorded autotune step. |
/api/autotune/sweep |
POST |
Walk a clock offset upward, measuring tok/s and watching for instability at each step. |
/api/autotune/cancel |
POST |
Stop the current sweep after the step in flight; the profile is restored either way. |
4. Model Context Protocol (MCP 2.0) Reference
HyperSwap includes a native MCP 2.0 server (mcp_server.py) exposing orchestration and telemetry tools to AI agents.
MCP Tools List
| Tool Name | Parameters | Description |
|---|---|---|
get_gpu_status |
None | Live NVIDIA GPU hardware telemetry, VRAM breakdown, temps, power, fan %, and PIDs. |
get_gpu_fan_status |
None | Current GPU fan mode (auto/manual) and target fan percentage. |
set_gpu_fan_speed |
mode (str), percent (optional int) |
Sets fan speed mode (auto|manual) and target PWM % (30–100%). |
get_host_memory_status |
None | 64GB host RAM breakdown, active page cache size, and cache ratio. |
switch_ollama_model |
model_name (str), keep_alive (str) |
Hot-swaps active LLM in VRAM, measures latency (ms) and tokens/sec. |
soft_yield_ollama_vram |
model_name (optional str) |
Yields Ollama VRAM to 0 MB in ~15ms while keeping model weights in RAM cache. |
purge_comfyui_vram |
None | Purges loaded diffusion models from ComfyUI pipeline VRAM. |
prewarm_all_models_to_ram |
None | Faults all local LLM and diffusion checkpoints into Linux OS page cache. |
prewarm_single_model |
model_name (optional str), filepath (optional str) |
Pre-warms a single GGUF or Safetensors file into RAM. |
list_available_models |
None | Lists all installed Ollama models and discovered ComfyUI Safetensors on disk. |
get_switch_history |
limit (int, default 20) |
Retrieves recent switch events, millisecond latencies, and RAM hit status. |
run_model_switch_benchmark |
iterations (int, default 2) |
Automated round-trip latency benchmark between installed models. |
MCP Resources List
gpu://metrics/live: Real-time snapshot of GPU sensors and RAM page cache.gpu://models/catalog: Catalog of all discovered GGUF and Safetensors models.gpu://history/switches: Event log of recent model transitions and swap speeds.
MCP Client Configurations
Antigravity Configuration (~/.gemini/antigravity-cli/mcp_config.json)
{
"mcpServers": {
"hyperswap": {
"command": "/home/drjones/comfy-mcp-venv/bin/python",
"args": ["/home/drjones/unified-model-manager/mcp_server.py", "--stdio"]
}
}
}
Claude Desktop Configuration (claude_desktop_config.json)
{
"mcpServers": {
"hyperswap": {
"command": "/home/drjones/comfy-mcp-venv/bin/python",
"args": ["/home/drjones/unified-model-manager/mcp_server.py", "--stdio"]
}
}
}
5. Linux Kernel & Host Tuning
To ensure that model weights remain permanently in RAM without kernel eviction:
# Set CPU scaling governor to performance
sudo cpupower frequency-set -g performance
# Configure sysctl optimizations in /etc/sysctl.d/99-hyperswap.conf
cat << 'EOF' | sudo tee /etc/sysctl.d/99-hyperswap.conf
# Retain model file cache aggressively in RAM
vm.vfs_cache_pressure = 50
# Prevent swapping cached models
vm.swappiness = 10
# Support large memory maps for high-parameter models
vm.max_map_count = 1048576
# Flush dirty pages quickly
vm.dirty_background_ratio = 5
vm.dirty_ratio = 10
EOF
# Apply sysctl settings immediately
sudo sysctl --system
6. Systemd Service Management
The HyperSwap server runs as a systemd service:
# Check service status
systemctl status hyperswap.service
# Restart service
sudo systemctl restart hyperswap.service
# View live telemetry and arbitration logs
journalctl -u hyperswap.service -f
7. License
MIT License. Developed for Google Antigravity & High-Throughput Linux AI Deployments.
