# HyperSwap // GPU Program Swapper & Memory Orchestrator [![FastAPI](https://img.shields.io/badge/FastAPI-0.141-009688.svg?style=flat&logo=fastapi)](https://fastapi.tiangolo.com) [![Model Context Protocol](https://img.shields.io/badge/MCP-2.0-8A2BE2.svg?style=flat)](https://modelcontextprotocol.io) [![NVIDIA CUDA](https://img.shields.io/badge/CUDA-13.2%20%7C%2012.8-76B900.svg?style=flat&logo=nvidia)](https://developer.nvidia.com/cuda-zone) [![Platform](https://img.shields.io/badge/Platform-Linux%20x86__64-orange.svg?style=flat&logo=linux)](https://ubuntu.com) **HyperSwap** is an ultra-low-latency VRAM arbitrator, host RAM cache pre-warmer, dynamic hardware overclocker, and real-time telemetry dashboard designed specifically for Linux deployment machines that simultaneously host **Ollama LLM workloads** and **ComfyUI Diffusion pipelines** on a single NVIDIA GPU. --- ## Real-Time Telemetry & Control Dashboard ![HyperSwap Dashboard](assets/dashboard.png) *The HyperSwap live dashboard running on `:9090`, demonstrating real-time VRAM allocation tracking (Ollama 14.14 GB, ComfyUI 0.38 GB, Desktop 0.6 GB), 33.89 GB of models resident in 64GB host RAM cache, live dual-axis memory charts, hardware fan control, and sub-25ms model VRAM purges and soft-yields.* --- ## 1. Feature Matrix ### ⚡ Bidirectional VRAM Hot-Swapping & Arbitration * **Confirmed Soft-Yield (barrier, not fire-and-forget)**: Releases Ollama VRAM allocations (`keep_alive: 0`) down to 0 MB, then **waits on NVML until the driver has actually freed the allocation** before letting ComfyUI proceed. Posting `keep_alive: 0` only *asks* Ollama to unload; on this box the HTTP call returns in ~63 ms while the driver takes a further ~77 ms to release 14.9 GB. Returning during that window is how diffusion ends up allocating into VRAM that is still occupied. * **Idle-Aware ComfyUI Purge**: Diffusion checkpoints are held for `COMFY_IDLE_PURGE_S` (30 s) of genuinely empty queue rather than purged 1.5 s after every prompt — iterating on a workflow no longer pays a full checkpoint reload per run. An immediate purge still happens the moment Ollama actually asks for VRAM (`POST /api/request-vram`). * **Real-Time ComfyUI WebSocket & Watchdog Listener**: Subscribes directly to `ws://127.0.0.1:8188/ws`. The WebSocket is the primary signal; a connection-pooled watchdog polls `/queue` at 1 Hz purely as a fallback, backing off to 3 s while the socket is healthy. * **Process-Level VRAM Attribution**: Live NVML process inspection attributes exact GPU memory usage across Ollama (`llama-server`), ComfyUI (`python`), and Desktop display servers (`gnome-shell`, `Xorg`). * **Bandwidth-Classified Transition History**: Every switch is classified by the bandwidth it actually achieved (`model size ÷ load duration`) rather than a fixed duration threshold: `RAM Cache Hit ⚡` (≥5 GB/s), `Partial Cache 🌤` (≥1.5 GB/s), `Cold Disk Load 💾` (below that). The previous `load_duration < 2500 ms` rule called a 12.9 GB model read at 2.9 GB/s a "cold disk load" and a 0.5 GB model read from NVMe a "cache hit". ### 🧠 64GB Host RAM Cache & Page Pre-warmer * **Zero-Latency Model Discovery**: Automatic cataloging of all local Ollama models (`/usr/share/ollama/.ollama/models`, `~/.ollama/models`) and ComfyUI model directories (`checkpoints`, `diffusion_models`, `unet`, `vae`, `clip`, `loras`, `controlnet`). * **POSIX `fadvise` & Pinned Pre-warmer**: Pre-faults multi-gigabyte GGUFs and Safetensors into the Linux OS Page Cache so that reloading models runs at page-cache speed rather than disk speed. * **Measured Residency via `cachestat(2)`**: Residency is measured, not assumed. `cachestat(2)` gives exact cached-page counts per file. Where the kernel refuses it — it only permits introspection of files you own, and Ollama's blobs are owned by uid `ollama` — HyperSwap falls back to a randomised read-rate probe and labels the result as such. Files it cannot measure are reported as unmeasurable rather than guessed at. * **Budgeted, Ranked Warming**: This box has 64 GB of RAM and >270 GB of model files; reading everything simply evicts whatever was warmed first. Files are ranked by recency/frequency (from the persisted event log) and warmed until a byte budget is spent, skipping anything already resident. `GET /api/warm-plan` previews the decision without executing it. * **Memory Telemetry**: Real-time breakdown of Total Host RAM, Applications Memory, Active Model Page Cache, Free Memory, and measured Cache Residency Ratio. ### 🎛️ Dynamic Overclocking & Thermal Management * **Workload-Aware Overclock Profiles**: * **`ollama` Profile (Memory-Bandwidth Bound)**: Max 370W power limit, +150 MHz Core Offset, +825 MHz Memory Offset, and 100% fan speed for maximum prompt eval / generation bandwidth. * **`comfy` Profile (Compute Bound)**: Max 370W power limit, +100 MHz Core Offset, +500 MHz Memory Offset, Core Clock locked to 2900–3105 MHz, and 75% fan speed for maximum diffusion compute. * **`balanced` Profile (Stock/General Purpose)**: Unlocked 370W power limit with stock dynamic boost curves and automatic fan control. * **Hardware Actuation Hierarchy**: * Level 1: Power Limit Control (`nvidia-smi -pl 370`). * Level 2: Core & Memory Clock Locking (`nvidia-smi -lgc` / `-lmc`). * Level 3: Clock Offsets via headless X display (`:8`) with Coolbits support (`nvidia-settings`). * **Hardware Fan Control**: Switch between `auto` and `manual` PWM control (30%–100%) with synchronized dual-fan actuation (`[fan:0]` and `[fan:1]`). * **Automated Lockstep Profile Switching**: AutoArbitrator automatically switches hardware profiles in lockstep with the active workload (`comfy` on generation start, `ollama` on completion). ### 🌡️ Thermal Governor (closed-loop de-escalation) * Every overclock lever here is sticky: a profile locks clocks and pins the fans to a manual PWM, and nothing used to undo that. The governor watches the telemetry the sampler already collects (so it costs no extra NVML calls) and walks the overclock back through a four-step derate ladder when the card runs hot or reports a hardware throttle. * **Hysteresis by design**: escalation needs 5 consecutive bad samples, recovery needs 30 consecutive good ones, with a 20 s cooldown between changes — a single spike during a diffusion step will not cause profile thrash. * **Guaranteed restore**: stock clocks, default power limit and automatic fans are restored by the server's shutdown hook *and* by a systemd `ExecStopPost=`, so a `SIGKILL` cannot leave the card with locked clocks and fans pinned at 100%. ### 🔬 Overclock Autotune (`autotune.py`) * Walks a clock offset upward, running a fixed decode benchmark at each step, and reports the **fastest stable** value with its measured gain over baseline. * **Instability detection**: kernel `Xid`/`NVRM` messages via `journalctl -k`, benchmark failure, degenerate output, and a temperature ceiling. The sweep stops climbing the moment a step looks unstable. * **Safety**: refuses to start while ComfyUI is executing, and restores the original profile in a `finally` block — including on exception or cancellation. ### 🗄️ Persistent Telemetry Store (`telemetry_store.py`) * Swap history used to be an in-memory `deque(maxlen=50)` that evaporated on every restart. Telemetry and events now persist to SQLite (WAL, single writer thread, batched 1 Hz inserts, automatic retention pruning) at roughly **0.4 MB per hour**. * This is what makes the app's central question answerable: **`GET /api/analytics/profiles` compares decode throughput per overclock profile**, joined against the thermals recorded while that profile was active. ### 📊 Real-Time Web Telemetry Dashboard (`:9090`) * **Live Hardware Telemetry**: GPU utilization %, GPU temperature (°C), power draw (W), fan speeds (%), and graphics/memory clock frequencies (MHz). * **Live Dual-Axis Time-Series Chart**: Real-time graphical visualization of VRAM usage (GB) and Host RAM Cache (GB) with zero frontend polling overhead. * **Interactive Control Center**: Trigger model hot-swaps, soft-yields, cache pre-warms, fan adjustments, and benchmarks directly from the web interface. * **Server-Sent Events (SSE)**: A single background sampler produces one 1 Hz snapshot and fans it out to every subscriber via `GET /api/stream`. Previously each connected client independently re-ran the whole snapshot — NVML, `/proc/meminfo`, an HTTP round-trip each to Ollama and ComfyUI, and a recursive walk of the ComfyUI models tree with a `stat()` per checkpoint — once per second, so opening the dashboard in three tabs tripled the load on the thing it was measuring. Slow clients drop stale frames instead of stalling the sampler. ### 🤖 Model Context Protocol (MCP 2.0) Server * **12 Native Agentic Tools**: Allows AI agents (Antigravity CLI, Claude Desktop, Cursor) to manage GPU resources, trigger model hot-swaps, tune fan curves, and inspect telemetry. * **3 Live MCP Resources**: Exposes live metrics, model catalogs, and switch logs as streamable resources (`gpu://metrics/live`, `gpu://models/catalog`, `gpu://history/switches`). * **Dual Transport Support**: Run via standard input/output (`--stdio`) or network Server-Sent Events (`--sse --port 8001`). ### ⏱️ Automated Latency & Throughput Benchmark Engine * Conducts automated round-trip model switching benchmarks to measure transition latency, model load time, tokens per second, and RAM cache effectiveness. --- ## 2. Architectural Overview ```mermaid flowchart TD subgraph HostRAM["64 GB Host System RAM (Page Cache & Staging Buffer)"] OllamaGGUFs["Ollama GGUF Weights
(Qwen, Gemma, Nemotron)"] ComfySafetensors["ComfyUI Safetensors & VAEs
(Wan2.1, Flux, SDXL)"] end subgraph GPU["NVIDIA GeForce RTX 4080 SUPER (16 GB VRAM)"] direction LR ActiveLLM["Active LLM
(0–15 GB VRAM)"] ActiveDiffusion["Active Diffusion Pipeline
(0–15 GB VRAM)"] end subgraph Orchestrator["HyperSwap Control Plane (:9090)"] REST["REST API & OpenAPI Docs"] MCP["Model Context Protocol (MCP 2.0)"] SSE["1Hz Real-Time SSE Stream"] Arbitrator["VRAM Arbitrator (15ms Soft-Yield)"] Overclock["Overclock & Fan Manager"] Warmer["Page Cache Pre-Warmer"] end HostRAM <== "PCIe 4.0 x16 Bus (~31.5 GB/s Hot-Swap)" ==> GPU Orchestrator --> GPU Orchestrator --> HostRAM ``` ### The Physics of Sub-Second Switching * **Host RAM as Staging**: Active LLMs and diffusion checkpoints remain resident in the 64GB Linux Page Cache. * **PCIe 4.0 x16 Hot-Swapping**: Transferring weights across PCIe 4.0 x16 achieves **~31.5 GB/s** bandwidth, reducing model loads from 30+ seconds (disk) to **under 1.5 seconds**. * **Soft-Yielding**: Dropping Ollama's VRAM allocation via `keep_alive: 0` preserves the weights in host RAM. Measured on this box: the HTTP request returns in **~63 ms**, and the driver finishes releasing 14.9 GB **~77 ms after that**. HyperSwap waits for the second number before handing VRAM to ComfyUI — the earlier "~15 ms" figure timed the request, not the release. --- ## 3. REST API Reference The HyperSwap server runs on port `9090` by default. Interactive OpenAPI/Swagger docs are available at `http://localhost:9090/docs`. ### Telemetry & Hardware Endpoints | Endpoint | Method | Description | | :--- | :--- | :--- | | `/api/stats` | `GET` | Complete unified JSON snapshot of hardware sensors, VRAM breakdown, host RAM, Ollama status, ComfyUI queue, and switch logs. | | `/api/gpu` | `GET` | NVIDIA GPU sensors (utilization %, temperature, power draw in Watts, fan speeds, clocks, and active PIDs). | | `/api/memory` | `GET` | Precise `/proc/meminfo` metrics (Total, Used, OS Page Cache containing models, Free memory). | | `/api/gpu/fan` | `GET` | Current GPU fan mode (`auto` vs `manual`), target speed %, and live fan RPM/PWM status. | | `/api/gpu/fan` | `POST` | Sets GPU fan speed mode (`auto` or `manual`) with target speed % (30–100%). | | `/api/overclock` | `GET` | Active overclock profile, configured profiles, GPU clock limits, and fan status. | | `/api/overclock/apply` | `POST` | Applies a named profile (`ollama`, `comfy`, `balanced`). | | `/api/overclock/profile` | `POST` | Creates or updates an overclock profile configuration. | | `/api/stream` | `GET` | Server-Sent Events (SSE) stream pushing full telemetry updates at 1Hz (`text/event-stream`). | ### Model Orchestration & Hot-Swap Endpoints | Endpoint | Method | Description | | :--- | :--- | :--- | | `/api/switch-model` | `POST` | Hot-swaps the active Ollama LLM in VRAM and tracks transition timing. | | `/api/free-vram` | `POST` | Soft-yields Ollama VRAM to 0 MB and **waits for NVML to confirm the release** (`?confirm=false` to skip). Returns `request_ms`, `confirm_ms` and the GB actually freed. | | `/api/comfy-free` | `POST` | Instructs ComfyUI to purge loaded diffusion weights and VRAM cache. | | `/api/request-vram` | `POST` | Ollama-priority path: purges ComfyUI immediately if there is not enough free VRAM. | | `/api/warm-all` | `POST` | Warms the highest-value models into page cache within a byte budget (`budget_gb`). | | `/api/warm-plan` | `GET` | Previews what warming would read, in what order, and what it would skip — without doing it. | | `/api/warm-model` | `POST` | Pre-warms a specific model or file into RAM (`blob_only` warms weights without touching VRAM). | | `/api/cache/report` | `GET` | Measured page-cache residency per model file, with the measurement method used for each. | | `/api/benchmark` | `POST` | Runs an automated back-and-forth model swap benchmark and calculates average latency. | ### Analytics Endpoints (persisted) | Endpoint | Method | Description | | :--- | :--- | :--- | | `/api/analytics/profiles` | `GET` | **Decode throughput per overclock profile**, joined with the thermals recorded under it. | | `/api/analytics/swaps` | `GET` | Aggregated swap/yield/purge latencies, cache-hit split, and per-model throughput. | | `/api/analytics/timeseries` | `GET` | Downsampled telemetry history for charts that outlive a page refresh. | | `/api/analytics/models` | `GET` | Recency/frequency model ranking used to prioritise the warm budget. | | `/api/history?durable=true` | `GET` | Swap history from the persistent store rather than the in-memory ring. | | `/api/db` | `GET` | Store location, row counts and how many hours of history are held. | ### Governor & Autotune Endpoints | Endpoint | Method | Description | | :--- | :--- | :--- | | `/api/governor` | `GET` / `POST` | Current derate level and why; enable/disable, or clear an active derate. | | `/api/overclock/restore` | `POST` | Drop all clock locks and offsets, restore default power limit and automatic fans. | | `/api/autotune` | `GET` | Sweep progress, last result, and every recorded autotune step. | | `/api/autotune/sweep` | `POST` | Walk a clock offset upward, measuring tok/s and watching for instability at each step. | | `/api/autotune/cancel` | `POST` | Stop the current sweep after the step in flight; the profile is restored either way. | --- ## 4. Model Context Protocol (MCP 2.0) Reference HyperSwap includes a native **MCP 2.0 server** (`mcp_server.py`) exposing orchestration and telemetry tools to AI agents. ### MCP Tools List | Tool Name | Parameters | Description | | :--- | :--- | :--- | | **`get_gpu_status`** | *None* | Live NVIDIA GPU hardware telemetry, VRAM breakdown, temps, power, fan %, and PIDs. | | **`get_gpu_fan_status`** | *None* | Current GPU fan mode (`auto`/`manual`) and target fan percentage. | | **`set_gpu_fan_speed`** | `mode` (str), `percent` (optional int) | Sets fan speed mode (`auto`\|`manual`) and target PWM % (30–100%). | | **`get_host_memory_status`** | *None* | 64GB host RAM breakdown, active page cache size, and cache ratio. | | **`switch_ollama_model`** | `model_name` (str), `keep_alive` (str) | Hot-swaps active LLM in VRAM, measures latency (ms) and tokens/sec. | | **`soft_yield_ollama_vram`** | `model_name` (optional str) | Yields Ollama VRAM to 0 MB in ~15ms while keeping model weights in RAM cache. | | **`purge_comfyui_vram`** | *None* | Purges loaded diffusion models from ComfyUI pipeline VRAM. | | **`prewarm_all_models_to_ram`** | *None* | Faults all local LLM and diffusion checkpoints into Linux OS page cache. | | **`prewarm_single_model`** | `model_name` (optional str), `filepath` (optional str) | Pre-warms a single GGUF or Safetensors file into RAM. | | **`list_available_models`** | *None* | Lists all installed Ollama models and discovered ComfyUI Safetensors on disk. | | **`get_switch_history`** | `limit` (int, default 20) | Retrieves recent switch events, millisecond latencies, and RAM hit status. | | **`run_model_switch_benchmark`**| `iterations` (int, default 2) | Automated round-trip latency benchmark between installed models. | ### MCP Resources List * `gpu://metrics/live`: Real-time snapshot of GPU sensors and RAM page cache. * `gpu://models/catalog`: Catalog of all discovered GGUF and Safetensors models. * `gpu://history/switches`: Event log of recent model transitions and swap speeds. --- ### MCP Client Configurations #### Antigravity Configuration (`~/.gemini/antigravity-cli/mcp_config.json`) ```json { "mcpServers": { "hyperswap": { "command": "/home/drjones/comfy-mcp-venv/bin/python", "args": ["/home/drjones/unified-model-manager/mcp_server.py", "--stdio"] } } } ``` #### Claude Desktop Configuration (`claude_desktop_config.json`) ```json { "mcpServers": { "hyperswap": { "command": "/home/drjones/comfy-mcp-venv/bin/python", "args": ["/home/drjones/unified-model-manager/mcp_server.py", "--stdio"] } } } ``` --- ## 5. Linux Kernel & Host Tuning To ensure that model weights remain permanently in RAM without kernel eviction: ```bash # Set CPU scaling governor to performance sudo cpupower frequency-set -g performance # Configure sysctl optimizations in /etc/sysctl.d/99-hyperswap.conf cat << 'EOF' | sudo tee /etc/sysctl.d/99-hyperswap.conf # Retain model file cache aggressively in RAM vm.vfs_cache_pressure = 50 # Prevent swapping cached models vm.swappiness = 10 # Support large memory maps for high-parameter models vm.max_map_count = 1048576 # Flush dirty pages quickly vm.dirty_background_ratio = 5 vm.dirty_ratio = 10 EOF # Apply sysctl settings immediately sudo sysctl --system ``` --- ## 6. Systemd Service Management The HyperSwap server runs as a systemd service: ```bash # Check service status systemctl status hyperswap.service # Restart service sudo systemctl restart hyperswap.service # View live telemetry and arbitration logs journalctl -u hyperswap.service -f ``` --- ## 7. License MIT License. Developed for Google Antigravity & High-Throughput Linux AI Deployments.