Nine changes, in rough order of how much they affect real behaviour: 1. VRAM yield is now a barrier. Posting keep_alive:0 only asks Ollama to unload; measured here, the HTTP call returns in 63ms while the driver takes a further 77ms to release 14.9GB. Returning inside that window is how ComfyUI ends up allocating into VRAM that is still occupied. instant_free_ollama_vram() polls NVML until the allocation is actually gone and reports request/confirm split. 2. ComfyUI VRAM is no longer purged 1.5s after every prompt, which forced a full checkpoint reload on each workflow iteration. It is held for 30s of genuinely empty queue, with an immediate purge when Ollama actually asks for the memory. 3. Cache-hit classification uses achieved bandwidth (size / load duration) rather than a fixed `load_duration < 2500ms`. That constant called a 12.9GB model read at 2.9GB/s a cold load, and a 0.5GB model read from NVMe a cache hit. 4. Page-cache residency is measured, not assumed. mincore(2) reported 128GB resident on a box with 46GB of page cache: the kernel only permits page-cache introspection on files you own, and the Ollama blobs are owned by uid ollama, for which mincore answers "all resident" instead of failing. Uses cachestat(2) where permitted and a randomised read-rate probe elsewhere, labelling which was used. Fixed-offset probing was self-fulfilling, so windows are random and cold ones are returned with FADV_DONTNEED. 5. Warming is budgeted and ranked by recency/frequency instead of reading every file top-to-bottom, which on 64GB of RAM just evicts whatever was warmed first. 6. Telemetry and events persist to SQLite (~0.38 MB/hour) instead of living in a 50-entry in-memory deque, so /api/analytics/profiles can finally answer whether an overclock profile actually delivers more tok/s. 7. Thermal governor walks the overclock back on sustained heat or hardware throttling, with hysteresis, fed from the existing sampler. 8. Autotune sweeps a clock offset, benchmarks decode at each step, watches for Xid errors and degenerate output, and restores the profile in a finally block. 9. Stock clocks/power/fans are restored on shutdown and via systemd ExecStopPost. Nothing previously undid a locked clock or a manually pinned fan. Also: one shared 1Hz telemetry sampler fanned out to SSE subscribers rather than every client re-running the whole snapshot; wall-clock timestamps in place of the event loop's monotonic clock; cached nvidia-smi shell-outs; quieter httpx logging. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
276 lines
18 KiB
Markdown
276 lines
18 KiB
Markdown
# HyperSwap // GPU Program Swapper & Memory Orchestrator
|
||
|
||
[](https://fastapi.tiangolo.com)
|
||
[](https://modelcontextprotocol.io)
|
||
[](https://developer.nvidia.com/cuda-zone)
|
||
[](https://ubuntu.com)
|
||
|
||
**HyperSwap** is an ultra-low-latency VRAM arbitrator, host RAM cache pre-warmer, dynamic hardware overclocker, and real-time telemetry dashboard designed specifically for Linux deployment machines that simultaneously host **Ollama LLM workloads** and **ComfyUI Diffusion pipelines** on a single NVIDIA GPU.
|
||
|
||
---
|
||
|
||
## Real-Time Telemetry & Control Dashboard
|
||
|
||

|
||
|
||
*The HyperSwap live dashboard running on `:9090`, demonstrating real-time VRAM allocation tracking (Ollama 14.14 GB, ComfyUI 0.38 GB, Desktop 0.6 GB), 33.89 GB of models resident in 64GB host RAM cache, live dual-axis memory charts, hardware fan control, and sub-25ms model VRAM purges and soft-yields.*
|
||
|
||
---
|
||
|
||
## 1. Feature Matrix
|
||
|
||
### ⚡ Bidirectional VRAM Hot-Swapping & Arbitration
|
||
* **Confirmed Soft-Yield (barrier, not fire-and-forget)**: Releases Ollama VRAM allocations (`keep_alive: 0`) down to 0 MB, then **waits on NVML until the driver has actually freed the allocation** before letting ComfyUI proceed. Posting `keep_alive: 0` only *asks* Ollama to unload; on this box the HTTP call returns in ~63 ms while the driver takes a further ~77 ms to release 14.9 GB. Returning during that window is how diffusion ends up allocating into VRAM that is still occupied.
|
||
* **Idle-Aware ComfyUI Purge**: Diffusion checkpoints are held for `COMFY_IDLE_PURGE_S` (30 s) of genuinely empty queue rather than purged 1.5 s after every prompt — iterating on a workflow no longer pays a full checkpoint reload per run. An immediate purge still happens the moment Ollama actually asks for VRAM (`POST /api/request-vram`).
|
||
* **Real-Time ComfyUI WebSocket & Watchdog Listener**: Subscribes directly to `ws://127.0.0.1:8188/ws`. The WebSocket is the primary signal; a connection-pooled watchdog polls `/queue` at 1 Hz purely as a fallback, backing off to 3 s while the socket is healthy.
|
||
* **Process-Level VRAM Attribution**: Live NVML process inspection attributes exact GPU memory usage across Ollama (`llama-server`), ComfyUI (`python`), and Desktop display servers (`gnome-shell`, `Xorg`).
|
||
* **Bandwidth-Classified Transition History**: Every switch is classified by the bandwidth it actually achieved (`model size ÷ load duration`) rather than a fixed duration threshold: `RAM Cache Hit ⚡` (≥5 GB/s), `Partial Cache 🌤` (≥1.5 GB/s), `Cold Disk Load 💾` (below that). The previous `load_duration < 2500 ms` rule called a 12.9 GB model read at 2.9 GB/s a "cold disk load" and a 0.5 GB model read from NVMe a "cache hit".
|
||
|
||
### 🧠 64GB Host RAM Cache & Page Pre-warmer
|
||
* **Zero-Latency Model Discovery**: Automatic cataloging of all local Ollama models (`/usr/share/ollama/.ollama/models`, `~/.ollama/models`) and ComfyUI model directories (`checkpoints`, `diffusion_models`, `unet`, `vae`, `clip`, `loras`, `controlnet`).
|
||
* **POSIX `fadvise` & Pinned Pre-warmer**: Pre-faults multi-gigabyte GGUFs and Safetensors into the Linux OS Page Cache so that reloading models runs at page-cache speed rather than disk speed.
|
||
* **Measured Residency via `cachestat(2)`**: Residency is measured, not assumed. `cachestat(2)` gives exact cached-page counts per file. Where the kernel refuses it — it only permits introspection of files you own, and Ollama's blobs are owned by uid `ollama` — HyperSwap falls back to a randomised read-rate probe and labels the result as such. Files it cannot measure are reported as unmeasurable rather than guessed at.
|
||
* **Budgeted, Ranked Warming**: This box has 64 GB of RAM and >270 GB of model files; reading everything simply evicts whatever was warmed first. Files are ranked by recency/frequency (from the persisted event log) and warmed until a byte budget is spent, skipping anything already resident. `GET /api/warm-plan` previews the decision without executing it.
|
||
* **Memory Telemetry**: Real-time breakdown of Total Host RAM, Applications Memory, Active Model Page Cache, Free Memory, and measured Cache Residency Ratio.
|
||
|
||
### 🎛️ Dynamic Overclocking & Thermal Management
|
||
* **Workload-Aware Overclock Profiles**:
|
||
* **`ollama` Profile (Memory-Bandwidth Bound)**: Max 370W power limit, +150 MHz Core Offset, +825 MHz Memory Offset, and 100% fan speed for maximum prompt eval / generation bandwidth.
|
||
* **`comfy` Profile (Compute Bound)**: Max 370W power limit, +100 MHz Core Offset, +500 MHz Memory Offset, Core Clock locked to 2900–3105 MHz, and 75% fan speed for maximum diffusion compute.
|
||
* **`balanced` Profile (Stock/General Purpose)**: Unlocked 370W power limit with stock dynamic boost curves and automatic fan control.
|
||
* **Hardware Actuation Hierarchy**:
|
||
* Level 1: Power Limit Control (`nvidia-smi -pl 370`).
|
||
* Level 2: Core & Memory Clock Locking (`nvidia-smi -lgc` / `-lmc`).
|
||
* Level 3: Clock Offsets via headless X display (`:8`) with Coolbits support (`nvidia-settings`).
|
||
* **Hardware Fan Control**: Switch between `auto` and `manual` PWM control (30%–100%) with synchronized dual-fan actuation (`[fan:0]` and `[fan:1]`).
|
||
* **Automated Lockstep Profile Switching**: AutoArbitrator automatically switches hardware profiles in lockstep with the active workload (`comfy` on generation start, `ollama` on completion).
|
||
|
||
### 🌡️ Thermal Governor (closed-loop de-escalation)
|
||
* Every overclock lever here is sticky: a profile locks clocks and pins the fans to a manual PWM, and nothing used to undo that. The governor watches the telemetry the sampler already collects (so it costs no extra NVML calls) and walks the overclock back through a four-step derate ladder when the card runs hot or reports a hardware throttle.
|
||
* **Hysteresis by design**: escalation needs 5 consecutive bad samples, recovery needs 30 consecutive good ones, with a 20 s cooldown between changes — a single spike during a diffusion step will not cause profile thrash.
|
||
* **Guaranteed restore**: stock clocks, default power limit and automatic fans are restored by the server's shutdown hook *and* by a systemd `ExecStopPost=`, so a `SIGKILL` cannot leave the card with locked clocks and fans pinned at 100%.
|
||
|
||
### 🔬 Overclock Autotune (`autotune.py`)
|
||
* Walks a clock offset upward, running a fixed decode benchmark at each step, and reports the **fastest stable** value with its measured gain over baseline.
|
||
* **Instability detection**: kernel `Xid`/`NVRM` messages via `journalctl -k`, benchmark failure, degenerate output, and a temperature ceiling. The sweep stops climbing the moment a step looks unstable.
|
||
* **Safety**: refuses to start while ComfyUI is executing, and restores the original profile in a `finally` block — including on exception or cancellation.
|
||
|
||
### 🗄️ Persistent Telemetry Store (`telemetry_store.py`)
|
||
* Swap history used to be an in-memory `deque(maxlen=50)` that evaporated on every restart. Telemetry and events now persist to SQLite (WAL, single writer thread, batched 1 Hz inserts, automatic retention pruning) at roughly **0.4 MB per hour**.
|
||
* This is what makes the app's central question answerable: **`GET /api/analytics/profiles` compares decode throughput per overclock profile**, joined against the thermals recorded while that profile was active.
|
||
|
||
### 📊 Real-Time Web Telemetry Dashboard (`:9090`)
|
||
* **Live Hardware Telemetry**: GPU utilization %, GPU temperature (°C), power draw (W), fan speeds (%), and graphics/memory clock frequencies (MHz).
|
||
* **Live Dual-Axis Time-Series Chart**: Real-time graphical visualization of VRAM usage (GB) and Host RAM Cache (GB) with zero frontend polling overhead.
|
||
* **Interactive Control Center**: Trigger model hot-swaps, soft-yields, cache pre-warms, fan adjustments, and benchmarks directly from the web interface.
|
||
* **Server-Sent Events (SSE)**: A single background sampler produces one 1 Hz snapshot and fans it out to every subscriber via `GET /api/stream`. Previously each connected client independently re-ran the whole snapshot — NVML, `/proc/meminfo`, an HTTP round-trip each to Ollama and ComfyUI, and a recursive walk of the ComfyUI models tree with a `stat()` per checkpoint — once per second, so opening the dashboard in three tabs tripled the load on the thing it was measuring. Slow clients drop stale frames instead of stalling the sampler.
|
||
|
||
### 🤖 Model Context Protocol (MCP 2.0) Server
|
||
* **12 Native Agentic Tools**: Allows AI agents (Antigravity CLI, Claude Desktop, Cursor) to manage GPU resources, trigger model hot-swaps, tune fan curves, and inspect telemetry.
|
||
* **3 Live MCP Resources**: Exposes live metrics, model catalogs, and switch logs as streamable resources (`gpu://metrics/live`, `gpu://models/catalog`, `gpu://history/switches`).
|
||
* **Dual Transport Support**: Run via standard input/output (`--stdio`) or network Server-Sent Events (`--sse --port 8001`).
|
||
|
||
### ⏱️ Automated Latency & Throughput Benchmark Engine
|
||
* Conducts automated round-trip model switching benchmarks to measure transition latency, model load time, tokens per second, and RAM cache effectiveness.
|
||
|
||
---
|
||
|
||
## 2. Architectural Overview
|
||
|
||
```mermaid
|
||
flowchart TD
|
||
subgraph HostRAM["64 GB Host System RAM (Page Cache & Staging Buffer)"]
|
||
OllamaGGUFs["Ollama GGUF Weights<br/>(Qwen, Gemma, Nemotron)"]
|
||
ComfySafetensors["ComfyUI Safetensors & VAEs<br/>(Wan2.1, Flux, SDXL)"]
|
||
end
|
||
|
||
subgraph GPU["NVIDIA GeForce RTX 4080 SUPER (16 GB VRAM)"]
|
||
direction LR
|
||
ActiveLLM["Active LLM<br/>(0–15 GB VRAM)"]
|
||
ActiveDiffusion["Active Diffusion Pipeline<br/>(0–15 GB VRAM)"]
|
||
end
|
||
|
||
subgraph Orchestrator["HyperSwap Control Plane (:9090)"]
|
||
REST["REST API & OpenAPI Docs"]
|
||
MCP["Model Context Protocol (MCP 2.0)"]
|
||
SSE["1Hz Real-Time SSE Stream"]
|
||
Arbitrator["VRAM Arbitrator (15ms Soft-Yield)"]
|
||
Overclock["Overclock & Fan Manager"]
|
||
Warmer["Page Cache Pre-Warmer"]
|
||
end
|
||
|
||
HostRAM <== "PCIe 4.0 x16 Bus (~31.5 GB/s Hot-Swap)" ==> GPU
|
||
Orchestrator --> GPU
|
||
Orchestrator --> HostRAM
|
||
```
|
||
|
||
### The Physics of Sub-Second Switching
|
||
* **Host RAM as Staging**: Active LLMs and diffusion checkpoints remain resident in the 64GB Linux Page Cache.
|
||
* **PCIe 4.0 x16 Hot-Swapping**: Transferring weights across PCIe 4.0 x16 achieves **~31.5 GB/s** bandwidth, reducing model loads from 30+ seconds (disk) to **under 1.5 seconds**.
|
||
* **Soft-Yielding**: Dropping Ollama's VRAM allocation via `keep_alive: 0` preserves the weights in host RAM. Measured on this box: the HTTP request returns in **~63 ms**, and the driver finishes releasing 14.9 GB **~77 ms after that**. HyperSwap waits for the second number before handing VRAM to ComfyUI — the earlier "~15 ms" figure timed the request, not the release.
|
||
|
||
---
|
||
|
||
## 3. REST API Reference
|
||
|
||
The HyperSwap server runs on port `9090` by default. Interactive OpenAPI/Swagger docs are available at `http://localhost:9090/docs`.
|
||
|
||
### Telemetry & Hardware Endpoints
|
||
|
||
| Endpoint | Method | Description |
|
||
| :--- | :--- | :--- |
|
||
| `/api/stats` | `GET` | Complete unified JSON snapshot of hardware sensors, VRAM breakdown, host RAM, Ollama status, ComfyUI queue, and switch logs. |
|
||
| `/api/gpu` | `GET` | NVIDIA GPU sensors (utilization %, temperature, power draw in Watts, fan speeds, clocks, and active PIDs). |
|
||
| `/api/memory` | `GET` | Precise `/proc/meminfo` metrics (Total, Used, OS Page Cache containing models, Free memory). |
|
||
| `/api/gpu/fan` | `GET` | Current GPU fan mode (`auto` vs `manual`), target speed %, and live fan RPM/PWM status. |
|
||
| `/api/gpu/fan` | `POST` | Sets GPU fan speed mode (`auto` or `manual`) with target speed % (30–100%). |
|
||
| `/api/overclock` | `GET` | Active overclock profile, configured profiles, GPU clock limits, and fan status. |
|
||
| `/api/overclock/apply` | `POST` | Applies a named profile (`ollama`, `comfy`, `balanced`). |
|
||
| `/api/overclock/profile` | `POST` | Creates or updates an overclock profile configuration. |
|
||
| `/api/stream` | `GET` | Server-Sent Events (SSE) stream pushing full telemetry updates at 1Hz (`text/event-stream`). |
|
||
|
||
### Model Orchestration & Hot-Swap Endpoints
|
||
|
||
| Endpoint | Method | Description |
|
||
| :--- | :--- | :--- |
|
||
| `/api/switch-model` | `POST` | Hot-swaps the active Ollama LLM in VRAM and tracks transition timing. |
|
||
| `/api/free-vram` | `POST` | Soft-yields Ollama VRAM to 0 MB and **waits for NVML to confirm the release** (`?confirm=false` to skip). Returns `request_ms`, `confirm_ms` and the GB actually freed. |
|
||
| `/api/comfy-free` | `POST` | Instructs ComfyUI to purge loaded diffusion weights and VRAM cache. |
|
||
| `/api/request-vram` | `POST` | Ollama-priority path: purges ComfyUI immediately if there is not enough free VRAM. |
|
||
| `/api/warm-all` | `POST` | Warms the highest-value models into page cache within a byte budget (`budget_gb`). |
|
||
| `/api/warm-plan` | `GET` | Previews what warming would read, in what order, and what it would skip — without doing it. |
|
||
| `/api/warm-model` | `POST` | Pre-warms a specific model or file into RAM (`blob_only` warms weights without touching VRAM). |
|
||
| `/api/cache/report` | `GET` | Measured page-cache residency per model file, with the measurement method used for each. |
|
||
| `/api/benchmark` | `POST` | Runs an automated back-and-forth model swap benchmark and calculates average latency. |
|
||
|
||
### Analytics Endpoints (persisted)
|
||
|
||
| Endpoint | Method | Description |
|
||
| :--- | :--- | :--- |
|
||
| `/api/analytics/profiles` | `GET` | **Decode throughput per overclock profile**, joined with the thermals recorded under it. |
|
||
| `/api/analytics/swaps` | `GET` | Aggregated swap/yield/purge latencies, cache-hit split, and per-model throughput. |
|
||
| `/api/analytics/timeseries` | `GET` | Downsampled telemetry history for charts that outlive a page refresh. |
|
||
| `/api/analytics/models` | `GET` | Recency/frequency model ranking used to prioritise the warm budget. |
|
||
| `/api/history?durable=true` | `GET` | Swap history from the persistent store rather than the in-memory ring. |
|
||
| `/api/db` | `GET` | Store location, row counts and how many hours of history are held. |
|
||
|
||
### Governor & Autotune Endpoints
|
||
|
||
| Endpoint | Method | Description |
|
||
| :--- | :--- | :--- |
|
||
| `/api/governor` | `GET` / `POST` | Current derate level and why; enable/disable, or clear an active derate. |
|
||
| `/api/overclock/restore` | `POST` | Drop all clock locks and offsets, restore default power limit and automatic fans. |
|
||
| `/api/autotune` | `GET` | Sweep progress, last result, and every recorded autotune step. |
|
||
| `/api/autotune/sweep` | `POST` | Walk a clock offset upward, measuring tok/s and watching for instability at each step. |
|
||
| `/api/autotune/cancel` | `POST` | Stop the current sweep after the step in flight; the profile is restored either way. |
|
||
|
||
---
|
||
|
||
## 4. Model Context Protocol (MCP 2.0) Reference
|
||
|
||
HyperSwap includes a native **MCP 2.0 server** (`mcp_server.py`) exposing orchestration and telemetry tools to AI agents.
|
||
|
||
### MCP Tools List
|
||
|
||
| Tool Name | Parameters | Description |
|
||
| :--- | :--- | :--- |
|
||
| **`get_gpu_status`** | *None* | Live NVIDIA GPU hardware telemetry, VRAM breakdown, temps, power, fan %, and PIDs. |
|
||
| **`get_gpu_fan_status`** | *None* | Current GPU fan mode (`auto`/`manual`) and target fan percentage. |
|
||
| **`set_gpu_fan_speed`** | `mode` (str), `percent` (optional int) | Sets fan speed mode (`auto`\|`manual`) and target PWM % (30–100%). |
|
||
| **`get_host_memory_status`** | *None* | 64GB host RAM breakdown, active page cache size, and cache ratio. |
|
||
| **`switch_ollama_model`** | `model_name` (str), `keep_alive` (str) | Hot-swaps active LLM in VRAM, measures latency (ms) and tokens/sec. |
|
||
| **`soft_yield_ollama_vram`** | `model_name` (optional str) | Yields Ollama VRAM to 0 MB in ~15ms while keeping model weights in RAM cache. |
|
||
| **`purge_comfyui_vram`** | *None* | Purges loaded diffusion models from ComfyUI pipeline VRAM. |
|
||
| **`prewarm_all_models_to_ram`** | *None* | Faults all local LLM and diffusion checkpoints into Linux OS page cache. |
|
||
| **`prewarm_single_model`** | `model_name` (optional str), `filepath` (optional str) | Pre-warms a single GGUF or Safetensors file into RAM. |
|
||
| **`list_available_models`** | *None* | Lists all installed Ollama models and discovered ComfyUI Safetensors on disk. |
|
||
| **`get_switch_history`** | `limit` (int, default 20) | Retrieves recent switch events, millisecond latencies, and RAM hit status. |
|
||
| **`run_model_switch_benchmark`**| `iterations` (int, default 2) | Automated round-trip latency benchmark between installed models. |
|
||
|
||
### MCP Resources List
|
||
|
||
* `gpu://metrics/live`: Real-time snapshot of GPU sensors and RAM page cache.
|
||
* `gpu://models/catalog`: Catalog of all discovered GGUF and Safetensors models.
|
||
* `gpu://history/switches`: Event log of recent model transitions and swap speeds.
|
||
|
||
---
|
||
|
||
### MCP Client Configurations
|
||
|
||
#### Antigravity Configuration (`~/.gemini/antigravity-cli/mcp_config.json`)
|
||
```json
|
||
{
|
||
"mcpServers": {
|
||
"hyperswap": {
|
||
"command": "/home/drjones/comfy-mcp-venv/bin/python",
|
||
"args": ["/home/drjones/unified-model-manager/mcp_server.py", "--stdio"]
|
||
}
|
||
}
|
||
}
|
||
```
|
||
|
||
#### Claude Desktop Configuration (`claude_desktop_config.json`)
|
||
```json
|
||
{
|
||
"mcpServers": {
|
||
"hyperswap": {
|
||
"command": "/home/drjones/comfy-mcp-venv/bin/python",
|
||
"args": ["/home/drjones/unified-model-manager/mcp_server.py", "--stdio"]
|
||
}
|
||
}
|
||
}
|
||
```
|
||
|
||
---
|
||
|
||
## 5. Linux Kernel & Host Tuning
|
||
|
||
To ensure that model weights remain permanently in RAM without kernel eviction:
|
||
|
||
```bash
|
||
# Set CPU scaling governor to performance
|
||
sudo cpupower frequency-set -g performance
|
||
|
||
# Configure sysctl optimizations in /etc/sysctl.d/99-hyperswap.conf
|
||
cat << 'EOF' | sudo tee /etc/sysctl.d/99-hyperswap.conf
|
||
# Retain model file cache aggressively in RAM
|
||
vm.vfs_cache_pressure = 50
|
||
|
||
# Prevent swapping cached models
|
||
vm.swappiness = 10
|
||
|
||
# Support large memory maps for high-parameter models
|
||
vm.max_map_count = 1048576
|
||
|
||
# Flush dirty pages quickly
|
||
vm.dirty_background_ratio = 5
|
||
vm.dirty_ratio = 10
|
||
EOF
|
||
|
||
# Apply sysctl settings immediately
|
||
sudo sysctl --system
|
||
```
|
||
|
||
---
|
||
|
||
## 6. Systemd Service Management
|
||
|
||
The HyperSwap server runs as a systemd service:
|
||
|
||
```bash
|
||
# Check service status
|
||
systemctl status hyperswap.service
|
||
|
||
# Restart service
|
||
sudo systemctl restart hyperswap.service
|
||
|
||
# View live telemetry and arbitration logs
|
||
journalctl -u hyperswap.service -f
|
||
```
|
||
|
||
---
|
||
|
||
## 7. License
|
||
|
||
MIT License. Developed for Google Antigravity & High-Throughput Linux AI Deployments.
|