drjones 5431144b2e Add barrier-confirmed yielding, measured residency, persistence and closed-loop tuning
Nine changes, in rough order of how much they affect real behaviour:

1. VRAM yield is now a barrier. Posting keep_alive:0 only asks Ollama to unload;
   measured here, the HTTP call returns in 63ms while the driver takes a further
   77ms to release 14.9GB. Returning inside that window is how ComfyUI ends up
   allocating into VRAM that is still occupied. instant_free_ollama_vram() polls
   NVML until the allocation is actually gone and reports request/confirm split.

2. ComfyUI VRAM is no longer purged 1.5s after every prompt, which forced a full
   checkpoint reload on each workflow iteration. It is held for 30s of genuinely
   empty queue, with an immediate purge when Ollama actually asks for the memory.

3. Cache-hit classification uses achieved bandwidth (size / load duration) rather
   than a fixed `load_duration < 2500ms`. That constant called a 12.9GB model read
   at 2.9GB/s a cold load, and a 0.5GB model read from NVMe a cache hit.

4. Page-cache residency is measured, not assumed. mincore(2) reported 128GB
   resident on a box with 46GB of page cache: the kernel only permits page-cache
   introspection on files you own, and the Ollama blobs are owned by uid ollama,
   for which mincore answers "all resident" instead of failing. Uses cachestat(2)
   where permitted and a randomised read-rate probe elsewhere, labelling which was
   used. Fixed-offset probing was self-fulfilling, so windows are random and cold
   ones are returned with FADV_DONTNEED.

5. Warming is budgeted and ranked by recency/frequency instead of reading every
   file top-to-bottom, which on 64GB of RAM just evicts whatever was warmed first.

6. Telemetry and events persist to SQLite (~0.38 MB/hour) instead of living in a
   50-entry in-memory deque, so /api/analytics/profiles can finally answer whether
   an overclock profile actually delivers more tok/s.

7. Thermal governor walks the overclock back on sustained heat or hardware
   throttling, with hysteresis, fed from the existing sampler.

8. Autotune sweeps a clock offset, benchmarks decode at each step, watches for Xid
   errors and degenerate output, and restores the profile in a finally block.

9. Stock clocks/power/fans are restored on shutdown and via systemd ExecStopPost.
   Nothing previously undid a locked clock or a manually pinned fan.

Also: one shared 1Hz telemetry sampler fanned out to SSE subscribers rather than
every client re-running the whole snapshot; wall-clock timestamps in place of the
event loop's monotonic clock; cached nvidia-smi shell-outs; quieter httpx logging.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-28 08:57:35 -07:00

HyperSwap // GPU Program Swapper & Memory Orchestrator

FastAPI Model Context Protocol NVIDIA CUDA Platform

HyperSwap is an ultra-low-latency VRAM arbitrator, host RAM cache pre-warmer, dynamic hardware overclocker, and real-time telemetry dashboard designed specifically for Linux deployment machines that simultaneously host Ollama LLM workloads and ComfyUI Diffusion pipelines on a single NVIDIA GPU.


Real-Time Telemetry & Control Dashboard

HyperSwap Dashboard

The HyperSwap live dashboard running on :9090, demonstrating real-time VRAM allocation tracking (Ollama 14.14 GB, ComfyUI 0.38 GB, Desktop 0.6 GB), 33.89 GB of models resident in 64GB host RAM cache, live dual-axis memory charts, hardware fan control, and sub-25ms model VRAM purges and soft-yields.


1. Feature Matrix

⚡ Bidirectional VRAM Hot-Swapping & Arbitration

  • Confirmed Soft-Yield (barrier, not fire-and-forget): Releases Ollama VRAM allocations (keep_alive: 0) down to 0 MB, then waits on NVML until the driver has actually freed the allocation before letting ComfyUI proceed. Posting keep_alive: 0 only asks Ollama to unload; on this box the HTTP call returns in ~63 ms while the driver takes a further ~77 ms to release 14.9 GB. Returning during that window is how diffusion ends up allocating into VRAM that is still occupied.
  • Idle-Aware ComfyUI Purge: Diffusion checkpoints are held for COMFY_IDLE_PURGE_S (30 s) of genuinely empty queue rather than purged 1.5 s after every prompt — iterating on a workflow no longer pays a full checkpoint reload per run. An immediate purge still happens the moment Ollama actually asks for VRAM (POST /api/request-vram).
  • Real-Time ComfyUI WebSocket & Watchdog Listener: Subscribes directly to ws://127.0.0.1:8188/ws. The WebSocket is the primary signal; a connection-pooled watchdog polls /queue at 1 Hz purely as a fallback, backing off to 3 s while the socket is healthy.
  • Process-Level VRAM Attribution: Live NVML process inspection attributes exact GPU memory usage across Ollama (llama-server), ComfyUI (python), and Desktop display servers (gnome-shell, Xorg).
  • Bandwidth-Classified Transition History: Every switch is classified by the bandwidth it actually achieved (model size ÷ load duration) rather than a fixed duration threshold: RAM Cache Hit ⚡ (≥5 GB/s), Partial Cache 🌤 (≥1.5 GB/s), Cold Disk Load 💾 (below that). The previous load_duration < 2500 ms rule called a 12.9 GB model read at 2.9 GB/s a "cold disk load" and a 0.5 GB model read from NVMe a "cache hit".

🧠 64GB Host RAM Cache & Page Pre-warmer

  • Zero-Latency Model Discovery: Automatic cataloging of all local Ollama models (/usr/share/ollama/.ollama/models, ~/.ollama/models) and ComfyUI model directories (checkpoints, diffusion_models, unet, vae, clip, loras, controlnet).
  • POSIX fadvise & Pinned Pre-warmer: Pre-faults multi-gigabyte GGUFs and Safetensors into the Linux OS Page Cache so that reloading models runs at page-cache speed rather than disk speed.
  • Measured Residency via cachestat(2): Residency is measured, not assumed. cachestat(2) gives exact cached-page counts per file. Where the kernel refuses it — it only permits introspection of files you own, and Ollama's blobs are owned by uid ollama — HyperSwap falls back to a randomised read-rate probe and labels the result as such. Files it cannot measure are reported as unmeasurable rather than guessed at.
  • Budgeted, Ranked Warming: This box has 64 GB of RAM and >270 GB of model files; reading everything simply evicts whatever was warmed first. Files are ranked by recency/frequency (from the persisted event log) and warmed until a byte budget is spent, skipping anything already resident. GET /api/warm-plan previews the decision without executing it.
  • Memory Telemetry: Real-time breakdown of Total Host RAM, Applications Memory, Active Model Page Cache, Free Memory, and measured Cache Residency Ratio.

🎛️ Dynamic Overclocking & Thermal Management

  • Workload-Aware Overclock Profiles:
    • ollama Profile (Memory-Bandwidth Bound): Max 370W power limit, +150 MHz Core Offset, +825 MHz Memory Offset, and 100% fan speed for maximum prompt eval / generation bandwidth.
    • comfy Profile (Compute Bound): Max 370W power limit, +100 MHz Core Offset, +500 MHz Memory Offset, Core Clock locked to 2900–3105 MHz, and 75% fan speed for maximum diffusion compute.
    • balanced Profile (Stock/General Purpose): Unlocked 370W power limit with stock dynamic boost curves and automatic fan control.
  • Hardware Actuation Hierarchy:
    • Level 1: Power Limit Control (nvidia-smi -pl 370).
    • Level 2: Core & Memory Clock Locking (nvidia-smi -lgc / -lmc).
    • Level 3: Clock Offsets via headless X display (:8) with Coolbits support (nvidia-settings).
  • Hardware Fan Control: Switch between auto and manual PWM control (30%–100%) with synchronized dual-fan actuation ([fan:0] and [fan:1]).
  • Automated Lockstep Profile Switching: AutoArbitrator automatically switches hardware profiles in lockstep with the active workload (comfy on generation start, ollama on completion).

🌡️ Thermal Governor (closed-loop de-escalation)

  • Every overclock lever here is sticky: a profile locks clocks and pins the fans to a manual PWM, and nothing used to undo that. The governor watches the telemetry the sampler already collects (so it costs no extra NVML calls) and walks the overclock back through a four-step derate ladder when the card runs hot or reports a hardware throttle.
  • Hysteresis by design: escalation needs 5 consecutive bad samples, recovery needs 30 consecutive good ones, with a 20 s cooldown between changes — a single spike during a diffusion step will not cause profile thrash.
  • Guaranteed restore: stock clocks, default power limit and automatic fans are restored by the server's shutdown hook and by a systemd ExecStopPost=, so a SIGKILL cannot leave the card with locked clocks and fans pinned at 100%.

🔬 Overclock Autotune (autotune.py)

  • Walks a clock offset upward, running a fixed decode benchmark at each step, and reports the fastest stable value with its measured gain over baseline.
  • Instability detection: kernel Xid/NVRM messages via journalctl -k, benchmark failure, degenerate output, and a temperature ceiling. The sweep stops climbing the moment a step looks unstable.
  • Safety: refuses to start while ComfyUI is executing, and restores the original profile in a finally block — including on exception or cancellation.

🗄️ Persistent Telemetry Store (telemetry_store.py)

  • Swap history used to be an in-memory deque(maxlen=50) that evaporated on every restart. Telemetry and events now persist to SQLite (WAL, single writer thread, batched 1 Hz inserts, automatic retention pruning) at roughly 0.4 MB per hour.
  • This is what makes the app's central question answerable: GET /api/analytics/profiles compares decode throughput per overclock profile, joined against the thermals recorded while that profile was active.

📊 Real-Time Web Telemetry Dashboard (:9090)

  • Live Hardware Telemetry: GPU utilization %, GPU temperature (°C), power draw (W), fan speeds (%), and graphics/memory clock frequencies (MHz).
  • Live Dual-Axis Time-Series Chart: Real-time graphical visualization of VRAM usage (GB) and Host RAM Cache (GB) with zero frontend polling overhead.
  • Interactive Control Center: Trigger model hot-swaps, soft-yields, cache pre-warms, fan adjustments, and benchmarks directly from the web interface.
  • Server-Sent Events (SSE): A single background sampler produces one 1 Hz snapshot and fans it out to every subscriber via GET /api/stream. Previously each connected client independently re-ran the whole snapshot — NVML, /proc/meminfo, an HTTP round-trip each to Ollama and ComfyUI, and a recursive walk of the ComfyUI models tree with a stat() per checkpoint — once per second, so opening the dashboard in three tabs tripled the load on the thing it was measuring. Slow clients drop stale frames instead of stalling the sampler.

🤖 Model Context Protocol (MCP 2.0) Server

  • 12 Native Agentic Tools: Allows AI agents (Antigravity CLI, Claude Desktop, Cursor) to manage GPU resources, trigger model hot-swaps, tune fan curves, and inspect telemetry.
  • 3 Live MCP Resources: Exposes live metrics, model catalogs, and switch logs as streamable resources (gpu://metrics/live, gpu://models/catalog, gpu://history/switches).
  • Dual Transport Support: Run via standard input/output (--stdio) or network Server-Sent Events (--sse --port 8001).

⏱️ Automated Latency & Throughput Benchmark Engine

  • Conducts automated round-trip model switching benchmarks to measure transition latency, model load time, tokens per second, and RAM cache effectiveness.

2. Architectural Overview

flowchart TD
    subgraph HostRAM["64 GB Host System RAM (Page Cache & Staging Buffer)"]
        OllamaGGUFs["Ollama GGUF Weights<br/>(Qwen, Gemma, Nemotron)"]
        ComfySafetensors["ComfyUI Safetensors & VAEs<br/>(Wan2.1, Flux, SDXL)"]
    end

    subgraph GPU["NVIDIA GeForce RTX 4080 SUPER (16 GB VRAM)"]
        direction LR
        ActiveLLM["Active LLM<br/>(0–15 GB VRAM)"]
        ActiveDiffusion["Active Diffusion Pipeline<br/>(0–15 GB VRAM)"]
    end

    subgraph Orchestrator["HyperSwap Control Plane (:9090)"]
        REST["REST API & OpenAPI Docs"]
        MCP["Model Context Protocol (MCP 2.0)"]
        SSE["1Hz Real-Time SSE Stream"]
        Arbitrator["VRAM Arbitrator (15ms Soft-Yield)"]
        Overclock["Overclock & Fan Manager"]
        Warmer["Page Cache Pre-Warmer"]
    end

    HostRAM <== "PCIe 4.0 x16 Bus (~31.5 GB/s Hot-Swap)" ==> GPU
    Orchestrator --> GPU
    Orchestrator --> HostRAM

The Physics of Sub-Second Switching

  • Host RAM as Staging: Active LLMs and diffusion checkpoints remain resident in the 64GB Linux Page Cache.
  • PCIe 4.0 x16 Hot-Swapping: Transferring weights across PCIe 4.0 x16 achieves ~31.5 GB/s bandwidth, reducing model loads from 30+ seconds (disk) to under 1.5 seconds.
  • Soft-Yielding: Dropping Ollama's VRAM allocation via keep_alive: 0 preserves the weights in host RAM. Measured on this box: the HTTP request returns in ~63 ms, and the driver finishes releasing 14.9 GB ~77 ms after that. HyperSwap waits for the second number before handing VRAM to ComfyUI — the earlier "~15 ms" figure timed the request, not the release.

3. REST API Reference

The HyperSwap server runs on port 9090 by default. Interactive OpenAPI/Swagger docs are available at http://localhost:9090/docs.

Telemetry & Hardware Endpoints

Endpoint Method Description
/api/stats GET Complete unified JSON snapshot of hardware sensors, VRAM breakdown, host RAM, Ollama status, ComfyUI queue, and switch logs.
/api/gpu GET NVIDIA GPU sensors (utilization %, temperature, power draw in Watts, fan speeds, clocks, and active PIDs).
/api/memory GET Precise /proc/meminfo metrics (Total, Used, OS Page Cache containing models, Free memory).
/api/gpu/fan GET Current GPU fan mode (auto vs manual), target speed %, and live fan RPM/PWM status.
/api/gpu/fan POST Sets GPU fan speed mode (auto or manual) with target speed % (30–100%).
/api/overclock GET Active overclock profile, configured profiles, GPU clock limits, and fan status.
/api/overclock/apply POST Applies a named profile (ollama, comfy, balanced).
/api/overclock/profile POST Creates or updates an overclock profile configuration.
/api/stream GET Server-Sent Events (SSE) stream pushing full telemetry updates at 1Hz (text/event-stream).

Model Orchestration & Hot-Swap Endpoints

Endpoint Method Description
/api/switch-model POST Hot-swaps the active Ollama LLM in VRAM and tracks transition timing.
/api/free-vram POST Soft-yields Ollama VRAM to 0 MB and waits for NVML to confirm the release (?confirm=false to skip). Returns request_ms, confirm_ms and the GB actually freed.
/api/comfy-free POST Instructs ComfyUI to purge loaded diffusion weights and VRAM cache.
/api/request-vram POST Ollama-priority path: purges ComfyUI immediately if there is not enough free VRAM.
/api/warm-all POST Warms the highest-value models into page cache within a byte budget (budget_gb).
/api/warm-plan GET Previews what warming would read, in what order, and what it would skip — without doing it.
/api/warm-model POST Pre-warms a specific model or file into RAM (blob_only warms weights without touching VRAM).
/api/cache/report GET Measured page-cache residency per model file, with the measurement method used for each.
/api/benchmark POST Runs an automated back-and-forth model swap benchmark and calculates average latency.

Analytics Endpoints (persisted)

Endpoint Method Description
/api/analytics/profiles GET Decode throughput per overclock profile, joined with the thermals recorded under it.
/api/analytics/swaps GET Aggregated swap/yield/purge latencies, cache-hit split, and per-model throughput.
/api/analytics/timeseries GET Downsampled telemetry history for charts that outlive a page refresh.
/api/analytics/models GET Recency/frequency model ranking used to prioritise the warm budget.
/api/history?durable=true GET Swap history from the persistent store rather than the in-memory ring.
/api/db GET Store location, row counts and how many hours of history are held.

Governor & Autotune Endpoints

Endpoint Method Description
/api/governor GET / POST Current derate level and why; enable/disable, or clear an active derate.
/api/overclock/restore POST Drop all clock locks and offsets, restore default power limit and automatic fans.
/api/autotune GET Sweep progress, last result, and every recorded autotune step.
/api/autotune/sweep POST Walk a clock offset upward, measuring tok/s and watching for instability at each step.
/api/autotune/cancel POST Stop the current sweep after the step in flight; the profile is restored either way.

4. Model Context Protocol (MCP 2.0) Reference

HyperSwap includes a native MCP 2.0 server (mcp_server.py) exposing orchestration and telemetry tools to AI agents.

MCP Tools List

Tool Name Parameters Description
get_gpu_status None Live NVIDIA GPU hardware telemetry, VRAM breakdown, temps, power, fan %, and PIDs.
get_gpu_fan_status None Current GPU fan mode (auto/manual) and target fan percentage.
set_gpu_fan_speed mode (str), percent (optional int) Sets fan speed mode (auto|manual) and target PWM % (30–100%).
get_host_memory_status None 64GB host RAM breakdown, active page cache size, and cache ratio.
switch_ollama_model model_name (str), keep_alive (str) Hot-swaps active LLM in VRAM, measures latency (ms) and tokens/sec.
soft_yield_ollama_vram model_name (optional str) Yields Ollama VRAM to 0 MB in ~15ms while keeping model weights in RAM cache.
purge_comfyui_vram None Purges loaded diffusion models from ComfyUI pipeline VRAM.
prewarm_all_models_to_ram None Faults all local LLM and diffusion checkpoints into Linux OS page cache.
prewarm_single_model model_name (optional str), filepath (optional str) Pre-warms a single GGUF or Safetensors file into RAM.
list_available_models None Lists all installed Ollama models and discovered ComfyUI Safetensors on disk.
get_switch_history limit (int, default 20) Retrieves recent switch events, millisecond latencies, and RAM hit status.
run_model_switch_benchmark iterations (int, default 2) Automated round-trip latency benchmark between installed models.

MCP Resources List

  • gpu://metrics/live: Real-time snapshot of GPU sensors and RAM page cache.
  • gpu://models/catalog: Catalog of all discovered GGUF and Safetensors models.
  • gpu://history/switches: Event log of recent model transitions and swap speeds.

MCP Client Configurations

Antigravity Configuration (~/.gemini/antigravity-cli/mcp_config.json)

{
  "mcpServers": {
    "hyperswap": {
      "command": "/home/drjones/comfy-mcp-venv/bin/python",
      "args": ["/home/drjones/unified-model-manager/mcp_server.py", "--stdio"]
    }
  }
}

Claude Desktop Configuration (claude_desktop_config.json)

{
  "mcpServers": {
    "hyperswap": {
      "command": "/home/drjones/comfy-mcp-venv/bin/python",
      "args": ["/home/drjones/unified-model-manager/mcp_server.py", "--stdio"]
    }
  }
}

5. Linux Kernel & Host Tuning

To ensure that model weights remain permanently in RAM without kernel eviction:

# Set CPU scaling governor to performance
sudo cpupower frequency-set -g performance

# Configure sysctl optimizations in /etc/sysctl.d/99-hyperswap.conf
cat << 'EOF' | sudo tee /etc/sysctl.d/99-hyperswap.conf
# Retain model file cache aggressively in RAM
vm.vfs_cache_pressure = 50

# Prevent swapping cached models
vm.swappiness = 10

# Support large memory maps for high-parameter models
vm.max_map_count = 1048576

# Flush dirty pages quickly
vm.dirty_background_ratio = 5
vm.dirty_ratio = 10
EOF

# Apply sysctl settings immediately
sudo sysctl --system

6. Systemd Service Management

The HyperSwap server runs as a systemd service:

# Check service status
systemctl status hyperswap.service

# Restart service
sudo systemctl restart hyperswap.service

# View live telemetry and arbitration logs
journalctl -u hyperswap.service -f

7. License

MIT License. Developed for Google Antigravity & High-Throughput Linux AI Deployments.

Description
a program to run ollama and comfy on one gpu
Readme 1.4 MiB
Languages
Python 78.1%
JavaScript 10.9%
HTML 10.8%
CSS 0.1%