Add test suite (164 tests); reclaim VRAM from ComfyUI when an LLM will not fit

Tests. First automated coverage for the project: 164 tests, 2.7s, no GPU or network.
An autouse fixture stubs overclock_manager._sh -- the single choke point for every
nvidia-smi/nvidia-settings write -- so no test can mutate the card. They deliberately
pin the empirically measured constants that would otherwise rot silently: the cold and
warm load figures behind the cache-hit thresholds, the warm_confident residency rule,
and the busy/stalled yield split. One test asserts RAM_HIT_GBPS stays at or below the
measured 2.63 GB/s warm load, so the old physically unreachable 5.0 GB/s bar cannot
come back.

Three bugs the suite surfaced, now fixed:
- autotune._subsample(values, 1) divided by zero; the early return only covered
  len(values) <= max_steps.
- telemetry_store.stop() flushed its local pending list but never drained the queue,
  silently losing rows submitted just before a shutdown -- exactly when the last
  events matter.
- ram_optimizer.page_residency's zero-byte short-circuit omitted keys every other
  return path provides, so a 0-byte file was planned for warming.

Reclaim. The README has claimed bidirectional arbitration from the start, but only one
direction was ever automatic. Establishing what actually happens took a controlled test
with the service stopped: with ComfyUI holding 6.83 GB, Ollama does not spill to the CPU
on this box -- it aborts with "cudaMalloc failed: out of memory", because n_gpu_layers is
pinned to 99 and it will not reduce the layer count. So both failure modes are handled:
_check_ollama_starved watches size_vram < size for the default configuration where Ollama
does spill, and switch_ollama_model catches the hard OOM, reclaims VRAM from an idle
ComfyUI and retries once. The request that returned HTTP 500 from Ollama directly now
succeeds through HyperSwap, loading at 3.85 GB/s after reclaiming 6.83 GB.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
drjones
2026-09-01 13:40:30 -07:00
parent bacaf50713
commit 868d82794d
16 changed files with 1972 additions and 7 deletions

View File

@@ -20,6 +20,21 @@
## 1. Feature Matrix
### ⚡ Bidirectional VRAM Hot-Swapping & Arbitration
* **Both directions are now automatic.** Yielding Ollama for ComfyUI always was; the
reverse was not, despite "bidirectional" in this heading. Which way an LLM fails when
it cannot fit depends on configuration: with `n_gpu_layers` left to Ollama it spills
layers to the CPU and reports `size_vram < size` (roughly an order of magnitude slower,
and silent). With `n_gpu_layers` pinned — 99 on this box — it refuses outright with
`cudaMalloc failed: out of memory`. Both are handled: the spill triggers a reclaim from
an idle ComfyUI, and the hard failure is caught by `switch_ollama_model`, which reclaims
and retries once. Measured: a 12.87 GB model that returned HTTP 500 from Ollama directly
now loads through HyperSwap after reclaiming 6.83 GB, at 3.85 GB/s.
* **A busy LLM is not a failed yield.** A model mid-generation cannot unload; the
`keep_alive: 0` request queues behind it and applies when it finishes. That is reported
as `busy` (returning in ~610 ms) rather than blocking, with per-model backoff and a
detached watcher that logs the eventual release. Only VRAM held while the GPU sits
*idle* counts as a fault.
* **Confirmed Soft-Yield (barrier, not fire-and-forget)**: Releases Ollama VRAM allocations (`keep_alive: 0`) down to 0 MB, then **waits on NVML until the driver has actually freed the allocation** before letting ComfyUI proceed. Posting `keep_alive: 0` only *asks* Ollama to unload; on this box the HTTP call returns in ~63 ms while the driver takes a further ~77 ms to release 14.9 GB. Returning during that window is how diffusion ends up allocating into VRAM that is still occupied.
* **Idle-Aware ComfyUI Purge**: Diffusion checkpoints are held for `COMFY_IDLE_PURGE_S` (30 s) of genuinely empty queue rather than purged 1.5 s after every prompt — iterating on a workflow no longer pays a full checkpoint reload per run. An immediate purge still happens the moment Ollama actually asks for VRAM (`POST /api/request-vram`).
* **Real-Time ComfyUI WebSocket & Watchdog Listener**: Subscribes directly to `ws://127.0.0.1:8188/ws`. The WebSocket is the primary signal; a connection-pooled watchdog polls `/queue` at 1 Hz purely as a fallback, backing off to 3 s while the socket is healthy.