Add test suite (164 tests); reclaim VRAM from ComfyUI when an LLM will not fit
Tests. First automated coverage for the project: 164 tests, 2.7s, no GPU or network. An autouse fixture stubs overclock_manager._sh -- the single choke point for every nvidia-smi/nvidia-settings write -- so no test can mutate the card. They deliberately pin the empirically measured constants that would otherwise rot silently: the cold and warm load figures behind the cache-hit thresholds, the warm_confident residency rule, and the busy/stalled yield split. One test asserts RAM_HIT_GBPS stays at or below the measured 2.63 GB/s warm load, so the old physically unreachable 5.0 GB/s bar cannot come back. Three bugs the suite surfaced, now fixed: - autotune._subsample(values, 1) divided by zero; the early return only covered len(values) <= max_steps. - telemetry_store.stop() flushed its local pending list but never drained the queue, silently losing rows submitted just before a shutdown -- exactly when the last events matter. - ram_optimizer.page_residency's zero-byte short-circuit omitted keys every other return path provides, so a 0-byte file was planned for warming. Reclaim. The README has claimed bidirectional arbitration from the start, but only one direction was ever automatic. Establishing what actually happens took a controlled test with the service stopped: with ComfyUI holding 6.83 GB, Ollama does not spill to the CPU on this box -- it aborts with "cudaMalloc failed: out of memory", because n_gpu_layers is pinned to 99 and it will not reduce the layer count. So both failure modes are handled: _check_ollama_starved watches size_vram < size for the default configuration where Ollama does spill, and switch_ollama_model catches the hard OOM, reclaims VRAM from an idle ComfyUI and retries once. The request that returned HTTP 500 from Ollama directly now succeeds through HyperSwap, loading at 3.85 GB/s after reclaiming 6.83 GB. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
15
README.md
15
README.md
@@ -20,6 +20,21 @@
|
||||
## 1. Feature Matrix
|
||||
|
||||
### ⚡ Bidirectional VRAM Hot-Swapping & Arbitration
|
||||
|
||||
* **Both directions are now automatic.** Yielding Ollama for ComfyUI always was; the
|
||||
reverse was not, despite "bidirectional" in this heading. Which way an LLM fails when
|
||||
it cannot fit depends on configuration: with `n_gpu_layers` left to Ollama it spills
|
||||
layers to the CPU and reports `size_vram < size` (roughly an order of magnitude slower,
|
||||
and silent). With `n_gpu_layers` pinned — 99 on this box — it refuses outright with
|
||||
`cudaMalloc failed: out of memory`. Both are handled: the spill triggers a reclaim from
|
||||
an idle ComfyUI, and the hard failure is caught by `switch_ollama_model`, which reclaims
|
||||
and retries once. Measured: a 12.87 GB model that returned HTTP 500 from Ollama directly
|
||||
now loads through HyperSwap after reclaiming 6.83 GB, at 3.85 GB/s.
|
||||
* **A busy LLM is not a failed yield.** A model mid-generation cannot unload; the
|
||||
`keep_alive: 0` request queues behind it and applies when it finishes. That is reported
|
||||
as `busy` (returning in ~610 ms) rather than blocking, with per-model backoff and a
|
||||
detached watcher that logs the eventual release. Only VRAM held while the GPU sits
|
||||
*idle* counts as a fault.
|
||||
* **Confirmed Soft-Yield (barrier, not fire-and-forget)**: Releases Ollama VRAM allocations (`keep_alive: 0`) down to 0 MB, then **waits on NVML until the driver has actually freed the allocation** before letting ComfyUI proceed. Posting `keep_alive: 0` only *asks* Ollama to unload; on this box the HTTP call returns in ~63 ms while the driver takes a further ~77 ms to release 14.9 GB. Returning during that window is how diffusion ends up allocating into VRAM that is still occupied.
|
||||
* **Idle-Aware ComfyUI Purge**: Diffusion checkpoints are held for `COMFY_IDLE_PURGE_S` (30 s) of genuinely empty queue rather than purged 1.5 s after every prompt — iterating on a workflow no longer pays a full checkpoint reload per run. An immediate purge still happens the moment Ollama actually asks for VRAM (`POST /api/request-vram`).
|
||||
* **Real-Time ComfyUI WebSocket & Watchdog Listener**: Subscribes directly to `ws://127.0.0.1:8188/ws`. The WebSocket is the primary signal; a connection-pooled watchdog polls `/queue` at 1 Hz purely as a fallback, backing off to 3 s while the socket is healthy.
|
||||
|
||||
Reference in New Issue
Block a user