Add test suite (164 tests); reclaim VRAM from ComfyUI when an LLM will not fit

Tests. First automated coverage for the project: 164 tests, 2.7s, no GPU or network.
An autouse fixture stubs overclock_manager._sh -- the single choke point for every
nvidia-smi/nvidia-settings write -- so no test can mutate the card. They deliberately
pin the empirically measured constants that would otherwise rot silently: the cold and
warm load figures behind the cache-hit thresholds, the warm_confident residency rule,
and the busy/stalled yield split. One test asserts RAM_HIT_GBPS stays at or below the
measured 2.63 GB/s warm load, so the old physically unreachable 5.0 GB/s bar cannot
come back.

Three bugs the suite surfaced, now fixed:
- autotune._subsample(values, 1) divided by zero; the early return only covered
  len(values) <= max_steps.
- telemetry_store.stop() flushed its local pending list but never drained the queue,
  silently losing rows submitted just before a shutdown -- exactly when the last
  events matter.
- ram_optimizer.page_residency's zero-byte short-circuit omitted keys every other
  return path provides, so a 0-byte file was planned for warming.

Reclaim. The README has claimed bidirectional arbitration from the start, but only one
direction was ever automatic. Establishing what actually happens took a controlled test
with the service stopped: with ComfyUI holding 6.83 GB, Ollama does not spill to the CPU
on this box -- it aborts with "cudaMalloc failed: out of memory", because n_gpu_layers is
pinned to 99 and it will not reduce the layer count. So both failure modes are handled:
_check_ollama_starved watches size_vram < size for the default configuration where Ollama
does spill, and switch_ollama_model catches the hard OOM, reclaims VRAM from an idle
ComfyUI and retries once. The request that returned HTTP 500 from Ollama directly now
succeeds through HyperSwap, loading at 3.85 GB/s after reclaiming 6.83 GB.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
drjones
2026-09-01 13:40:30 -07:00
parent bacaf50713
commit 868d82794d
16 changed files with 1972 additions and 7 deletions

74
tests/README.md Normal file
View File

@@ -0,0 +1,74 @@
# HyperSwap test suite
Fast, hermetic unit tests. No GPU is touched, no network call is made, no systemd unit is
poked, and the production `hyperswap.db` is never opened.
## Running
```bash
/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/ -q
```
Single file / single test:
```bash
/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/test_classify_load.py -q
/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/ -q -k warm_confident
```
Whole suite runs in about 3 seconds.
## Safety rails
These matter, because this repo drives a live 4080 SUPER that a running service is using.
* `tests/conftest.py` installs an **autouse** `no_gpu_mutation` fixture that replaces
`overclock_manager._sh` (the single choke point for every `nvidia-smi` /
`nvidia-settings` write) plus `apply_profile`, `apply_fan_control`, `set_fan_speed`,
`set_fan_auto` and `restore_safe` with recording stubs. Even a test that accidentally
reaches an actuation path can only reach the stub. The fixture yields a dict of
recorded calls, which the thermal tests assert against.
* `HYPERSWAP_DB` is set to a non-existent path before `telemetry_store` is imported, so no
import can bind `DB_PATH` to the production database. Tests that need a DB use the
`temp_db` fixture, which monkeypatches `telemetry_store.DB_PATH` to a `tmp_path` file
and stops the writer thread afterwards.
* All file IO happens against files the tests create in `tmp_path`. No real model blob is
read and `warm_file_to_ram` is never called.
* Nothing sweeps, and nothing sends HTTP to Ollama, ComfyUI or :9090.
## Measured constants pinned here
These numbers came from measurement on this box, not from taste. If a change makes one of
these tests fail, the constant is probably wrong, not the test.
| Constant | Value | Where pinned |
| --- | --- | --- |
| Cold load of a 12.87 GB model, 3.1% resident | 34267 ms → 0.38 GB/s | `test_classify_load.py::test_measured_cold_load_classifies_as_cold_disk` |
| Warm load of the same model, 100% resident | 4901 ms → 2.63 GB/s | `test_classify_load.py::test_measured_warm_load_classifies_as_ram_hit` |
| `RAM_HIT_GBPS = 2.0` must stay below the fastest achievable warm load (2.63 GB/s) | — | `test_classify_load.py::test_ram_hit_threshold_is_physically_achievable` |
| `PARTIAL_HIT_GBPS = 0.8` must stay above the measured cold rate (0.38 GB/s) | — | same test |
| Size-unknown fallback splits at 8000 ms (between 4.9 s warm and 34.3 s cold) | — | `test_classify_load.py::test_unknown_size_guess_boundary_is_8s` |
| `WARM_SKIP_THRESHOLD_PCT = 90.0` | — | `test_ram_optimizer.py::test_warm_skip_threshold_constant_unchanged` |
| A *probe* reading may only be trusted at exactly 100% (a 12-window probe once cleared 90% on a mostly-cold 12.87 GB blob that then loaded at 2.44 GB/s) | — | `test_ram_optimizer.py::test_probe_reading_is_only_trusted_at_exactly_100_percent` |
| `PROBE_CACHED_GBPS = 1.5` sits in the gap between cold NVMe (0.35–0.5 GB/s) and page cache (3.2–13 GB/s) | — | `test_ram_optimizer.py::test_probe_cached_threshold_sits_between_measured_disk_and_cache_rates` |
| Card power envelope: 320 W stock, 370 W max, sweeps never go below 60% of max | — | `test_autotune_helpers.py::test_supported_power_limits_parses_min_default_max` |
| `_supported_clocks` must always query the `mem,gr` pair (a single-field query returns one column and silently yielded `[]`) | — | `test_autotune_helpers.py::test_supported_clocks_always_queries_the_mem_gr_pair` |
| ComfyUI benchmark seed must vary per call (a fixed seed made ComfyUI serve a cached result in ~1 ms) | — | `test_autotune_helpers.py::test_comfy_workflow_seed_varies_between_calls` |
| Governor hysteresis: `HOT_SAMPLES = 5`, `COOL_SAMPLES = 30`, `REAPPLY_COOLDOWN_S = 20` | — | `test_thermal_governor.py` (escalation, recovery, cooldown, alternating-sample tests) |
| Model usage score: frequency decayed with a ~24 h half-life | — | `test_telemetry_store.py::test_model_usage_ranking_scores_recent_use_higher` |
## What is deliberately not covered
* `vram_arbitrator.instant_free_ollama_vram`, the `AutoArbitrator` yield/purge paths and
the SSE broker — under active edit, contract changing.
* `overclock_manager.apply_profile` and every other actuation path, `autotune.sweep`,
`ram_optimizer.warm_file_to_ram` — these mutate hardware or do heavy IO.
* `server.py` HTTP routes and `mcp_server.py` — would need the app wired to live
subsystems.
## Known rough edge the tests work around
`telemetry_store.stop()` flushes the writer's pending *batch* but does not drain the
submission queue, so a `stop()` racing a just-submitted row can drop it. The writer tests
call a local `_drain()` helper to wait for the queue to empty before stopping, rather than
encoding the race into an assertion.