Files
gpu-program-swapper/tests
drjones c3d9b36035 Stop a stale ComfyUI queue entry from disabling half the arbitration
Chasing why the reverse-direction reclaim never fired turned up something worse than
the reclaim itself.

The starvation check was never running. Instrumenting the watchdog showed busy=6,
idle_check=0: every poll took the "ComfyUI is busy" branch. ComfyUI's /queue was
reporting a WAN 2.1 i2v job in queue_running while the GPU sat at 0% and ComfyUI held
0.56 GB. The job was dead; ComfyUI had simply never cleared the row.

Believing that flag meant this service thought ComfyUI was permanently busy, so it
yielded the LLM's VRAM on every poll, never ran the idle purge, and never checked
whether the LLM had been squeezed onto the CPU. One stale row disabled half of the
arbitration, and it very likely explains the earlier burst of yields against a
cron-driven model.

A running entry is now corroborated before it is believed. The first attempt used GPU
utilisation, which does not work: utilisation is shared with Ollama and with the
third-party process on this box, so peak utilisation stayed above any sensible
threshold and a stuck entry never looked stale. ComfyUI's own VRAM is the right
signal -- a real diffusion job loads gigabytes of checkpoint, a dead one holds only
its CUDA context. After the fix the same watchdog reports busy=3, idle_check=32.

Every early return in the starvation check now records why it bailed, because with
four of them there was no way to tell which had fired. /api/health reports a stale
queue entry with its impact and how to clear it.

Also confirmed, contradicting an earlier conclusion in this branch: Ollama on this box
*does* spill to the CPU. smtek/Qwen3.8-27B:Q2_K_XL held steady at 29.2% on GPU
(size=15.59 GB, size_vram=4.56 GB) across twelve seconds of polling -- a stable
placement, not a progressive load. Both failure modes are real; which one occurs
depends on the model.

Tests: 206 (was 199).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 08:48:37 -07:00
..

HyperSwap test suite

Fast, hermetic unit tests. No GPU is touched, no network call is made, no systemd unit is poked, and the production hyperswap.db is never opened.

Running

/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/ -q

Single file / single test:

/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/test_classify_load.py -q
/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/ -q -k warm_confident

Whole suite runs in about 3 seconds.

Safety rails

These matter, because this repo drives a live 4080 SUPER that a running service is using.

  • tests/conftest.py installs an autouse no_gpu_mutation fixture that replaces overclock_manager._sh (the single choke point for every nvidia-smi / nvidia-settings write) plus apply_profile, apply_fan_control, set_fan_speed, set_fan_auto and restore_safe with recording stubs. Even a test that accidentally reaches an actuation path can only reach the stub. The fixture yields a dict of recorded calls, which the thermal tests assert against.
  • HYPERSWAP_DB is set to a non-existent path before telemetry_store is imported, so no import can bind DB_PATH to the production database. Tests that need a DB use the temp_db fixture, which monkeypatches telemetry_store.DB_PATH to a tmp_path file and stops the writer thread afterwards.
  • All file IO happens against files the tests create in tmp_path. No real model blob is read and warm_file_to_ram is never called.
  • Nothing sweeps, and nothing sends HTTP to Ollama, ComfyUI or :9090.

Measured constants pinned here

These numbers came from measurement on this box, not from taste. If a change makes one of these tests fail, the constant is probably wrong, not the test.

Constant Value Where pinned
Cold load of a 12.87 GB model, 3.1% resident 34267 ms → 0.38 GB/s test_classify_load.py::test_measured_cold_load_classifies_as_cold_disk
Warm load of the same model, 100% resident 4901 ms → 2.63 GB/s test_classify_load.py::test_measured_warm_load_classifies_as_ram_hit
RAM_HIT_GBPS = 2.0 must stay below the fastest achievable warm load (2.63 GB/s) — test_classify_load.py::test_ram_hit_threshold_is_physically_achievable
PARTIAL_HIT_GBPS = 0.8 must stay above the measured cold rate (0.38 GB/s) — same test
Size-unknown fallback splits at 8000 ms (between 4.9 s warm and 34.3 s cold) — test_classify_load.py::test_unknown_size_guess_boundary_is_8s
WARM_SKIP_THRESHOLD_PCT = 90.0 — test_ram_optimizer.py::test_warm_skip_threshold_constant_unchanged
A probe reading may only be trusted at exactly 100% (a 12-window probe once cleared 90% on a mostly-cold 12.87 GB blob that then loaded at 2.44 GB/s) — test_ram_optimizer.py::test_probe_reading_is_only_trusted_at_exactly_100_percent
PROBE_CACHED_GBPS = 1.5 sits in the gap between cold NVMe (0.35–0.5 GB/s) and page cache (3.2–13 GB/s) — test_ram_optimizer.py::test_probe_cached_threshold_sits_between_measured_disk_and_cache_rates
Card power envelope: 320 W stock, 370 W max, sweeps never go below 60% of max — test_autotune_helpers.py::test_supported_power_limits_parses_min_default_max
_supported_clocks must always query the mem,gr pair (a single-field query returns one column and silently yielded []) — test_autotune_helpers.py::test_supported_clocks_always_queries_the_mem_gr_pair
ComfyUI benchmark seed must vary per call (a fixed seed made ComfyUI serve a cached result in ~1 ms) — test_autotune_helpers.py::test_comfy_workflow_seed_varies_between_calls
Governor hysteresis: HOT_SAMPLES = 5, COOL_SAMPLES = 30, REAPPLY_COOLDOWN_S = 20 — test_thermal_governor.py (escalation, recovery, cooldown, alternating-sample tests)
Model usage score: frequency decayed with a ~24 h half-life — test_telemetry_store.py::test_model_usage_ranking_scores_recent_use_higher

What is deliberately not covered

  • vram_arbitrator.instant_free_ollama_vram, the AutoArbitrator yield/purge paths and the SSE broker — under active edit, contract changing.
  • overclock_manager.apply_profile and every other actuation path, autotune.sweep, ram_optimizer.warm_file_to_ram — these mutate hardware or do heavy IO.
  • server.py HTTP routes and mcp_server.py — would need the app wired to live subsystems.

Known rough edge the tests work around

telemetry_store.stop() flushes the writer's pending batch but does not drain the submission queue, so a stop() racing a just-submitted row can drop it. The writer tests call a local _drain() helper to wait for the queue to empty before stopping, rather than encoding the race into an assertion.