The queue shipped with no tests despite having produced four bugs during development, which is the wrong order. 21 tests now cover the parts whose failure modes are not obvious from reading the code: ordering by priority then FIFO, that depth is genuinely unbounded and survives a restart, that running work is never cancelled, that a job abandoned mid-run is requeued rather than left RUNNING forever, and that an LLM job's VRAM requirement comes from its own model rather than a tenant-wide figure -- the mistake that dispatched a 14.9 GB model into 8 GB of free memory and killed llama-server three times. The dashboard had no view of the queue at all, so a stuck queue was indistinguishable from an empty one. The Job Queue panel shows what is running and what it released to get there, pending jobs in execution order with their wait time and a cancel control, and -- when the scheduler is blocked -- how long it has been waiting and why. MCP gains queue_job, get_job_queue and cancel_job, so an agent can line work up rather than firing a request and hoping the GPU is free. Tests: 271. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
HyperSwap test suite
Fast, hermetic unit tests. No GPU is touched, no network call is made, no systemd unit is
poked, and the production hyperswap.db is never opened.
Running
/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/ -q
Single file / single test:
/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/test_classify_load.py -q
/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/ -q -k warm_confident
Whole suite runs in about 3 seconds.
Safety rails
These matter, because this repo drives a live 4080 SUPER that a running service is using.
tests/conftest.pyinstalls an autouseno_gpu_mutationfixture that replacesoverclock_manager._sh(the single choke point for everynvidia-smi/nvidia-settingswrite) plusapply_profile,apply_fan_control,set_fan_speed,set_fan_autoandrestore_safewith recording stubs. Even a test that accidentally reaches an actuation path can only reach the stub. The fixture yields a dict of recorded calls, which the thermal tests assert against.HYPERSWAP_DBis set to a non-existent path beforetelemetry_storeis imported, so no import can bindDB_PATHto the production database. Tests that need a DB use thetemp_dbfixture, which monkeypatchestelemetry_store.DB_PATHto atmp_pathfile and stops the writer thread afterwards.- All file IO happens against files the tests create in
tmp_path. No real model blob is read andwarm_file_to_ramis never called. - Nothing sweeps, and nothing sends HTTP to Ollama, ComfyUI or :9090.
Measured constants pinned here
These numbers came from measurement on this box, not from taste. If a change makes one of these tests fail, the constant is probably wrong, not the test.
| Constant | Value | Where pinned |
|---|---|---|
| Cold load of a 12.87 GB model, 3.1% resident | 34267 ms → 0.38 GB/s | test_classify_load.py::test_measured_cold_load_classifies_as_cold_disk |
| Warm load of the same model, 100% resident | 4901 ms → 2.63 GB/s | test_classify_load.py::test_measured_warm_load_classifies_as_ram_hit |
RAM_HIT_GBPS = 2.0 must stay below the fastest achievable warm load (2.63 GB/s) |
— | test_classify_load.py::test_ram_hit_threshold_is_physically_achievable |
PARTIAL_HIT_GBPS = 0.8 must stay above the measured cold rate (0.38 GB/s) |
— | same test |
| Size-unknown fallback splits at 8000 ms (between 4.9 s warm and 34.3 s cold) | — | test_classify_load.py::test_unknown_size_guess_boundary_is_8s |
WARM_SKIP_THRESHOLD_PCT = 90.0 |
— | test_ram_optimizer.py::test_warm_skip_threshold_constant_unchanged |
| A probe reading may only be trusted at exactly 100% (a 12-window probe once cleared 90% on a mostly-cold 12.87 GB blob that then loaded at 2.44 GB/s) | — | test_ram_optimizer.py::test_probe_reading_is_only_trusted_at_exactly_100_percent |
PROBE_CACHED_GBPS = 1.5 sits in the gap between cold NVMe (0.35–0.5 GB/s) and page cache (3.2–13 GB/s) |
— | test_ram_optimizer.py::test_probe_cached_threshold_sits_between_measured_disk_and_cache_rates |
| Card power envelope: 320 W stock, 370 W max, sweeps never go below 60% of max | — | test_autotune_helpers.py::test_supported_power_limits_parses_min_default_max |
_supported_clocks must always query the mem,gr pair (a single-field query returns one column and silently yielded []) |
— | test_autotune_helpers.py::test_supported_clocks_always_queries_the_mem_gr_pair |
| ComfyUI benchmark seed must vary per call (a fixed seed made ComfyUI serve a cached result in ~1 ms) | — | test_autotune_helpers.py::test_comfy_workflow_seed_varies_between_calls |
Governor hysteresis: HOT_SAMPLES = 5, COOL_SAMPLES = 30, REAPPLY_COOLDOWN_S = 20 |
— | test_thermal_governor.py (escalation, recovery, cooldown, alternating-sample tests) |
| Model usage score: frequency decayed with a ~24 h half-life | — | test_telemetry_store.py::test_model_usage_ranking_scores_recent_use_higher |
What is deliberately not covered
vram_arbitrator.instant_free_ollama_vram, theAutoArbitratoryield/purge paths and the SSE broker — under active edit, contract changing.overclock_manager.apply_profileand every other actuation path,autotune.sweep,ram_optimizer.warm_file_to_ram— these mutate hardware or do heavy IO.server.pyHTTP routes andmcp_server.py— would need the app wired to live subsystems.
Known rough edge the tests work around
telemetry_store.stop() flushes the writer's pending batch but does not drain the
submission queue, so a stop() racing a just-submitted row can drop it. The writer tests
call a local _drain() helper to wait for the queue to empty before stopping, rather than
encoding the race into an assertion.