Measured failure: a diffusion job took 46 s instead of 3 s, squeezed into 1.6 GB. The LLM yielded correctly, an inference reloaded it two seconds later, and plan_release then refused to touch it because it was "busy" -- so ComfyUI crawled while Ollama held 13 GB for the whole run. "Never interrupt busy work" looks like the safe rule and is not. Preempting a lower-priority tenant is safe precisely because releasing is asynchronous: an Ollama unload queues behind its running request and applies when that finishes, so nothing is killed mid-flight. That is what makes fast handoff possible at all, and refusing to do it defeats the purpose of the service. A tenant may now be asked for memory if it is idle, whatever its rank, or if it is busy and ranks strictly below the demander. Equal or higher priority is never interrupted, so peers cannot fight. Idle tenants are still preferred over preempting busy ones. Verified against the real contention: with Ollama at 12.38 GB and 98% utilisation, a diffusion job released it within three seconds and completed in 18 s rather than 46 s. Tests updated to the corrected rule, and the fixture's priorities aligned with what actually ships -- it still had the LLM outranking diffusion from before that was swapped. Tests: 246. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
HyperSwap test suite
Fast, hermetic unit tests. No GPU is touched, no network call is made, no systemd unit is
poked, and the production hyperswap.db is never opened.
Running
/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/ -q
Single file / single test:
/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/test_classify_load.py -q
/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/ -q -k warm_confident
Whole suite runs in about 3 seconds.
Safety rails
These matter, because this repo drives a live 4080 SUPER that a running service is using.
tests/conftest.pyinstalls an autouseno_gpu_mutationfixture that replacesoverclock_manager._sh(the single choke point for everynvidia-smi/nvidia-settingswrite) plusapply_profile,apply_fan_control,set_fan_speed,set_fan_autoandrestore_safewith recording stubs. Even a test that accidentally reaches an actuation path can only reach the stub. The fixture yields a dict of recorded calls, which the thermal tests assert against.HYPERSWAP_DBis set to a non-existent path beforetelemetry_storeis imported, so no import can bindDB_PATHto the production database. Tests that need a DB use thetemp_dbfixture, which monkeypatchestelemetry_store.DB_PATHto atmp_pathfile and stops the writer thread afterwards.- All file IO happens against files the tests create in
tmp_path. No real model blob is read andwarm_file_to_ramis never called. - Nothing sweeps, and nothing sends HTTP to Ollama, ComfyUI or :9090.
Measured constants pinned here
These numbers came from measurement on this box, not from taste. If a change makes one of these tests fail, the constant is probably wrong, not the test.
| Constant | Value | Where pinned |
|---|---|---|
| Cold load of a 12.87 GB model, 3.1% resident | 34267 ms → 0.38 GB/s | test_classify_load.py::test_measured_cold_load_classifies_as_cold_disk |
| Warm load of the same model, 100% resident | 4901 ms → 2.63 GB/s | test_classify_load.py::test_measured_warm_load_classifies_as_ram_hit |
RAM_HIT_GBPS = 2.0 must stay below the fastest achievable warm load (2.63 GB/s) |
— | test_classify_load.py::test_ram_hit_threshold_is_physically_achievable |
PARTIAL_HIT_GBPS = 0.8 must stay above the measured cold rate (0.38 GB/s) |
— | same test |
| Size-unknown fallback splits at 8000 ms (between 4.9 s warm and 34.3 s cold) | — | test_classify_load.py::test_unknown_size_guess_boundary_is_8s |
WARM_SKIP_THRESHOLD_PCT = 90.0 |
— | test_ram_optimizer.py::test_warm_skip_threshold_constant_unchanged |
| A probe reading may only be trusted at exactly 100% (a 12-window probe once cleared 90% on a mostly-cold 12.87 GB blob that then loaded at 2.44 GB/s) | — | test_ram_optimizer.py::test_probe_reading_is_only_trusted_at_exactly_100_percent |
PROBE_CACHED_GBPS = 1.5 sits in the gap between cold NVMe (0.35–0.5 GB/s) and page cache (3.2–13 GB/s) |
— | test_ram_optimizer.py::test_probe_cached_threshold_sits_between_measured_disk_and_cache_rates |
| Card power envelope: 320 W stock, 370 W max, sweeps never go below 60% of max | — | test_autotune_helpers.py::test_supported_power_limits_parses_min_default_max |
_supported_clocks must always query the mem,gr pair (a single-field query returns one column and silently yielded []) |
— | test_autotune_helpers.py::test_supported_clocks_always_queries_the_mem_gr_pair |
| ComfyUI benchmark seed must vary per call (a fixed seed made ComfyUI serve a cached result in ~1 ms) | — | test_autotune_helpers.py::test_comfy_workflow_seed_varies_between_calls |
Governor hysteresis: HOT_SAMPLES = 5, COOL_SAMPLES = 30, REAPPLY_COOLDOWN_S = 20 |
— | test_thermal_governor.py (escalation, recovery, cooldown, alternating-sample tests) |
| Model usage score: frequency decayed with a ~24 h half-life | — | test_telemetry_store.py::test_model_usage_ranking_scores_recent_use_higher |
What is deliberately not covered
vram_arbitrator.instant_free_ollama_vram, theAutoArbitratoryield/purge paths and the SSE broker — under active edit, contract changing.overclock_manager.apply_profileand every other actuation path,autotune.sweep,ram_optimizer.warm_file_to_ram— these mutate hardware or do heavy IO.server.pyHTTP routes andmcp_server.py— would need the app wired to live subsystems.
Known rough edge the tests work around
telemetry_store.stop() flushes the writer's pending batch but does not drain the
submission queue, so a stop() racing a just-submitted row can drop it. The writer tests
call a local _drain() helper to wait for the queue to empty before stopping, rather than
encoding the race into an assertion.