The point of this service is fast handoff of one GPU between applications. It grew up
around the two on this box, and their names ended up compiled into process matching,
VRAM attribution, busy detection and release calls alike -- about 385 references
across five modules. That made it a script for Ollama and ComfyUI rather than a GPU
arbitrator.
tenants.py describes an application as data: how to recognise its processes, how to
tell whether it is genuinely working, how to ask it for VRAM back, and how much it
matters when two want the card. Ollama, ComfyUI and the desktop compositor ship as
defaults in tenants.json, so behaviour is unchanged, but the arbitration logic no
longer knows any particular name. Endpoints are generic: GET /api/tenants,
GET /api/tenants/{name}, POST /api/tenants/{name}/release -- the last being the
general form of both the Ollama soft-yield and the ComfyUI purge.
Verified by registering a third application on this machine with no code change: the
speech relay that had been showing up only as anonymous "unmanaged VRAM" is now named,
attributed, and probed by the VRAM it holds rather than by an API it does not have.
Because it declares no release strategy, a release request returns 409 explaining that
its memory cannot be reclaimed, instead of reporting a success that did nothing.
Busy probes deliberately cannot use GPU utilisation. It is shared by every tenant, so
it cannot attribute work to one of them -- the mistake that made a stale ComfyUI queue
entry undetectable earlier in this branch. A tenant's own VRAM is the signal.
Writing the tests exposed that the suite had become non-hermetic: classification is now
configuration, so a test asserting "a third-party process is unmanaged" started failing
the moment the speech relay was registered on this machine. An autouse fixture now
isolates every test from the operator's live tenants.json.
Tests: 231 (was 206).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
HyperSwap test suite
Fast, hermetic unit tests. No GPU is touched, no network call is made, no systemd unit is
poked, and the production hyperswap.db is never opened.
Running
/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/ -q
Single file / single test:
/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/test_classify_load.py -q
/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/ -q -k warm_confident
Whole suite runs in about 3 seconds.
Safety rails
These matter, because this repo drives a live 4080 SUPER that a running service is using.
tests/conftest.pyinstalls an autouseno_gpu_mutationfixture that replacesoverclock_manager._sh(the single choke point for everynvidia-smi/nvidia-settingswrite) plusapply_profile,apply_fan_control,set_fan_speed,set_fan_autoandrestore_safewith recording stubs. Even a test that accidentally reaches an actuation path can only reach the stub. The fixture yields a dict of recorded calls, which the thermal tests assert against.HYPERSWAP_DBis set to a non-existent path beforetelemetry_storeis imported, so no import can bindDB_PATHto the production database. Tests that need a DB use thetemp_dbfixture, which monkeypatchestelemetry_store.DB_PATHto atmp_pathfile and stops the writer thread afterwards.- All file IO happens against files the tests create in
tmp_path. No real model blob is read andwarm_file_to_ramis never called. - Nothing sweeps, and nothing sends HTTP to Ollama, ComfyUI or :9090.
Measured constants pinned here
These numbers came from measurement on this box, not from taste. If a change makes one of these tests fail, the constant is probably wrong, not the test.
| Constant | Value | Where pinned |
|---|---|---|
| Cold load of a 12.87 GB model, 3.1% resident | 34267 ms → 0.38 GB/s | test_classify_load.py::test_measured_cold_load_classifies_as_cold_disk |
| Warm load of the same model, 100% resident | 4901 ms → 2.63 GB/s | test_classify_load.py::test_measured_warm_load_classifies_as_ram_hit |
RAM_HIT_GBPS = 2.0 must stay below the fastest achievable warm load (2.63 GB/s) |
— | test_classify_load.py::test_ram_hit_threshold_is_physically_achievable |
PARTIAL_HIT_GBPS = 0.8 must stay above the measured cold rate (0.38 GB/s) |
— | same test |
| Size-unknown fallback splits at 8000 ms (between 4.9 s warm and 34.3 s cold) | — | test_classify_load.py::test_unknown_size_guess_boundary_is_8s |
WARM_SKIP_THRESHOLD_PCT = 90.0 |
— | test_ram_optimizer.py::test_warm_skip_threshold_constant_unchanged |
| A probe reading may only be trusted at exactly 100% (a 12-window probe once cleared 90% on a mostly-cold 12.87 GB blob that then loaded at 2.44 GB/s) | — | test_ram_optimizer.py::test_probe_reading_is_only_trusted_at_exactly_100_percent |
PROBE_CACHED_GBPS = 1.5 sits in the gap between cold NVMe (0.35–0.5 GB/s) and page cache (3.2–13 GB/s) |
— | test_ram_optimizer.py::test_probe_cached_threshold_sits_between_measured_disk_and_cache_rates |
| Card power envelope: 320 W stock, 370 W max, sweeps never go below 60% of max | — | test_autotune_helpers.py::test_supported_power_limits_parses_min_default_max |
_supported_clocks must always query the mem,gr pair (a single-field query returns one column and silently yielded []) |
— | test_autotune_helpers.py::test_supported_clocks_always_queries_the_mem_gr_pair |
| ComfyUI benchmark seed must vary per call (a fixed seed made ComfyUI serve a cached result in ~1 ms) | — | test_autotune_helpers.py::test_comfy_workflow_seed_varies_between_calls |
Governor hysteresis: HOT_SAMPLES = 5, COOL_SAMPLES = 30, REAPPLY_COOLDOWN_S = 20 |
— | test_thermal_governor.py (escalation, recovery, cooldown, alternating-sample tests) |
| Model usage score: frequency decayed with a ~24 h half-life | — | test_telemetry_store.py::test_model_usage_ranking_scores_recent_use_higher |
What is deliberately not covered
vram_arbitrator.instant_free_ollama_vram, theAutoArbitratoryield/purge paths and the SSE broker — under active edit, contract changing.overclock_manager.apply_profileand every other actuation path,autotune.sweep,ram_optimizer.warm_file_to_ram— these mutate hardware or do heavy IO.server.pyHTTP routes andmcp_server.py— would need the app wired to live subsystems.
Known rough edge the tests work around
telemetry_store.stop() flushes the writer's pending batch but does not drain the
submission queue, so a stop() racing a just-submitted row can drop it. The writer tests
call a local _drain() helper to wait for the queue to empty before stopping, rather than
encoding the race into an assertion.