Files
gpu-program-swapper/tests
drjones ca97f18be6 Let any tenant declare its GPU profile and event source; fix priority semantics
Two remaining pieces of the two-application coupling are gone.

Overclock profiles were switched by naming 'comfy' and 'ollama' directly, so a third
application could never get tuned clocks. A tenant declares overclock_profile and the
arbitrator applies whichever the highest-priority *working* tenant asks for, falling
back to the idle profile when nothing is running.

The websocket listener parsed ComfyUI's message schema -- status, execution_start,
executing, execution_success -- which tied the fast path to one application. An event
source is now declarative and the messages are not parsed at all: any message means
"look now", and the tenant's own busy probe decides what is true. That gives the same
sub-second reaction to any application that emits anything on state change, with no
knowledge of what it emits.

Generalising this exposed a design error in the priority rule I had introduced.
plan_release excluded candidates ranking above the demander, which broke both
directions in turn. With the LLM at priority 60 and diffusion at 50, ComfyUI could
never reclaim from Ollama -- the premise the whole service is built on, and preserved
until now only by the ComfyUI-specific trigger that was about to be removed. Swapping
the ranks then broke the reverse: a starved Ollama could no longer reclaim from an
idle ComfyUI.

Priority now orders rather than vetoes. Any idle reclaimable tenant is a candidate,
because an idle tenant is not using its VRAM; priority decides who is asked first, and
busy tenants are never interrupted whatever their rank. Diffusion outranks the LLM,
whose weights reload from page cache in seconds. All three cases are pinned by tests,
including that busy work is never interrupted even by a far higher-priority demander.

Tests: 244 (was 242).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 15:23:48 -07:00
..

HyperSwap test suite

Fast, hermetic unit tests. No GPU is touched, no network call is made, no systemd unit is poked, and the production hyperswap.db is never opened.

Running

/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/ -q

Single file / single test:

/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/test_classify_load.py -q
/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/ -q -k warm_confident

Whole suite runs in about 3 seconds.

Safety rails

These matter, because this repo drives a live 4080 SUPER that a running service is using.

  • tests/conftest.py installs an autouse no_gpu_mutation fixture that replaces overclock_manager._sh (the single choke point for every nvidia-smi / nvidia-settings write) plus apply_profile, apply_fan_control, set_fan_speed, set_fan_auto and restore_safe with recording stubs. Even a test that accidentally reaches an actuation path can only reach the stub. The fixture yields a dict of recorded calls, which the thermal tests assert against.
  • HYPERSWAP_DB is set to a non-existent path before telemetry_store is imported, so no import can bind DB_PATH to the production database. Tests that need a DB use the temp_db fixture, which monkeypatches telemetry_store.DB_PATH to a tmp_path file and stops the writer thread afterwards.
  • All file IO happens against files the tests create in tmp_path. No real model blob is read and warm_file_to_ram is never called.
  • Nothing sweeps, and nothing sends HTTP to Ollama, ComfyUI or :9090.

Measured constants pinned here

These numbers came from measurement on this box, not from taste. If a change makes one of these tests fail, the constant is probably wrong, not the test.

Constant Value Where pinned
Cold load of a 12.87 GB model, 3.1% resident 34267 ms → 0.38 GB/s test_classify_load.py::test_measured_cold_load_classifies_as_cold_disk
Warm load of the same model, 100% resident 4901 ms → 2.63 GB/s test_classify_load.py::test_measured_warm_load_classifies_as_ram_hit
RAM_HIT_GBPS = 2.0 must stay below the fastest achievable warm load (2.63 GB/s) — test_classify_load.py::test_ram_hit_threshold_is_physically_achievable
PARTIAL_HIT_GBPS = 0.8 must stay above the measured cold rate (0.38 GB/s) — same test
Size-unknown fallback splits at 8000 ms (between 4.9 s warm and 34.3 s cold) — test_classify_load.py::test_unknown_size_guess_boundary_is_8s
WARM_SKIP_THRESHOLD_PCT = 90.0 — test_ram_optimizer.py::test_warm_skip_threshold_constant_unchanged
A probe reading may only be trusted at exactly 100% (a 12-window probe once cleared 90% on a mostly-cold 12.87 GB blob that then loaded at 2.44 GB/s) — test_ram_optimizer.py::test_probe_reading_is_only_trusted_at_exactly_100_percent
PROBE_CACHED_GBPS = 1.5 sits in the gap between cold NVMe (0.35–0.5 GB/s) and page cache (3.2–13 GB/s) — test_ram_optimizer.py::test_probe_cached_threshold_sits_between_measured_disk_and_cache_rates
Card power envelope: 320 W stock, 370 W max, sweeps never go below 60% of max — test_autotune_helpers.py::test_supported_power_limits_parses_min_default_max
_supported_clocks must always query the mem,gr pair (a single-field query returns one column and silently yielded []) — test_autotune_helpers.py::test_supported_clocks_always_queries_the_mem_gr_pair
ComfyUI benchmark seed must vary per call (a fixed seed made ComfyUI serve a cached result in ~1 ms) — test_autotune_helpers.py::test_comfy_workflow_seed_varies_between_calls
Governor hysteresis: HOT_SAMPLES = 5, COOL_SAMPLES = 30, REAPPLY_COOLDOWN_S = 20 — test_thermal_governor.py (escalation, recovery, cooldown, alternating-sample tests)
Model usage score: frequency decayed with a ~24 h half-life — test_telemetry_store.py::test_model_usage_ranking_scores_recent_use_higher

What is deliberately not covered

  • vram_arbitrator.instant_free_ollama_vram, the AutoArbitrator yield/purge paths and the SSE broker — under active edit, contract changing.
  • overclock_manager.apply_profile and every other actuation path, autotune.sweep, ram_optimizer.warm_file_to_ram — these mutate hardware or do heavy IO.
  • server.py HTTP routes and mcp_server.py — would need the app wired to live subsystems.

Known rough edge the tests work around

telemetry_store.stop() flushes the writer's pending batch but does not drain the submission queue, so a stop() racing a just-submitted row can drop it. The writer tests call a local _drain() helper to wait for the queue to empty before stopping, rather than encoding the race into an assertion.