Add test suite (164 tests); reclaim VRAM from ComfyUI when an LLM will not fit

Tests. First automated coverage for the project: 164 tests, 2.7s, no GPU or network.
An autouse fixture stubs overclock_manager._sh -- the single choke point for every
nvidia-smi/nvidia-settings write -- so no test can mutate the card. They deliberately
pin the empirically measured constants that would otherwise rot silently: the cold and
warm load figures behind the cache-hit thresholds, the warm_confident residency rule,
and the busy/stalled yield split. One test asserts RAM_HIT_GBPS stays at or below the
measured 2.63 GB/s warm load, so the old physically unreachable 5.0 GB/s bar cannot
come back.

Three bugs the suite surfaced, now fixed:
- autotune._subsample(values, 1) divided by zero; the early return only covered
  len(values) <= max_steps.
- telemetry_store.stop() flushed its local pending list but never drained the queue,
  silently losing rows submitted just before a shutdown -- exactly when the last
  events matter.
- ram_optimizer.page_residency's zero-byte short-circuit omitted keys every other
  return path provides, so a 0-byte file was planned for warming.

Reclaim. The README has claimed bidirectional arbitration from the start, but only one
direction was ever automatic. Establishing what actually happens took a controlled test
with the service stopped: with ComfyUI holding 6.83 GB, Ollama does not spill to the CPU
on this box -- it aborts with "cudaMalloc failed: out of memory", because n_gpu_layers is
pinned to 99 and it will not reduce the layer count. So both failure modes are handled:
_check_ollama_starved watches size_vram < size for the default configuration where Ollama
does spill, and switch_ollama_model catches the hard OOM, reclaims VRAM from an idle
ComfyUI and retries once. The request that returned HTTP 500 from Ollama directly now
succeeds through HyperSwap, loading at 3.85 GB/s after reclaiming 6.83 GB.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
drjones
2026-09-01 13:40:30 -07:00
parent bacaf50713
commit 868d82794d
16 changed files with 1972 additions and 7 deletions

69
tests/conftest.py Normal file
View File

@@ -0,0 +1,69 @@
"""Shared fixtures and — more importantly — hardware safety rails for the suite.
This repo drives a live GPU and a running systemd service. Every test here must be
hermetic: no NVML mutation, no nvidia-smi/nvidia-settings writes, no touching the
production telemetry DB, no HTTP to Ollama/ComfyUI/:9090.
The `no_gpu_mutation` fixture below is autouse, so even a test that accidentally
reaches an actuation path can only reach a recording stub.
"""
import os
import sys
import pytest
REPO_ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
if REPO_ROOT not in sys.path:
sys.path.insert(0, REPO_ROOT)
# telemetry_store resolves DB_PATH from the environment *at import time*. Point it at a
# path that does not exist before anything imports it, so no import of this suite can
# ever open the production hyperswap.db. Individual tests monkeypatch DB_PATH to a
# tmp_path file when they actually need a database.
os.environ.setdefault("HYPERSWAP_DB", os.path.join(REPO_ROOT, "tests", "_never_created.db"))
import overclock_manager # noqa: E402 (must follow the sys.path/env setup above)
@pytest.fixture(autouse=True)
def no_gpu_mutation(monkeypatch):
"""Hard block on every code path that can physically change GPU state.
Yields a dict of call recorders so a test can assert that actuation *would* have
happened without any of it reaching the card.
"""
calls = {"apply_profile": [], "fan": [], "restore_safe": [], "sh": []}
def _blocked_sh(cmd, use_sudo=True, timeout=10):
# Catch-all: every nvidia-smi / nvidia-settings write in overclock_manager
# funnels through _sh. Nothing in the suite may shell out to the driver.
calls["sh"].append(list(cmd))
return {"rc": -1, "out": "", "err": "blocked by test suite"}
monkeypatch.setattr(overclock_manager, "_sh", _blocked_sh)
monkeypatch.setattr(overclock_manager, "apply_profile",
lambda name, overrides=None: calls["apply_profile"].append((name, overrides)))
monkeypatch.setattr(overclock_manager, "apply_fan_control",
lambda mode, speed_pct: calls["fan"].append((mode, speed_pct)))
monkeypatch.setattr(overclock_manager, "set_fan_speed",
lambda percent: calls["fan"].append(("manual", percent)))
monkeypatch.setattr(overclock_manager, "set_fan_auto",
lambda: calls["fan"].append(("auto", None)))
monkeypatch.setattr(overclock_manager, "restore_safe",
lambda reason="shutdown": calls["restore_safe"].append(reason))
return calls
@pytest.fixture
def temp_db(tmp_path, monkeypatch):
"""Point telemetry_store at a throwaway SQLite file for the duration of one test."""
import telemetry_store
db = tmp_path / "test_hyperswap.db"
monkeypatch.setattr(telemetry_store, "DB_PATH", str(db))
yield str(db)
# Never leave a writer thread running against a tmp path that is about to vanish.
try:
telemetry_store.stop()
except Exception:
pass