The point of this service is fast handoff of one GPU between applications. It grew up
around the two on this box, and their names ended up compiled into process matching,
VRAM attribution, busy detection and release calls alike -- about 385 references
across five modules. That made it a script for Ollama and ComfyUI rather than a GPU
arbitrator.
tenants.py describes an application as data: how to recognise its processes, how to
tell whether it is genuinely working, how to ask it for VRAM back, and how much it
matters when two want the card. Ollama, ComfyUI and the desktop compositor ship as
defaults in tenants.json, so behaviour is unchanged, but the arbitration logic no
longer knows any particular name. Endpoints are generic: GET /api/tenants,
GET /api/tenants/{name}, POST /api/tenants/{name}/release -- the last being the
general form of both the Ollama soft-yield and the ComfyUI purge.
Verified by registering a third application on this machine with no code change: the
speech relay that had been showing up only as anonymous "unmanaged VRAM" is now named,
attributed, and probed by the VRAM it holds rather than by an API it does not have.
Because it declares no release strategy, a release request returns 409 explaining that
its memory cannot be reclaimed, instead of reporting a success that did nothing.
Busy probes deliberately cannot use GPU utilisation. It is shared by every tenant, so
it cannot attribute work to one of them -- the mistake that made a stale ComfyUI queue
entry undetectable earlier in this branch. A tenant's own VRAM is the signal.
Writing the tests exposed that the suite had become non-hermetic: classification is now
configuration, so a test asserting "a third-party process is unmanaged" started failing
the moment the speech relay was registered on this machine. An autouse fixture now
isolates every test from the operator's live tenants.json.
Tests: 231 (was 206).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
91 lines
3.9 KiB
Python
91 lines
3.9 KiB
Python
"""Shared fixtures and — more importantly — hardware safety rails for the suite.
|
|
|
|
This repo drives a live GPU and a running systemd service. Every test here must be
|
|
hermetic: no NVML mutation, no nvidia-smi/nvidia-settings writes, no touching the
|
|
production telemetry DB, no HTTP to Ollama/ComfyUI/:9090.
|
|
|
|
The `no_gpu_mutation` fixture below is autouse, so even a test that accidentally
|
|
reaches an actuation path can only reach a recording stub.
|
|
"""
|
|
import os
|
|
import sys
|
|
|
|
import pytest
|
|
|
|
REPO_ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
|
|
if REPO_ROOT not in sys.path:
|
|
sys.path.insert(0, REPO_ROOT)
|
|
|
|
# telemetry_store resolves DB_PATH from the environment *at import time*. Point it at a
|
|
# path that does not exist before anything imports it, so no import of this suite can
|
|
# ever open the production hyperswap.db. Individual tests monkeypatch DB_PATH to a
|
|
# tmp_path file when they actually need a database.
|
|
os.environ.setdefault("HYPERSWAP_DB", os.path.join(REPO_ROOT, "tests", "_never_created.db"))
|
|
|
|
import overclock_manager # noqa: E402 (must follow the sys.path/env setup above)
|
|
|
|
|
|
@pytest.fixture(autouse=True)
|
|
def no_gpu_mutation(monkeypatch):
|
|
"""Hard block on every code path that can physically change GPU state.
|
|
|
|
Yields a dict of call recorders so a test can assert that actuation *would* have
|
|
happened without any of it reaching the card.
|
|
"""
|
|
calls = {"apply_profile": [], "fan": [], "restore_safe": [], "sh": []}
|
|
|
|
def _blocked_sh(cmd, use_sudo=True, timeout=10):
|
|
# Catch-all: every nvidia-smi / nvidia-settings write in overclock_manager
|
|
# funnels through _sh. Nothing in the suite may shell out to the driver.
|
|
calls["sh"].append(list(cmd))
|
|
return {"rc": -1, "out": "", "err": "blocked by test suite"}
|
|
|
|
monkeypatch.setattr(overclock_manager, "_sh", _blocked_sh)
|
|
monkeypatch.setattr(overclock_manager, "apply_profile",
|
|
lambda name, overrides=None: calls["apply_profile"].append((name, overrides)))
|
|
monkeypatch.setattr(overclock_manager, "apply_fan_control",
|
|
lambda mode, speed_pct: calls["fan"].append((mode, speed_pct)))
|
|
monkeypatch.setattr(overclock_manager, "set_fan_speed",
|
|
lambda percent: calls["fan"].append(("manual", percent)))
|
|
monkeypatch.setattr(overclock_manager, "set_fan_auto",
|
|
lambda: calls["fan"].append(("auto", None)))
|
|
monkeypatch.setattr(overclock_manager, "restore_safe",
|
|
lambda reason="shutdown": calls["restore_safe"].append(reason))
|
|
return calls
|
|
|
|
|
|
@pytest.fixture
|
|
def temp_db(tmp_path, monkeypatch):
|
|
"""Point telemetry_store at a throwaway SQLite file for the duration of one test."""
|
|
import telemetry_store
|
|
|
|
db = tmp_path / "test_hyperswap.db"
|
|
monkeypatch.setattr(telemetry_store, "DB_PATH", str(db))
|
|
yield str(db)
|
|
# Never leave a writer thread running against a tmp path that is about to vanish.
|
|
try:
|
|
telemetry_store.stop()
|
|
except Exception:
|
|
pass
|
|
|
|
|
|
@pytest.fixture(autouse=True)
|
|
def isolated_tenant_registry(tmp_path, monkeypatch):
|
|
"""Never let tests read the operator's live tenants.json.
|
|
|
|
Classification is now configuration, which means a test that reads the real config
|
|
changes result when someone adds an application to their own machine -- exactly what
|
|
happened when stt-relay was registered and a "third party is unmanaged" test started
|
|
seeing it as a named tenant. Every test gets the shipped defaults unless it opts out
|
|
by pointing CONFIG_PATH somewhere itself.
|
|
"""
|
|
import json as _json
|
|
import tenants as _tenants
|
|
|
|
path = tmp_path / "tenants-default.json"
|
|
path.write_text(_json.dumps(_tenants.DEFAULT_TENANTS))
|
|
monkeypatch.setattr(_tenants, "CONFIG_PATH", str(path))
|
|
_tenants._cache.update({"ts": 0.0, "tenants": None, "mtime": None})
|
|
yield
|
|
_tenants._cache.update({"ts": 0.0, "tenants": None, "mtime": None})
|