Detect fan-mode drift; document health checks, VRAM accounting and tests

The reconciler only compared the power limit, so a profile whose fan setting never
applied stayed wrong indefinitely. The headless X server that owns the GPU can be
starting when this unit does; in-process retries cover a short delay, but if X
arrives later nothing noticed that the fan mode had never been set. profile_drift()
now compares fan mode too, gated on fan control having worked at least once so the
check does not fire forever on a machine without it.

README documents the self-check endpoint, the ollama/comfy/desktop/unmanaged
bucketing and why unreclaimable VRAM is reported separately, the SSE trimming, and
a tests section listing the measured constants the suite pins.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
drjones
2026-09-05 19:17:16 -07:00
parent fb104ac9c0
commit 2948b0b440
2 changed files with 77 additions and 8 deletions

View File

@@ -111,6 +111,28 @@ only if the card actually needs it.
* **Honest gain reporting**: gain against the profile's *current* setting is reported separately from the spread across values tried. Conflating them turns a flat result into a headline "+102%".
* **Safety**: refuses to start while ComfyUI is executing, suspends the arbitrator's automatic profile switching for the duration (otherwise a diffusion benchmark trips `trigger_comfy_priority`, which reapplies the whole profile and overwrites the clock being measured), and restores the original profile in a `finally` block — including on exception or cancellation.
### 🩺 Dependency Self-Check (`health.py`)
* `GET /api/health` verifies **everything this service depends on**: NVML, passwordless
sudo for `nvidia-smi`, fan control through the headless X server, overclock drift, the
telemetry store, residency-measurement capability, model directories, the ComfyUI
WebSocket, and both upstream HTTP services.
* Each check reports **what is broken, what that breaks, and how to fix it** — not just a
red light. Shown on the dashboard as a badge that expands only when something is wrong.
* It exists because fan control once failed for an entire session, recoverably and
silently: the unit started before the X server that owns the GPU was accepting
connections, the assignment failed with `Error resolving target specification 'gpu:0'`,
nothing retried, and nothing ever asked whether fans worked. That failure now shows up
in three places — a retry, a drift check, and this endpoint.
### 🧮 Honest VRAM Accounting
* Processes are bucketed **ollama / comfy / desktop / unmanaged** rather than into one
catch-all. On this machine a long-running `stt_relay.py` held 842 MB for three days
while the compositor held 3.9 MB; a single "system" number reported them as one figure.
* That distinction matters because ComfyUI's memory **can** be reclaimed and a third
party's **cannot**. `unmanaged_gb` is headroom the arbitrator can never give back, so it
is reported explicitly, shown on the dashboard, and named in the error when a
reclaim-and-retry still cannot fit a model.
### 🗄️ Persistent Telemetry Store (`telemetry_store.py`)
* Swap history used to be an in-memory `deque(maxlen=50)` that evaporated on every restart. Telemetry and events now persist to SQLite (WAL, single writer thread, batched 1 Hz inserts, automatic retention pruning) at roughly **0.4 MB per hour**.
* This is what makes the app's central question answerable: **`GET /api/analytics/profiles` compares decode throughput per overclock profile**, joined against the thermals recorded while that profile was active.
@@ -119,7 +141,7 @@ only if the card actually needs it.
* **Live Hardware Telemetry**: GPU utilization %, GPU temperature (°C), power draw (W), fan speeds (%), and graphics/memory clock frequencies (MHz).
* **Live Dual-Axis Time-Series Chart**: Real-time graphical visualization of VRAM usage (GB) and Host RAM Cache (GB) with zero frontend polling overhead.
* **Interactive Control Center**: Trigger model hot-swaps, soft-yields, cache pre-warms, fan adjustments, and benchmarks directly from the web interface.
* **Server-Sent Events (SSE)**: A single background sampler produces one 1 Hz snapshot and fans it out to every subscriber via `GET /api/stream`. Previously each connected client independently re-ran the whole snapshot — NVML, `/proc/meminfo`, an HTTP round-trip each to Ollama and ComfyUI, and a recursive walk of the ComfyUI models tree with a `stat()` per checkpoint — once per second, so opening the dashboard in three tabs tripled the load on the thing it was measuring. Slow clients drop stale frames instead of stalling the sampler.
* **Server-Sent Events (SSE)**: A single background sampler produces one 1 Hz snapshot and fans it out to every subscriber via `GET /api/stream`. Frames are trimmed: the installed-model catalog was 81% of a 13.1 KB payload and changes only when a model is pulled, so it is sent on a subscriber's first frame and whenever it changes. Steady-state frames dropped 14041 → 3664 bytes (**74% smaller**; 135 → 38 MB/hour across three tabs), while `/api/stats` still returns the complete snapshot. Previously each connected client independently re-ran the whole snapshot — NVML, `/proc/meminfo`, an HTTP round-trip each to Ollama and ComfyUI, and a recursive walk of the ComfyUI models tree with a `stat()` per checkpoint — once per second, so opening the dashboard in three tabs tripled the load on the thing it was measuring. Slow clients drop stale frames instead of stalling the sampler.
### 🤖 Model Context Protocol (MCP 2.0) Server
* **23 Native Agentic Tools**: Allows AI agents (Antigravity CLI, Claude Desktop, Cursor) to manage GPU resources, trigger model hot-swaps, measure page-cache residency, read persisted performance analytics, drive the thermal governor, and run overclock sweeps.
@@ -131,6 +153,26 @@ only if the card actually needs it.
---
## 1a. Tests
```bash
/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/ -q # 182 passed in ~3.7s
```
Hermetic: no GPU, no network, no sleeps. An autouse fixture stubs `overclock_manager._sh`
— the single choke point for every `nvidia-smi`/`nvidia-settings` write — so no test can
mutate the card, and `HYPERSWAP_DB` is redirected before `telemetry_store` imports.
The suite deliberately **pins empirically measured constants**, so that a future edit
which contradicts the hardware fails loudly rather than silently:
| Pinned fact | Measured value | Why it is pinned |
| :--- | :--- | :--- |
| Warm model load | 12.87 GB in 4901 ms = 2.63 GB/s | The cache-hit threshold must stay below this, or no load can ever qualify |
| Cold model load | 12.87 GB in 34267 ms = 0.38 GB/s | Separates a genuine cold read from a partial hit |
| Busy yield | VRAM held at ≥50% GPU utilisation | A mid-generation model is finishing, not failing |
| Residency confidence | probe trusted only at 100% | A 12-window probe once cleared 90% on a mostly-cold file |
## 2. Architectural Overview
```mermaid
@@ -222,6 +264,7 @@ The HyperSwap server runs on port `9090` by default. Interactive OpenAPI/Swagger
| `/api/analytics/models` | `GET` | Recency/frequency model ranking used to prioritise the warm budget. |
| `/api/history?durable=true` | `GET` | Swap history from the persistent store rather than the in-memory ring. |
| `/api/db` | `GET` | Store location, row counts and how many hours of history are held. |
| `/api/health` | `GET` | Dependency self-check: NVML, sudo, fan control, drift, store, upstreams — each with impact and remediation. |
### Governor & Autotune Endpoints

View File

@@ -308,9 +308,11 @@ def apply_profile(name: str, overrides: Optional[Dict[str, Any]] = None) -> Dict
logger.warning(f"Profile '{name}' could not set fans: "
f"{result['fan'].get('detail')}")
global _APPLIED_ONCE
global _APPLIED_ONCE, _FAN_AVAILABLE
if result["verified"]["power_limit_ok"]:
_APPLIED_ONCE = True
if result["verified"]["fan_ok"]:
_FAN_AVAILABLE = True
ACTIVE_PROFILE = name
_LAST_RESULT = result
logger.info(f"Overclock profile applied: {name} -> {json.dumps(result, default=str)}")
@@ -475,6 +477,9 @@ def restore_safe(reason: str = "shutdown") -> Dict[str, Any]:
_APPLIED_ONCE = False
# Set once fan control has worked at least once, so drift checks do not fire forever on
# a machine that simply has no fan control available.
_FAN_AVAILABLE = False
def profile_drift() -> Dict[str, Any]:
@@ -490,17 +495,38 @@ def profile_drift() -> Dict[str, Any]:
state = get_gpu_state()
intended = int(cfg.get("power_limit_w", 0) or 0)
actual = state.get("power_limit_w")
drifted = bool(intended and actual is not None
power_drift = bool(intended and actual is not None
and abs(float(actual) - intended) >= 1.0)
# Fan mode is checked too. The headless X server that owns the GPU can still be
# starting when this unit does, and the fan assignment then fails; the in-process
# retries cover a short delay, but if X arrives later nothing else would ever notice
# that the profile's fan setting was never applied.
fan_intended = cfg.get("fan_mode", "auto")
fan_actual = None
fan_drift = False
if _FAN_AVAILABLE:
fan_actual = get_fan_status().get("mode")
fan_drift = bool(fan_actual and fan_actual != fan_intended)
drifted = power_drift or fan_drift or not _APPLIED_ONCE
reasons = []
if not _APPLIED_ONCE:
reasons.append("no profile has been successfully applied since startup")
if power_drift:
reasons.append(f"card reports {actual}W, profile asks {intended}W")
if fan_drift:
reasons.append(f"fans are {fan_actual}, profile asks {fan_intended}")
return {
"profile": ACTIVE_PROFILE,
"applied_since_start": _APPLIED_ONCE,
"power_limit_intended_w": intended,
"power_limit_actual_w": actual,
"drifted": drifted or not _APPLIED_ONCE,
"reason": ("no profile has been successfully applied since startup"
if not _APPLIED_ONCE else
f"card reports {actual}W, profile asks {intended}W" if drifted else None),
"fan_mode_intended": fan_intended,
"fan_mode_actual": fan_actual,
"fan_available": _FAN_AVAILABLE,
"drifted": drifted,
"reason": "; ".join(reasons) or None,
}