Detect fan-mode drift; document health checks, VRAM accounting and tests

The reconciler only compared the power limit, so a profile whose fan setting never
applied stayed wrong indefinitely. The headless X server that owns the GPU can be
starting when this unit does; in-process retries cover a short delay, but if X
arrives later nothing noticed that the fan mode had never been set. profile_drift()
now compares fan mode too, gated on fan control having worked at least once so the
check does not fire forever on a machine without it.

README documents the self-check endpoint, the ollama/comfy/desktop/unmanaged
bucketing and why unreclaimable VRAM is reported separately, the SSE trimming, and
a tests section listing the measured constants the suite pins.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
drjones
2026-09-05 19:17:16 -07:00
parent fb104ac9c0
commit 2948b0b440
2 changed files with 77 additions and 8 deletions

View File

@@ -111,6 +111,28 @@ only if the card actually needs it.
* **Honest gain reporting**: gain against the profile's *current* setting is reported separately from the spread across values tried. Conflating them turns a flat result into a headline "+102%".
* **Safety**: refuses to start while ComfyUI is executing, suspends the arbitrator's automatic profile switching for the duration (otherwise a diffusion benchmark trips `trigger_comfy_priority`, which reapplies the whole profile and overwrites the clock being measured), and restores the original profile in a `finally` block — including on exception or cancellation.
### 🩺 Dependency Self-Check (`health.py`)
* `GET /api/health` verifies **everything this service depends on**: NVML, passwordless
sudo for `nvidia-smi`, fan control through the headless X server, overclock drift, the
telemetry store, residency-measurement capability, model directories, the ComfyUI
WebSocket, and both upstream HTTP services.
* Each check reports **what is broken, what that breaks, and how to fix it** — not just a
red light. Shown on the dashboard as a badge that expands only when something is wrong.
* It exists because fan control once failed for an entire session, recoverably and
silently: the unit started before the X server that owns the GPU was accepting
connections, the assignment failed with `Error resolving target specification 'gpu:0'`,
nothing retried, and nothing ever asked whether fans worked. That failure now shows up
in three places — a retry, a drift check, and this endpoint.
### 🧮 Honest VRAM Accounting
* Processes are bucketed **ollama / comfy / desktop / unmanaged** rather than into one
catch-all. On this machine a long-running `stt_relay.py` held 842 MB for three days
while the compositor held 3.9 MB; a single "system" number reported them as one figure.
* That distinction matters because ComfyUI's memory **can** be reclaimed and a third
party's **cannot**. `unmanaged_gb` is headroom the arbitrator can never give back, so it
is reported explicitly, shown on the dashboard, and named in the error when a
reclaim-and-retry still cannot fit a model.
### 🗄️ Persistent Telemetry Store (`telemetry_store.py`)
* Swap history used to be an in-memory `deque(maxlen=50)` that evaporated on every restart. Telemetry and events now persist to SQLite (WAL, single writer thread, batched 1 Hz inserts, automatic retention pruning) at roughly **0.4 MB per hour**.
* This is what makes the app's central question answerable: **`GET /api/analytics/profiles` compares decode throughput per overclock profile**, joined against the thermals recorded while that profile was active.
@@ -119,7 +141,7 @@ only if the card actually needs it.
* **Live Hardware Telemetry**: GPU utilization %, GPU temperature (°C), power draw (W), fan speeds (%), and graphics/memory clock frequencies (MHz).
* **Live Dual-Axis Time-Series Chart**: Real-time graphical visualization of VRAM usage (GB) and Host RAM Cache (GB) with zero frontend polling overhead.
* **Interactive Control Center**: Trigger model hot-swaps, soft-yields, cache pre-warms, fan adjustments, and benchmarks directly from the web interface.
* **Server-Sent Events (SSE)**: A single background sampler produces one 1 Hz snapshot and fans it out to every subscriber via `GET /api/stream`. Previously each connected client independently re-ran the whole snapshot — NVML, `/proc/meminfo`, an HTTP round-trip each to Ollama and ComfyUI, and a recursive walk of the ComfyUI models tree with a `stat()` per checkpoint — once per second, so opening the dashboard in three tabs tripled the load on the thing it was measuring. Slow clients drop stale frames instead of stalling the sampler.
* **Server-Sent Events (SSE)**: A single background sampler produces one 1 Hz snapshot and fans it out to every subscriber via `GET /api/stream`. Frames are trimmed: the installed-model catalog was 81% of a 13.1 KB payload and changes only when a model is pulled, so it is sent on a subscriber's first frame and whenever it changes. Steady-state frames dropped 14041 → 3664 bytes (**74% smaller**; 135 → 38 MB/hour across three tabs), while `/api/stats` still returns the complete snapshot. Previously each connected client independently re-ran the whole snapshot — NVML, `/proc/meminfo`, an HTTP round-trip each to Ollama and ComfyUI, and a recursive walk of the ComfyUI models tree with a `stat()` per checkpoint — once per second, so opening the dashboard in three tabs tripled the load on the thing it was measuring. Slow clients drop stale frames instead of stalling the sampler.
### 🤖 Model Context Protocol (MCP 2.0) Server
* **23 Native Agentic Tools**: Allows AI agents (Antigravity CLI, Claude Desktop, Cursor) to manage GPU resources, trigger model hot-swaps, measure page-cache residency, read persisted performance analytics, drive the thermal governor, and run overclock sweeps.
@@ -131,6 +153,26 @@ only if the card actually needs it.
---
## 1a. Tests
```bash
/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/ -q # 182 passed in ~3.7s
```
Hermetic: no GPU, no network, no sleeps. An autouse fixture stubs `overclock_manager._sh`
— the single choke point for every `nvidia-smi`/`nvidia-settings` write — so no test can
mutate the card, and `HYPERSWAP_DB` is redirected before `telemetry_store` imports.
The suite deliberately **pins empirically measured constants**, so that a future edit
which contradicts the hardware fails loudly rather than silently:
| Pinned fact | Measured value | Why it is pinned |
| :--- | :--- | :--- |
| Warm model load | 12.87 GB in 4901 ms = 2.63 GB/s | The cache-hit threshold must stay below this, or no load can ever qualify |
| Cold model load | 12.87 GB in 34267 ms = 0.38 GB/s | Separates a genuine cold read from a partial hit |
| Busy yield | VRAM held at ≥50% GPU utilisation | A mid-generation model is finishing, not failing |
| Residency confidence | probe trusted only at 100% | A 12-window probe once cleared 90% on a mostly-cold file |
## 2. Architectural Overview
```mermaid
@@ -222,6 +264,7 @@ The HyperSwap server runs on port `9090` by default. Interactive OpenAPI/Swagger
| `/api/analytics/models` | `GET` | Recency/frequency model ranking used to prioritise the warm budget. |
| `/api/history?durable=true` | `GET` | Swap history from the persistent store rather than the in-memory ring. |
| `/api/db` | `GET` | Store location, row counts and how many hours of history are held. |
| `/api/health` | `GET` | Dependency self-check: NVML, sudo, fan control, drift, store, upstreams — each with impact and remediation. |
### Governor & Autotune Endpoints