From 2948b0b4401ce14d2717b469d7df962678e0b379 Mon Sep 17 00:00:00 2001 From: drjones Date: Sat, 5 Sep 2026 19:17:16 -0700 Subject: [PATCH] Detect fan-mode drift; document health checks, VRAM accounting and tests The reconciler only compared the power limit, so a profile whose fan setting never applied stayed wrong indefinitely. The headless X server that owns the GPU can be starting when this unit does; in-process retries cover a short delay, but if X arrives later nothing noticed that the fan mode had never been set. profile_drift() now compares fan mode too, gated on fan control having worked at least once so the check does not fire forever on a machine without it. README documents the self-check endpoint, the ollama/comfy/desktop/unmanaged bucketing and why unreclaimable VRAM is reported separately, the SSE trimming, and a tests section listing the measured constants the suite pins. Co-Authored-By: Claude Opus 5 --- README.md | 45 +++++++++++++++++++++++++++++++++++++++++++- overclock_manager.py | 40 ++++++++++++++++++++++++++++++++------- 2 files changed, 77 insertions(+), 8 deletions(-) diff --git a/README.md b/README.md index 08c121e..9c7649e 100644 --- a/README.md +++ b/README.md @@ -111,6 +111,28 @@ only if the card actually needs it. * **Honest gain reporting**: gain against the profile's *current* setting is reported separately from the spread across values tried. Conflating them turns a flat result into a headline "+102%". * **Safety**: refuses to start while ComfyUI is executing, suspends the arbitrator's automatic profile switching for the duration (otherwise a diffusion benchmark trips `trigger_comfy_priority`, which reapplies the whole profile and overwrites the clock being measured), and restores the original profile in a `finally` block — including on exception or cancellation. +### 🩺 Dependency Self-Check (`health.py`) +* `GET /api/health` verifies **everything this service depends on**: NVML, passwordless + sudo for `nvidia-smi`, fan control through the headless X server, overclock drift, the + telemetry store, residency-measurement capability, model directories, the ComfyUI + WebSocket, and both upstream HTTP services. +* Each check reports **what is broken, what that breaks, and how to fix it** — not just a + red light. Shown on the dashboard as a badge that expands only when something is wrong. +* It exists because fan control once failed for an entire session, recoverably and + silently: the unit started before the X server that owns the GPU was accepting + connections, the assignment failed with `Error resolving target specification 'gpu:0'`, + nothing retried, and nothing ever asked whether fans worked. That failure now shows up + in three places — a retry, a drift check, and this endpoint. + +### 🧮 Honest VRAM Accounting +* Processes are bucketed **ollama / comfy / desktop / unmanaged** rather than into one + catch-all. On this machine a long-running `stt_relay.py` held 842 MB for three days + while the compositor held 3.9 MB; a single "system" number reported them as one figure. +* That distinction matters because ComfyUI's memory **can** be reclaimed and a third + party's **cannot**. `unmanaged_gb` is headroom the arbitrator can never give back, so it + is reported explicitly, shown on the dashboard, and named in the error when a + reclaim-and-retry still cannot fit a model. + ### 🗄️ Persistent Telemetry Store (`telemetry_store.py`) * Swap history used to be an in-memory `deque(maxlen=50)` that evaporated on every restart. Telemetry and events now persist to SQLite (WAL, single writer thread, batched 1 Hz inserts, automatic retention pruning) at roughly **0.4 MB per hour**. * This is what makes the app's central question answerable: **`GET /api/analytics/profiles` compares decode throughput per overclock profile**, joined against the thermals recorded while that profile was active. @@ -119,7 +141,7 @@ only if the card actually needs it. * **Live Hardware Telemetry**: GPU utilization %, GPU temperature (°C), power draw (W), fan speeds (%), and graphics/memory clock frequencies (MHz). * **Live Dual-Axis Time-Series Chart**: Real-time graphical visualization of VRAM usage (GB) and Host RAM Cache (GB) with zero frontend polling overhead. * **Interactive Control Center**: Trigger model hot-swaps, soft-yields, cache pre-warms, fan adjustments, and benchmarks directly from the web interface. -* **Server-Sent Events (SSE)**: A single background sampler produces one 1 Hz snapshot and fans it out to every subscriber via `GET /api/stream`. Previously each connected client independently re-ran the whole snapshot — NVML, `/proc/meminfo`, an HTTP round-trip each to Ollama and ComfyUI, and a recursive walk of the ComfyUI models tree with a `stat()` per checkpoint — once per second, so opening the dashboard in three tabs tripled the load on the thing it was measuring. Slow clients drop stale frames instead of stalling the sampler. +* **Server-Sent Events (SSE)**: A single background sampler produces one 1 Hz snapshot and fans it out to every subscriber via `GET /api/stream`. Frames are trimmed: the installed-model catalog was 81% of a 13.1 KB payload and changes only when a model is pulled, so it is sent on a subscriber's first frame and whenever it changes. Steady-state frames dropped 14041 → 3664 bytes (**74% smaller**; 135 → 38 MB/hour across three tabs), while `/api/stats` still returns the complete snapshot. Previously each connected client independently re-ran the whole snapshot — NVML, `/proc/meminfo`, an HTTP round-trip each to Ollama and ComfyUI, and a recursive walk of the ComfyUI models tree with a `stat()` per checkpoint — once per second, so opening the dashboard in three tabs tripled the load on the thing it was measuring. Slow clients drop stale frames instead of stalling the sampler. ### 🤖 Model Context Protocol (MCP 2.0) Server * **23 Native Agentic Tools**: Allows AI agents (Antigravity CLI, Claude Desktop, Cursor) to manage GPU resources, trigger model hot-swaps, measure page-cache residency, read persisted performance analytics, drive the thermal governor, and run overclock sweeps. @@ -131,6 +153,26 @@ only if the card actually needs it. --- +## 1a. Tests + +```bash +/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/ -q # 182 passed in ~3.7s +``` + +Hermetic: no GPU, no network, no sleeps. An autouse fixture stubs `overclock_manager._sh` +— the single choke point for every `nvidia-smi`/`nvidia-settings` write — so no test can +mutate the card, and `HYPERSWAP_DB` is redirected before `telemetry_store` imports. + +The suite deliberately **pins empirically measured constants**, so that a future edit +which contradicts the hardware fails loudly rather than silently: + +| Pinned fact | Measured value | Why it is pinned | +| :--- | :--- | :--- | +| Warm model load | 12.87 GB in 4901 ms = 2.63 GB/s | The cache-hit threshold must stay below this, or no load can ever qualify | +| Cold model load | 12.87 GB in 34267 ms = 0.38 GB/s | Separates a genuine cold read from a partial hit | +| Busy yield | VRAM held at ≥50% GPU utilisation | A mid-generation model is finishing, not failing | +| Residency confidence | probe trusted only at 100% | A 12-window probe once cleared 90% on a mostly-cold file | + ## 2. Architectural Overview ```mermaid @@ -222,6 +264,7 @@ The HyperSwap server runs on port `9090` by default. Interactive OpenAPI/Swagger | `/api/analytics/models` | `GET` | Recency/frequency model ranking used to prioritise the warm budget. | | `/api/history?durable=true` | `GET` | Swap history from the persistent store rather than the in-memory ring. | | `/api/db` | `GET` | Store location, row counts and how many hours of history are held. | +| `/api/health` | `GET` | Dependency self-check: NVML, sudo, fan control, drift, store, upstreams — each with impact and remediation. | ### Governor & Autotune Endpoints diff --git a/overclock_manager.py b/overclock_manager.py index 05f09ab..4cb67c1 100644 --- a/overclock_manager.py +++ b/overclock_manager.py @@ -308,9 +308,11 @@ def apply_profile(name: str, overrides: Optional[Dict[str, Any]] = None) -> Dict logger.warning(f"Profile '{name}' could not set fans: " f"{result['fan'].get('detail')}") - global _APPLIED_ONCE + global _APPLIED_ONCE, _FAN_AVAILABLE if result["verified"]["power_limit_ok"]: _APPLIED_ONCE = True + if result["verified"]["fan_ok"]: + _FAN_AVAILABLE = True ACTIVE_PROFILE = name _LAST_RESULT = result logger.info(f"Overclock profile applied: {name} -> {json.dumps(result, default=str)}") @@ -475,6 +477,9 @@ def restore_safe(reason: str = "shutdown") -> Dict[str, Any]: _APPLIED_ONCE = False +# Set once fan control has worked at least once, so drift checks do not fire forever on +# a machine that simply has no fan control available. +_FAN_AVAILABLE = False def profile_drift() -> Dict[str, Any]: @@ -490,17 +495,38 @@ def profile_drift() -> Dict[str, Any]: state = get_gpu_state() intended = int(cfg.get("power_limit_w", 0) or 0) actual = state.get("power_limit_w") - drifted = bool(intended and actual is not None - and abs(float(actual) - intended) >= 1.0) + power_drift = bool(intended and actual is not None + and abs(float(actual) - intended) >= 1.0) + + # Fan mode is checked too. The headless X server that owns the GPU can still be + # starting when this unit does, and the fan assignment then fails; the in-process + # retries cover a short delay, but if X arrives later nothing else would ever notice + # that the profile's fan setting was never applied. + fan_intended = cfg.get("fan_mode", "auto") + fan_actual = None + fan_drift = False + if _FAN_AVAILABLE: + fan_actual = get_fan_status().get("mode") + fan_drift = bool(fan_actual and fan_actual != fan_intended) + + drifted = power_drift or fan_drift or not _APPLIED_ONCE + reasons = [] + if not _APPLIED_ONCE: + reasons.append("no profile has been successfully applied since startup") + if power_drift: + reasons.append(f"card reports {actual}W, profile asks {intended}W") + if fan_drift: + reasons.append(f"fans are {fan_actual}, profile asks {fan_intended}") return { "profile": ACTIVE_PROFILE, "applied_since_start": _APPLIED_ONCE, "power_limit_intended_w": intended, "power_limit_actual_w": actual, - "drifted": drifted or not _APPLIED_ONCE, - "reason": ("no profile has been successfully applied since startup" - if not _APPLIED_ONCE else - f"card reports {actual}W, profile asks {intended}W" if drifted else None), + "fan_mode_intended": fan_intended, + "fan_mode_actual": fan_actual, + "fan_available": _FAN_AVAILABLE, + "drifted": drifted, + "reason": "; ".join(reasons) or None, }