Detect fan-mode drift; document health checks, VRAM accounting and tests
The reconciler only compared the power limit, so a profile whose fan setting never applied stayed wrong indefinitely. The headless X server that owns the GPU can be starting when this unit does; in-process retries cover a short delay, but if X arrives later nothing noticed that the fan mode had never been set. profile_drift() now compares fan mode too, gated on fan control having worked at least once so the check does not fire forever on a machine without it. README documents the self-check endpoint, the ollama/comfy/desktop/unmanaged bucketing and why unreclaimable VRAM is reported separately, the SSE trimming, and a tests section listing the measured constants the suite pins. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
45
README.md
45
README.md
@@ -111,6 +111,28 @@ only if the card actually needs it.
|
||||
* **Honest gain reporting**: gain against the profile's *current* setting is reported separately from the spread across values tried. Conflating them turns a flat result into a headline "+102%".
|
||||
* **Safety**: refuses to start while ComfyUI is executing, suspends the arbitrator's automatic profile switching for the duration (otherwise a diffusion benchmark trips `trigger_comfy_priority`, which reapplies the whole profile and overwrites the clock being measured), and restores the original profile in a `finally` block — including on exception or cancellation.
|
||||
|
||||
### 🩺 Dependency Self-Check (`health.py`)
|
||||
* `GET /api/health` verifies **everything this service depends on**: NVML, passwordless
|
||||
sudo for `nvidia-smi`, fan control through the headless X server, overclock drift, the
|
||||
telemetry store, residency-measurement capability, model directories, the ComfyUI
|
||||
WebSocket, and both upstream HTTP services.
|
||||
* Each check reports **what is broken, what that breaks, and how to fix it** — not just a
|
||||
red light. Shown on the dashboard as a badge that expands only when something is wrong.
|
||||
* It exists because fan control once failed for an entire session, recoverably and
|
||||
silently: the unit started before the X server that owns the GPU was accepting
|
||||
connections, the assignment failed with `Error resolving target specification 'gpu:0'`,
|
||||
nothing retried, and nothing ever asked whether fans worked. That failure now shows up
|
||||
in three places — a retry, a drift check, and this endpoint.
|
||||
|
||||
### 🧮 Honest VRAM Accounting
|
||||
* Processes are bucketed **ollama / comfy / desktop / unmanaged** rather than into one
|
||||
catch-all. On this machine a long-running `stt_relay.py` held 842 MB for three days
|
||||
while the compositor held 3.9 MB; a single "system" number reported them as one figure.
|
||||
* That distinction matters because ComfyUI's memory **can** be reclaimed and a third
|
||||
party's **cannot**. `unmanaged_gb` is headroom the arbitrator can never give back, so it
|
||||
is reported explicitly, shown on the dashboard, and named in the error when a
|
||||
reclaim-and-retry still cannot fit a model.
|
||||
|
||||
### 🗄️ Persistent Telemetry Store (`telemetry_store.py`)
|
||||
* Swap history used to be an in-memory `deque(maxlen=50)` that evaporated on every restart. Telemetry and events now persist to SQLite (WAL, single writer thread, batched 1 Hz inserts, automatic retention pruning) at roughly **0.4 MB per hour**.
|
||||
* This is what makes the app's central question answerable: **`GET /api/analytics/profiles` compares decode throughput per overclock profile**, joined against the thermals recorded while that profile was active.
|
||||
@@ -119,7 +141,7 @@ only if the card actually needs it.
|
||||
* **Live Hardware Telemetry**: GPU utilization %, GPU temperature (°C), power draw (W), fan speeds (%), and graphics/memory clock frequencies (MHz).
|
||||
* **Live Dual-Axis Time-Series Chart**: Real-time graphical visualization of VRAM usage (GB) and Host RAM Cache (GB) with zero frontend polling overhead.
|
||||
* **Interactive Control Center**: Trigger model hot-swaps, soft-yields, cache pre-warms, fan adjustments, and benchmarks directly from the web interface.
|
||||
* **Server-Sent Events (SSE)**: A single background sampler produces one 1 Hz snapshot and fans it out to every subscriber via `GET /api/stream`. Previously each connected client independently re-ran the whole snapshot — NVML, `/proc/meminfo`, an HTTP round-trip each to Ollama and ComfyUI, and a recursive walk of the ComfyUI models tree with a `stat()` per checkpoint — once per second, so opening the dashboard in three tabs tripled the load on the thing it was measuring. Slow clients drop stale frames instead of stalling the sampler.
|
||||
* **Server-Sent Events (SSE)**: A single background sampler produces one 1 Hz snapshot and fans it out to every subscriber via `GET /api/stream`. Frames are trimmed: the installed-model catalog was 81% of a 13.1 KB payload and changes only when a model is pulled, so it is sent on a subscriber's first frame and whenever it changes. Steady-state frames dropped 14041 → 3664 bytes (**74% smaller**; 135 → 38 MB/hour across three tabs), while `/api/stats` still returns the complete snapshot. Previously each connected client independently re-ran the whole snapshot — NVML, `/proc/meminfo`, an HTTP round-trip each to Ollama and ComfyUI, and a recursive walk of the ComfyUI models tree with a `stat()` per checkpoint — once per second, so opening the dashboard in three tabs tripled the load on the thing it was measuring. Slow clients drop stale frames instead of stalling the sampler.
|
||||
|
||||
### 🤖 Model Context Protocol (MCP 2.0) Server
|
||||
* **23 Native Agentic Tools**: Allows AI agents (Antigravity CLI, Claude Desktop, Cursor) to manage GPU resources, trigger model hot-swaps, measure page-cache residency, read persisted performance analytics, drive the thermal governor, and run overclock sweeps.
|
||||
@@ -131,6 +153,26 @@ only if the card actually needs it.
|
||||
|
||||
---
|
||||
|
||||
## 1a. Tests
|
||||
|
||||
```bash
|
||||
/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/ -q # 182 passed in ~3.7s
|
||||
```
|
||||
|
||||
Hermetic: no GPU, no network, no sleeps. An autouse fixture stubs `overclock_manager._sh`
|
||||
— the single choke point for every `nvidia-smi`/`nvidia-settings` write — so no test can
|
||||
mutate the card, and `HYPERSWAP_DB` is redirected before `telemetry_store` imports.
|
||||
|
||||
The suite deliberately **pins empirically measured constants**, so that a future edit
|
||||
which contradicts the hardware fails loudly rather than silently:
|
||||
|
||||
| Pinned fact | Measured value | Why it is pinned |
|
||||
| :--- | :--- | :--- |
|
||||
| Warm model load | 12.87 GB in 4901 ms = 2.63 GB/s | The cache-hit threshold must stay below this, or no load can ever qualify |
|
||||
| Cold model load | 12.87 GB in 34267 ms = 0.38 GB/s | Separates a genuine cold read from a partial hit |
|
||||
| Busy yield | VRAM held at ≥50% GPU utilisation | A mid-generation model is finishing, not failing |
|
||||
| Residency confidence | probe trusted only at 100% | A 12-window probe once cleared 90% on a mostly-cold file |
|
||||
|
||||
## 2. Architectural Overview
|
||||
|
||||
```mermaid
|
||||
@@ -222,6 +264,7 @@ The HyperSwap server runs on port `9090` by default. Interactive OpenAPI/Swagger
|
||||
| `/api/analytics/models` | `GET` | Recency/frequency model ranking used to prioritise the warm budget. |
|
||||
| `/api/history?durable=true` | `GET` | Swap history from the persistent store rather than the in-memory ring. |
|
||||
| `/api/db` | `GET` | Store location, row counts and how many hours of history are held. |
|
||||
| `/api/health` | `GET` | Dependency self-check: NVML, sudo, fan control, drift, store, upstreams — each with impact and remediation. |
|
||||
|
||||
### Governor & Autotune Endpoints
|
||||
|
||||
|
||||
@@ -308,9 +308,11 @@ def apply_profile(name: str, overrides: Optional[Dict[str, Any]] = None) -> Dict
|
||||
logger.warning(f"Profile '{name}' could not set fans: "
|
||||
f"{result['fan'].get('detail')}")
|
||||
|
||||
global _APPLIED_ONCE
|
||||
global _APPLIED_ONCE, _FAN_AVAILABLE
|
||||
if result["verified"]["power_limit_ok"]:
|
||||
_APPLIED_ONCE = True
|
||||
if result["verified"]["fan_ok"]:
|
||||
_FAN_AVAILABLE = True
|
||||
ACTIVE_PROFILE = name
|
||||
_LAST_RESULT = result
|
||||
logger.info(f"Overclock profile applied: {name} -> {json.dumps(result, default=str)}")
|
||||
@@ -475,6 +477,9 @@ def restore_safe(reason: str = "shutdown") -> Dict[str, Any]:
|
||||
|
||||
|
||||
_APPLIED_ONCE = False
|
||||
# Set once fan control has worked at least once, so drift checks do not fire forever on
|
||||
# a machine that simply has no fan control available.
|
||||
_FAN_AVAILABLE = False
|
||||
|
||||
|
||||
def profile_drift() -> Dict[str, Any]:
|
||||
@@ -490,17 +495,38 @@ def profile_drift() -> Dict[str, Any]:
|
||||
state = get_gpu_state()
|
||||
intended = int(cfg.get("power_limit_w", 0) or 0)
|
||||
actual = state.get("power_limit_w")
|
||||
drifted = bool(intended and actual is not None
|
||||
and abs(float(actual) - intended) >= 1.0)
|
||||
power_drift = bool(intended and actual is not None
|
||||
and abs(float(actual) - intended) >= 1.0)
|
||||
|
||||
# Fan mode is checked too. The headless X server that owns the GPU can still be
|
||||
# starting when this unit does, and the fan assignment then fails; the in-process
|
||||
# retries cover a short delay, but if X arrives later nothing else would ever notice
|
||||
# that the profile's fan setting was never applied.
|
||||
fan_intended = cfg.get("fan_mode", "auto")
|
||||
fan_actual = None
|
||||
fan_drift = False
|
||||
if _FAN_AVAILABLE:
|
||||
fan_actual = get_fan_status().get("mode")
|
||||
fan_drift = bool(fan_actual and fan_actual != fan_intended)
|
||||
|
||||
drifted = power_drift or fan_drift or not _APPLIED_ONCE
|
||||
reasons = []
|
||||
if not _APPLIED_ONCE:
|
||||
reasons.append("no profile has been successfully applied since startup")
|
||||
if power_drift:
|
||||
reasons.append(f"card reports {actual}W, profile asks {intended}W")
|
||||
if fan_drift:
|
||||
reasons.append(f"fans are {fan_actual}, profile asks {fan_intended}")
|
||||
return {
|
||||
"profile": ACTIVE_PROFILE,
|
||||
"applied_since_start": _APPLIED_ONCE,
|
||||
"power_limit_intended_w": intended,
|
||||
"power_limit_actual_w": actual,
|
||||
"drifted": drifted or not _APPLIED_ONCE,
|
||||
"reason": ("no profile has been successfully applied since startup"
|
||||
if not _APPLIED_ONCE else
|
||||
f"card reports {actual}W, profile asks {intended}W" if drifted else None),
|
||||
"fan_mode_intended": fan_intended,
|
||||
"fan_mode_actual": fan_actual,
|
||||
"fan_available": _FAN_AVAILABLE,
|
||||
"drifted": drifted,
|
||||
"reason": "; ".join(reasons) or None,
|
||||
}
|
||||
|
||||
|
||||
|
||||
Reference in New Issue
Block a user