Add a dependency self-check; cut the SSE payload by 74%

Self-check. Fan control failed for an entire session -- recoverably, and completely
invisibly. It appeared once, inside one field of one log line, and nothing ever
asked whether fan control worked. health.py now checks everything this service
depends on (NVML, passwordless sudo for nvidia-smi, fan control via the headless X
server, overclock drift, the telemetry store, residency measurement capability,
model directories, the ComfyUI websocket, and both upstream HTTP services) and
reports for each one what is broken, what that breaks, and how to fix it. Exposed at
GET /api/health, as an MCP tool, and as a dashboard panel that collapses to a badge
when healthy and expands to impact-and-fix when not. Current state: 9 ok, 1 degraded
(the known cachestat permission limit on Ollama's blobs).

A self-check that returns ok while a dependency is broken is worse than none, so the
tests drive each check to its failure state -- including the exact "Error resolving
target specification 'gpu:0'" string from the original incident -- and assert that a
check which raises surfaces as failed rather than taking down the endpoint.

SSE payload. The installed-model catalog was 10.6 KB of a 13.1 KB frame, 81% of the
stream, re-sent to every subscriber every second despite changing only when a model
is pulled or removed: 135 MB/hour across three tabs. It is now sent on a
subscriber's first frame and whenever the set changes; the client keeps the last
known list. Steady-state frames dropped from 14041 to 3664 bytes, a 74% reduction,
and /api/stats still returns the complete snapshot for API consumers.

Tests: 182 (was 169).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
drjones
2026-09-05 19:15:52 -07:00
parent 25b601e24a
commit fb104ac9c0
6 changed files with 437 additions and 3 deletions

View File

@@ -16,6 +16,7 @@ from fastapi.middleware.cors import CORSMiddleware
from pydantic import BaseModel, Field
import autotune
import health
import overclock_manager
import ram_optimizer
import telemetry_store
@@ -59,6 +60,7 @@ class TelemetryBroker:
self.running = False
self.samples = 0
self.last_sample_ms = 0.0
self._models_fp = None
# Set on shutdown so open SSE generators finish instead of holding the server up.
self.closing = False
@@ -89,6 +91,26 @@ class TelemetryBroker:
with contextlib.suppress(asyncio.CancelledError):
await self.task
def _stream_frame(self, snap: Dict[str, Any]) -> Dict[str, Any]:
"""Trim the snapshot for streaming.
The installed-model catalog is 10.6 KB of a 13.1 KB payload -- 81% -- and it
changes only when a model is pulled or removed, yet it was re-sent to every
subscriber every second (135 MB/hour across three tabs). It is sent on the first
frame and whenever it changes; otherwise the client keeps what it has.
/api/stats still returns the complete snapshot, so API consumers are unaffected.
"""
ollama = snap.get("ollama", {})
models = ollama.get("installed_models") or []
fp = hash(tuple(sorted(m.get("name", "") for m in models)))
if fp == self._models_fp:
trimmed_ollama = {k: v for k, v in ollama.items() if k != "installed_models"}
trimmed_ollama["installed_models_unchanged"] = True
return {**snap, "ollama": trimmed_ollama}
self._models_fp = fp
return snap
def subscribe(self) -> asyncio.Queue:
q: asyncio.Queue = asyncio.Queue(maxsize=2)
self.subscribers.add(q)
@@ -122,13 +144,14 @@ class TelemetryBroker:
throttle_reasons=",".join(snap.get("gpu", {}).get("throttle_reasons") or []),
)
frame = self._stream_frame(snap)
for q in list(self.subscribers):
if q.full():
# Slow client: drop the stale frame rather than stalling the sampler.
with contextlib.suppress(asyncio.QueueEmpty):
q.get_nowait()
with contextlib.suppress(asyncio.QueueFull):
q.put_nowait(snap)
q.put_nowait(frame)
except asyncio.CancelledError:
raise
except Exception as e:
@@ -288,6 +311,16 @@ async def get_all_stats() -> Dict[str, Any]:
return await broker.get()
@app.get("/api/health", summary="Dependency Self-Check", tags=["Telemetry"])
async def api_health():
"""Check everything HyperSwap depends on, with impact and remediation for each.
Returns overall `status` of ok | degraded | failed. Exists because fan control once
failed for a whole session -- recoverably, and completely silently.
"""
return await health.run_health_checks()
@app.get("/api/gpu", summary="GPU Sensors and VRAM Breakdown", tags=["Telemetry"])
async def get_gpu_metrics() -> Dict[str, Any]:
"""Detailed NVML sensors (utilization, temp, power, fan, clocks, throttle reasons, per-process VRAM)."""
@@ -305,6 +338,7 @@ async def sse_telemetry_stream(request: Request):
q = broker.subscribe()
try:
snap = await broker.get()
# Full snapshot first: a new subscriber has no cached catalog yet.
yield f"data: {json.dumps(snap)}\n\n"
while not broker.closing:
if await request.is_disconnected():