Add a dependency self-check; cut the SSE payload by 74%
Self-check. Fan control failed for an entire session -- recoverably, and completely invisibly. It appeared once, inside one field of one log line, and nothing ever asked whether fan control worked. health.py now checks everything this service depends on (NVML, passwordless sudo for nvidia-smi, fan control via the headless X server, overclock drift, the telemetry store, residency measurement capability, model directories, the ComfyUI websocket, and both upstream HTTP services) and reports for each one what is broken, what that breaks, and how to fix it. Exposed at GET /api/health, as an MCP tool, and as a dashboard panel that collapses to a badge when healthy and expands to impact-and-fix when not. Current state: 9 ok, 1 degraded (the known cachestat permission limit on Ollama's blobs). A self-check that returns ok while a dependency is broken is worse than none, so the tests drive each check to its failure state -- including the exact "Error resolving target specification 'gpu:0'" string from the original incident -- and assert that a check which raises surfaces as failed rather than taking down the endpoint. SSE payload. The installed-model catalog was 10.6 KB of a 13.1 KB frame, 81% of the stream, re-sent to every subscriber every second despite changing only when a model is pulled or removed: 135 MB/hour across three tabs. It is now sent on a subscriber's first frame and whenever the set changes; the client keeps the last known list. Steady-state frames dropped from 14041 to 3664 bytes, a 74% reduction, and /api/stats still returns the complete snapshot for API consumers. Tests: 182 (was 169). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -8,6 +8,7 @@ from typing import Dict, List, Any, Optional
|
||||
|
||||
from mcp.server import MCPServer
|
||||
import autotune
|
||||
import health
|
||||
import overclock_manager
|
||||
import ram_optimizer
|
||||
import telemetry_store
|
||||
@@ -136,6 +137,14 @@ def set_gpu_fan_speed(mode: str = "auto", percent: Optional[int] = None) -> str:
|
||||
res = overclock_manager.set_fan_auto()
|
||||
return json.dumps(res, indent=2)
|
||||
|
||||
@mcp.tool()
|
||||
async def check_system_health() -> str:
|
||||
"""Check every dependency HyperSwap needs (NVML, sudo nvidia-smi, fan control via the
|
||||
headless X server, Ollama, ComfyUI, the telemetry store, model directories) and report
|
||||
what is broken, what it breaks, and how to fix it."""
|
||||
return json.dumps(await health.run_health_checks(), indent=2, default=str)
|
||||
|
||||
|
||||
@mcp.tool()
|
||||
def get_page_cache_residency(include_files: bool = True) -> str:
|
||||
"""Measure how much of each model on disk is genuinely resident in the Linux page cache.
|
||||
|
||||
Reference in New Issue
Block a user