Read engine configuration live instead of asserting it in the dashboard

The engine subtitles were hardcoded: "FlashAttention + Q4 KV Cache" and "DynamicVRAM
+ Pinned Async Offload". The first turned out to be accurate -- OLLAMA_FLASH_ATTENTION
and OLLAMA_KV_CACHE_TYPE really are set -- which is worse than being wrong, because it
would have gone on looking accurate after the settings changed.

engines.py reads both engines' real configuration: the ollama service environment via
systemd, and ComfyUI's own /system_stats for version, allocator, VRAM mode and argv.
Exposed at GET /api/engines, as an MCP tool, and in the dashboard subtitles with the
full settings list as a tooltip.

The settings worth surfacing are the ones that dictate how this service must behave
and that previously had to be discovered by reading journald: OLLAMA_NUM_PARALLEL=1
is why an unload queues behind a running generation and is reported as deferred
rather than failed, and OLLAMA_MAX_LOADED_MODELS=1 is why every swap evicts the
previous model. Each is reported with that explanation attached.

Writing the tests found a bug in the new code: (system.get("python_version") or
"").split()[0] raises IndexError when ComfyUI omits the field, and the surrounding
except would have swallowed it and reported ComfyUI as entirely offline.

Tests: 192 (was 182).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
drjones
2026-09-06 10:56:04 -07:00
parent b53b026ced
commit 043d61722b
7 changed files with 315 additions and 3 deletions

View File

@@ -111,6 +111,18 @@ only if the card actually needs it.
* **Honest gain reporting**: gain against the profile's *current* setting is reported separately from the spread across values tried. Conflating them turns a flat result into a headline "+102%". * **Honest gain reporting**: gain against the profile's *current* setting is reported separately from the spread across values tried. Conflating them turns a flat result into a headline "+102%".
* **Safety**: refuses to start while ComfyUI is executing, suspends the arbitrator's automatic profile switching for the duration (otherwise a diffusion benchmark trips `trigger_comfy_priority`, which reapplies the whole profile and overwrites the clock being measured), and restores the original profile in a `finally` block — including on exception or cancellation. * **Safety**: refuses to start while ComfyUI is executing, suspends the arbitrator's automatic profile switching for the duration (otherwise a diffusion benchmark trips `trigger_comfy_priority`, which reapplies the whole profile and overwrites the clock being measured), and restores the original profile in a `finally` block — including on exception or cancellation.
### 🔧 Live Engine Configuration (`engines.py`)
* `GET /api/engines` reports the **real, current** configuration of both engines and what
each setting implies for arbitration — because the settings that dictate this service's
behaviour live outside its own codebase.
* `OLLAMA_NUM_PARALLEL=1` is why a `keep_alive: 0` unload queues behind a running
generation and is reported as *deferred* rather than failed. `OLLAMA_MAX_LOADED_MODELS=1`
is why every swap evicts the previous model. Working these out originally meant reading
journald and the systemd unit by hand.
* The dashboard's engine subtitles now come from this endpoint. They were previously
hardcoded — and happened to be accurate, which is worse than being wrong, since they
would have kept looking accurate after the configuration changed.
### 🩺 Dependency Self-Check (`health.py`) ### 🩺 Dependency Self-Check (`health.py`)
* `GET /api/health` verifies **everything this service depends on**: NVML, passwordless * `GET /api/health` verifies **everything this service depends on**: NVML, passwordless
sudo for `nvidia-smi`, fan control through the headless X server, overclock drift, the sudo for `nvidia-smi`, fan control through the headless X server, overclock drift, the
@@ -156,7 +168,7 @@ only if the card actually needs it.
## 1a. Tests ## 1a. Tests
```bash ```bash
/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/ -q # 182 passed in ~3.7s /home/drjones/comfy-mcp-venv/bin/python -m pytest tests/ -q # 192 passed in ~3.7s
``` ```
Hermetic: no GPU, no network, no sleeps. An autouse fixture stubs `overclock_manager._sh` Hermetic: no GPU, no network, no sleeps. An autouse fixture stubs `overclock_manager._sh`
@@ -264,6 +276,7 @@ The HyperSwap server runs on port `9090` by default. Interactive OpenAPI/Swagger
| `/api/analytics/models` | `GET` | Recency/frequency model ranking used to prioritise the warm budget. | | `/api/analytics/models` | `GET` | Recency/frequency model ranking used to prioritise the warm budget. |
| `/api/history?durable=true` | `GET` | Swap history from the persistent store rather than the in-memory ring. | | `/api/history?durable=true` | `GET` | Swap history from the persistent store rather than the in-memory ring. |
| `/api/db` | `GET` | Store location, row counts and how many hours of history are held. | | `/api/db` | `GET` | Store location, row counts and how many hours of history are held. |
| `/api/engines` | `GET` | Live Ollama and ComfyUI configuration, with what each setting implies for arbitration. |
| `/api/health` | `GET` | Dependency self-check: NVML, sudo, fan control, drift, store, upstreams — each with impact and remediation. | | `/api/health` | `GET` | Dependency self-check: NVML, sudo, fan control, drift, store, upstreams — each with impact and remediation. |
### Governor & Autotune Endpoints ### Governor & Autotune Endpoints

117
engines.py Normal file
View File

@@ -0,0 +1,117 @@
"""Live configuration of the two engines HyperSwap arbitrates between.
Arbitration behaviour is largely dictated by settings that live outside this codebase.
Working out why a yield behaved the way it did meant reading journald and the ollama
unit by hand: OLLAMA_NUM_PARALLEL decides whether an unload queues behind a running
generation, OLLAMA_MAX_LOADED_MODELS decides whether more than one model can be
resident, and a pinned n_gpu_layers decides whether a model that will not fit spills to
the CPU or fails outright. Those are worth reading and explaining rather than hardcoding
into a dashboard subtitle that silently goes stale.
"""
import json
import logging
import subprocess
from typing import Any, Dict, List, Optional
import httpx
import vram_arbitrator
logger = logging.getLogger("engines")
# Settings that change how the arbitrator must behave, with what they imply.
OLLAMA_SETTING_NOTES = {
"OLLAMA_NUM_PARALLEL": (
"Requests per model. At 1, a keep_alive:0 unload queues behind any running "
"generation and applies when it finishes — which is why a busy model is "
"reported as deferred rather than failed."),
"OLLAMA_MAX_LOADED_MODELS": (
"How many models may be resident at once. At 1, Ollama evicts the previous "
"model on every swap."),
"OLLAMA_KEEP_ALIVE": (
"Default residency after a request. Long values keep VRAM occupied and make "
"ComfyUI wait for an explicit yield."),
"OLLAMA_FLASH_ATTENTION": "FlashAttention kernels for attention.",
"OLLAMA_KV_CACHE_TYPE": "KV cache quantisation; smaller types cut VRAM per context.",
"OLLAMA_NUM_BATCH": "Prompt-evaluation batch size.",
}
def _ollama_unit_environment() -> Dict[str, str]:
"""Read the ollama service's environment. Empty if it is not a systemd unit."""
env: Dict[str, str] = {}
try:
proc = subprocess.run(["systemctl", "show", "ollama", "-p", "Environment",
"--value"], capture_output=True, text=True, timeout=8)
for token in proc.stdout.split():
if "=" in token and token.startswith("OLLAMA"):
k, _, v = token.partition("=")
env[k] = v
except Exception as e:
logger.debug(f"could not read ollama unit environment: {e}")
return env
async def get_engine_config() -> Dict[str, Any]:
"""Real, live configuration of both engines, with arbitration implications."""
ollama_env = _ollama_unit_environment()
ollama_settings = [
{"key": k, "value": v, "means": OLLAMA_SETTING_NOTES.get(k, "")}
for k, v in sorted(ollama_env.items())
]
# A short, honest summary line to replace the dashboard's hardcoded subtitle.
feature_bits: List[str] = []
if ollama_env.get("OLLAMA_FLASH_ATTENTION") == "1":
feature_bits.append("FlashAttention")
kv = ollama_env.get("OLLAMA_KV_CACHE_TYPE")
if kv:
feature_bits.append(f"{kv} KV cache")
host = ollama_env.get("OLLAMA_HOST", "")
port = host.rsplit(":", 1)[-1] if ":" in host else "11434"
ollama = {
"port": port,
"settings": ollama_settings,
"summary": " + ".join(feature_bits) if feature_bits else "default configuration",
"max_loaded_models": ollama_env.get("OLLAMA_MAX_LOADED_MODELS"),
"num_parallel": ollama_env.get("OLLAMA_NUM_PARALLEL"),
"keep_alive": ollama_env.get("OLLAMA_KEEP_ALIVE"),
"config_source": "systemd unit environment" if ollama_env else "unavailable",
}
comfy: Dict[str, Any] = {"online": False}
try:
async with httpx.AsyncClient(timeout=4.0) as c:
r = await c.get(f"{vram_arbitrator.COMFY_API_BASE}/system_stats")
if r.status_code == 200:
data = r.json()
system = data.get("system", {})
argv = system.get("argv") or []
devices = data.get("devices") or []
dev = devices[0] if devices else {}
# The allocator is named in the device string; it is the closest thing
# ComfyUI reports to the "async offload" the old subtitle asserted.
dev_name = dev.get("name", "")
allocator = ("cudaMallocAsync" if "cudaMallocAsync" in dev_name
else "cudaMalloc" if "cudaMalloc" in dev_name else "unknown")
vram_flags = [a for a in argv
if a in ("--lowvram", "--novram", "--highvram", "--normalvram",
"--gpu-only", "--cpu")]
comfy = {
"online": True,
"version": system.get("comfyui_version"),
"pytorch": system.get("pytorch_version"),
# split()[0] on an absent version raises IndexError, which would have
# been swallowed by the except below and reported ComfyUI as offline.
"python": ((system.get("python_version") or "").split() or [None])[0],
"argv": argv,
"vram_mode": vram_flags[0] if vram_flags else "default (auto)",
"allocator": allocator,
"device": dev_name,
"summary": f"{allocator}, {vram_flags[0] if vram_flags else 'auto VRAM'}",
}
except Exception as e:
comfy = {"online": False, "error": str(e)[:120]}
return {"ollama": ollama, "comfyui": comfy}

View File

@@ -8,6 +8,7 @@ from typing import Dict, List, Any, Optional
from mcp.server import MCPServer from mcp.server import MCPServer
import autotune import autotune
import engines
import health import health
import overclock_manager import overclock_manager
import ram_optimizer import ram_optimizer
@@ -137,6 +138,14 @@ def set_gpu_fan_speed(mode: str = "auto", percent: Optional[int] = None) -> str:
res = overclock_manager.set_fan_auto() res = overclock_manager.set_fan_auto()
return json.dumps(res, indent=2) return json.dumps(res, indent=2)
@mcp.tool()
async def get_engine_config() -> str:
"""Live configuration of Ollama and ComfyUI (parallelism, max loaded models,
keep-alive, KV cache type, ComfyUI VRAM mode and allocator), with what each setting
implies for VRAM arbitration."""
return json.dumps(await engines.get_engine_config(), indent=2, default=str)
@mcp.tool() @mcp.tool()
async def check_system_health() -> str: async def check_system_health() -> str:
"""Check every dependency HyperSwap needs (NVML, sudo nvidia-smi, fan control via the """Check every dependency HyperSwap needs (NVML, sudo nvidia-smi, fan control via the

View File

@@ -16,6 +16,7 @@ from fastapi.middleware.cors import CORSMiddleware
from pydantic import BaseModel, Field from pydantic import BaseModel, Field
import autotune import autotune
import engines
import health import health
import overclock_manager import overclock_manager
import ram_optimizer import ram_optimizer
@@ -321,6 +322,13 @@ async def api_health():
return await health.run_health_checks() return await health.run_health_checks()
@app.get("/api/engines", summary="Live Engine Configuration", tags=["Telemetry"])
async def api_engines():
"""Real configuration of Ollama and ComfyUI, with what each setting implies for
arbitration. These live outside this codebase but dictate how it must behave."""
return await engines.get_engine_config()
@app.get("/api/gpu", summary="GPU Sensors and VRAM Breakdown", tags=["Telemetry"]) @app.get("/api/gpu", summary="GPU Sensors and VRAM Breakdown", tags=["Telemetry"])
async def get_gpu_metrics() -> Dict[str, Any]: async def get_gpu_metrics() -> Dict[str, Any]:
"""Detailed NVML sensors (utilization, temp, power, fan, clocks, throttle reasons, per-process VRAM).""" """Detailed NVML sensors (utilization, temp, power, fan, clocks, throttle reasons, per-process VRAM)."""

View File

@@ -1069,3 +1069,35 @@ async function fetchLastSwapFromStore() {
document.addEventListener('DOMContentLoaded', () => { document.addEventListener('DOMContentLoaded', () => {
setTimeout(fetchLastSwapFromStore, 1500); setTimeout(fetchLastSwapFromStore, 1500);
}); });
// ---------------------------------------------------------------- engine config
async function fetchEngineConfig() {
// These subtitles used to be hardcoded. They happened to be accurate, which is worse
// than being wrong: they would have stayed accurate-looking after the settings changed.
try {
const d = await (await fetch('/api/engines')).json();
const o = document.getElementById('ollama-engine-sub');
if (o && d.ollama) {
const bits = [`Port :${d.ollama.port}`, d.ollama.summary];
if (d.ollama.max_loaded_models) bits.push(`${d.ollama.max_loaded_models} model resident`);
if (d.ollama.keep_alive) bits.push(`keep-alive ${d.ollama.keep_alive}`);
o.textContent = bits.join(' // ');
o.title = (d.ollama.settings || [])
.filter(s => s.means)
.map(s => `${s.key}=${s.value} — ${s.means}`)
.join('\n');
}
const c = document.getElementById('comfy-engine-sub');
if (c && d.comfyui && d.comfyui.online) {
c.textContent = `Port :8188 // v${d.comfyui.version} // ${d.comfyui.summary}`;
c.title = `torch ${d.comfyui.pytorch}\n${d.comfyui.device || ''}`;
}
} catch (e) { /* subtitles are cosmetic; never break the page over them */ }
}
document.addEventListener('DOMContentLoaded', () => {
fetchEngineConfig();
setInterval(fetchEngineConfig, 120000);
});

View File

@@ -194,7 +194,7 @@
</div> </div>
<div> <div>
<h3 class="font-bold text-slate-100 text-sm">Ollama LLM Engine</h3> <h3 class="font-bold text-slate-100 text-sm">Ollama LLM Engine</h3>
<p class="text-xs text-slate-400">Port :11434 // FlashAttention + Q4 KV Cache</p> <p class="text-xs text-slate-400" id="ollama-engine-sub" title="Read live from the ollama service environment">Port :11434</p>
</div> </div>
</div> </div>
<button onclick="freeOllamaVRAM()" class="px-2.5 py-1 text-xs font-semibold rounded-lg bg-rose-950/70 border border-rose-800 text-rose-300 hover:bg-rose-900 transition flex items-center space-x-1"> <button onclick="freeOllamaVRAM()" class="px-2.5 py-1 text-xs font-semibold rounded-lg bg-rose-950/70 border border-rose-800 text-rose-300 hover:bg-rose-900 transition flex items-center space-x-1">
@@ -265,7 +265,7 @@
</div> </div>
<div> <div>
<h3 class="font-bold text-slate-100 text-sm">ComfyUI Diffusion Engine</h3> <h3 class="font-bold text-slate-100 text-sm">ComfyUI Diffusion Engine</h3>
<p class="text-xs text-slate-400">Port :8188 // DynamicVRAM + Pinned Async Offload</p> <p class="text-xs text-slate-400" id="comfy-engine-sub" title="Read live from ComfyUI's /system_stats">Port :8188</p>
</div> </div>
</div> </div>
<button onclick="freeComfyVRAM()" class="px-2.5 py-1 text-xs font-semibold rounded-lg bg-rose-950/70 border border-rose-800 text-rose-300 hover:bg-rose-900 transition flex items-center space-x-1"> <button onclick="freeComfyVRAM()" class="px-2.5 py-1 text-xs font-semibold rounded-lg bg-rose-950/70 border border-rose-800 text-rose-300 hover:bg-rose-900 transition flex items-center space-x-1">

133
tests/test_engines.py Normal file
View File

@@ -0,0 +1,133 @@
"""Tests for live engine-configuration reporting.
These settings live outside this codebase but dictate how arbitration must behave, and
working out why a yield behaved a certain way once meant reading journald by hand. The
dashboard previously asserted them as hardcoded text, which happened to be accurate --
worse than being wrong, because it would have stayed accurate-looking after the settings
changed.
"""
import asyncio
import pytest
import engines
class _Proc:
def __init__(self, stdout=""):
self.stdout = stdout
self.returncode = 0
class TestOllamaEnvironmentParsing:
def test_parses_the_real_unit_environment(self, monkeypatch):
# Verbatim from `systemctl show ollama -p Environment --value` on this machine.
raw = ("OLLAMA_HOST=0.0.0.0:11434 OLLAMA_FLASH_ATTENTION=1 "
"OLLAMA_KV_CACHE_TYPE=q4_0 OLLAMA_KEEP_ALIVE=30m "
"OLLAMA_MAX_LOADED_MODELS=1 OLLAMA_NUM_PARALLEL=1 OLLAMA_NUM_BATCH=2048")
monkeypatch.setattr(engines.subprocess, "run", lambda *a, **k: _Proc(raw))
env = engines._ollama_unit_environment()
assert env["OLLAMA_NUM_PARALLEL"] == "1"
assert env["OLLAMA_MAX_LOADED_MODELS"] == "1"
assert env["OLLAMA_KV_CACHE_TYPE"] == "q4_0"
def test_ignores_non_ollama_variables(self, monkeypatch):
monkeypatch.setattr(engines.subprocess, "run",
lambda *a, **k: _Proc("PATH=/usr/bin OLLAMA_HOST=x:1 HOME=/root"))
env = engines._ollama_unit_environment()
assert set(env) == {"OLLAMA_HOST"}
def test_returns_empty_rather_than_raising_when_systemctl_fails(self, monkeypatch):
def boom(*a, **k):
raise FileNotFoundError("systemctl")
monkeypatch.setattr(engines.subprocess, "run", boom)
assert engines._ollama_unit_environment() == {}
class TestEngineConfigReport:
def _run(self, monkeypatch, env, comfy_ok=True):
monkeypatch.setattr(engines, "_ollama_unit_environment", lambda: env)
class _Resp:
status_code = 200 if comfy_ok else 500
def json(self):
return {"system": {"comfyui_version": "0.33.1",
"pytorch_version": "2.11.0+cu128",
"python_version": "3.14.4 (main)",
"argv": ["main.py", "--listen", "0.0.0.0"]},
"devices": [{"name": "cuda:0 NVIDIA GeForce RTX 4080 SUPER "
": cudaMallocAsync"}]}
class _Client:
async def __aenter__(self): return self
async def __aexit__(self, *a): return False
async def get(self, url): return _Resp()
monkeypatch.setattr(engines.httpx, "AsyncClient", lambda **k: _Client())
return asyncio.run(engines.get_engine_config())
def test_surfaces_the_settings_that_drive_arbitration(self, monkeypatch):
d = self._run(monkeypatch, {"OLLAMA_NUM_PARALLEL": "1",
"OLLAMA_MAX_LOADED_MODELS": "1",
"OLLAMA_KEEP_ALIVE": "30m"})
assert d["ollama"]["num_parallel"] == "1"
assert d["ollama"]["max_loaded_models"] == "1"
assert d["ollama"]["keep_alive"] == "30m"
def test_num_parallel_explains_the_deferred_yield_behaviour(self, monkeypatch):
d = self._run(monkeypatch, {"OLLAMA_NUM_PARALLEL": "1"})
note = next(s["means"] for s in d["ollama"]["settings"]
if s["key"] == "OLLAMA_NUM_PARALLEL")
# The explanation is the point: it is why a busy model is deferred, not failed.
assert "queue" in note.lower()
def test_summary_reflects_actual_flags_not_a_fixed_string(self, monkeypatch):
on = self._run(monkeypatch, {"OLLAMA_FLASH_ATTENTION": "1",
"OLLAMA_KV_CACHE_TYPE": "q4_0"})
assert "FlashAttention" in on["ollama"]["summary"]
assert "q4_0" in on["ollama"]["summary"]
off = self._run(monkeypatch, {})
assert "FlashAttention" not in off["ollama"]["summary"]
assert off["ollama"]["config_source"] == "unavailable"
def test_port_comes_from_ollama_host(self, monkeypatch):
d = self._run(monkeypatch, {"OLLAMA_HOST": "0.0.0.0:11500"})
assert d["ollama"]["port"] == "11500"
def test_comfy_allocator_and_vram_mode_are_read_not_asserted(self, monkeypatch):
d = self._run(monkeypatch, {})
assert d["comfyui"]["allocator"] == "cudaMallocAsync"
assert d["comfyui"]["vram_mode"] == "default (auto)"
assert d["comfyui"]["version"] == "0.33.1"
def test_comfy_vram_flag_is_detected_when_present(self, monkeypatch):
monkeypatch.setattr(engines, "_ollama_unit_environment", lambda: {})
class _Resp:
status_code = 200
def json(self):
return {"system": {"argv": ["main.py", "--lowvram"]},
"devices": [{"name": "cuda:0 X : cudaMalloc"}]}
class _Client:
async def __aenter__(self): return self
async def __aexit__(self, *a): return False
async def get(self, url): return _Resp()
monkeypatch.setattr(engines.httpx, "AsyncClient", lambda **k: _Client())
d = asyncio.run(engines.get_engine_config())
assert d["comfyui"]["vram_mode"] == "--lowvram"
assert d["comfyui"]["allocator"] == "cudaMalloc"
def test_offline_comfy_is_reported_not_raised(self, monkeypatch):
monkeypatch.setattr(engines, "_ollama_unit_environment", lambda: {})
class _Client:
async def __aenter__(self): return self
async def __aexit__(self, *a): return False
async def get(self, url): raise ConnectionError("refused")
monkeypatch.setattr(engines.httpx, "AsyncClient", lambda **k: _Client())
d = asyncio.run(engines.get_engine_config())
assert d["comfyui"]["online"] is False
assert "error" in d["comfyui"]