Read engine configuration live instead of asserting it in the dashboard
The engine subtitles were hardcoded: "FlashAttention + Q4 KV Cache" and "DynamicVRAM
+ Pinned Async Offload". The first turned out to be accurate -- OLLAMA_FLASH_ATTENTION
and OLLAMA_KV_CACHE_TYPE really are set -- which is worse than being wrong, because it
would have gone on looking accurate after the settings changed.
engines.py reads both engines' real configuration: the ollama service environment via
systemd, and ComfyUI's own /system_stats for version, allocator, VRAM mode and argv.
Exposed at GET /api/engines, as an MCP tool, and in the dashboard subtitles with the
full settings list as a tooltip.
The settings worth surfacing are the ones that dictate how this service must behave
and that previously had to be discovered by reading journald: OLLAMA_NUM_PARALLEL=1
is why an unload queues behind a running generation and is reported as deferred
rather than failed, and OLLAMA_MAX_LOADED_MODELS=1 is why every swap evicts the
previous model. Each is reported with that explanation attached.
Writing the tests found a bug in the new code: (system.get("python_version") or
"").split()[0] raises IndexError when ComfyUI omits the field, and the surrounding
except would have swallowed it and reported ComfyUI as entirely offline.
Tests: 192 (was 182).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
15
README.md
15
README.md
@@ -111,6 +111,18 @@ only if the card actually needs it.
|
||||
* **Honest gain reporting**: gain against the profile's *current* setting is reported separately from the spread across values tried. Conflating them turns a flat result into a headline "+102%".
|
||||
* **Safety**: refuses to start while ComfyUI is executing, suspends the arbitrator's automatic profile switching for the duration (otherwise a diffusion benchmark trips `trigger_comfy_priority`, which reapplies the whole profile and overwrites the clock being measured), and restores the original profile in a `finally` block — including on exception or cancellation.
|
||||
|
||||
### 🔧 Live Engine Configuration (`engines.py`)
|
||||
* `GET /api/engines` reports the **real, current** configuration of both engines and what
|
||||
each setting implies for arbitration — because the settings that dictate this service's
|
||||
behaviour live outside its own codebase.
|
||||
* `OLLAMA_NUM_PARALLEL=1` is why a `keep_alive: 0` unload queues behind a running
|
||||
generation and is reported as *deferred* rather than failed. `OLLAMA_MAX_LOADED_MODELS=1`
|
||||
is why every swap evicts the previous model. Working these out originally meant reading
|
||||
journald and the systemd unit by hand.
|
||||
* The dashboard's engine subtitles now come from this endpoint. They were previously
|
||||
hardcoded — and happened to be accurate, which is worse than being wrong, since they
|
||||
would have kept looking accurate after the configuration changed.
|
||||
|
||||
### 🩺 Dependency Self-Check (`health.py`)
|
||||
* `GET /api/health` verifies **everything this service depends on**: NVML, passwordless
|
||||
sudo for `nvidia-smi`, fan control through the headless X server, overclock drift, the
|
||||
@@ -156,7 +168,7 @@ only if the card actually needs it.
|
||||
## 1a. Tests
|
||||
|
||||
```bash
|
||||
/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/ -q # 182 passed in ~3.7s
|
||||
/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/ -q # 192 passed in ~3.7s
|
||||
```
|
||||
|
||||
Hermetic: no GPU, no network, no sleeps. An autouse fixture stubs `overclock_manager._sh`
|
||||
@@ -264,6 +276,7 @@ The HyperSwap server runs on port `9090` by default. Interactive OpenAPI/Swagger
|
||||
| `/api/analytics/models` | `GET` | Recency/frequency model ranking used to prioritise the warm budget. |
|
||||
| `/api/history?durable=true` | `GET` | Swap history from the persistent store rather than the in-memory ring. |
|
||||
| `/api/db` | `GET` | Store location, row counts and how many hours of history are held. |
|
||||
| `/api/engines` | `GET` | Live Ollama and ComfyUI configuration, with what each setting implies for arbitration. |
|
||||
| `/api/health` | `GET` | Dependency self-check: NVML, sudo, fan control, drift, store, upstreams — each with impact and remediation. |
|
||||
|
||||
### Governor & Autotune Endpoints
|
||||
|
||||
117
engines.py
Normal file
117
engines.py
Normal file
@@ -0,0 +1,117 @@
|
||||
"""Live configuration of the two engines HyperSwap arbitrates between.
|
||||
|
||||
Arbitration behaviour is largely dictated by settings that live outside this codebase.
|
||||
Working out why a yield behaved the way it did meant reading journald and the ollama
|
||||
unit by hand: OLLAMA_NUM_PARALLEL decides whether an unload queues behind a running
|
||||
generation, OLLAMA_MAX_LOADED_MODELS decides whether more than one model can be
|
||||
resident, and a pinned n_gpu_layers decides whether a model that will not fit spills to
|
||||
the CPU or fails outright. Those are worth reading and explaining rather than hardcoding
|
||||
into a dashboard subtitle that silently goes stale.
|
||||
"""
|
||||
import json
|
||||
import logging
|
||||
import subprocess
|
||||
from typing import Any, Dict, List, Optional
|
||||
|
||||
import httpx
|
||||
|
||||
import vram_arbitrator
|
||||
|
||||
logger = logging.getLogger("engines")
|
||||
|
||||
# Settings that change how the arbitrator must behave, with what they imply.
|
||||
OLLAMA_SETTING_NOTES = {
|
||||
"OLLAMA_NUM_PARALLEL": (
|
||||
"Requests per model. At 1, a keep_alive:0 unload queues behind any running "
|
||||
"generation and applies when it finishes — which is why a busy model is "
|
||||
"reported as deferred rather than failed."),
|
||||
"OLLAMA_MAX_LOADED_MODELS": (
|
||||
"How many models may be resident at once. At 1, Ollama evicts the previous "
|
||||
"model on every swap."),
|
||||
"OLLAMA_KEEP_ALIVE": (
|
||||
"Default residency after a request. Long values keep VRAM occupied and make "
|
||||
"ComfyUI wait for an explicit yield."),
|
||||
"OLLAMA_FLASH_ATTENTION": "FlashAttention kernels for attention.",
|
||||
"OLLAMA_KV_CACHE_TYPE": "KV cache quantisation; smaller types cut VRAM per context.",
|
||||
"OLLAMA_NUM_BATCH": "Prompt-evaluation batch size.",
|
||||
}
|
||||
|
||||
|
||||
def _ollama_unit_environment() -> Dict[str, str]:
|
||||
"""Read the ollama service's environment. Empty if it is not a systemd unit."""
|
||||
env: Dict[str, str] = {}
|
||||
try:
|
||||
proc = subprocess.run(["systemctl", "show", "ollama", "-p", "Environment",
|
||||
"--value"], capture_output=True, text=True, timeout=8)
|
||||
for token in proc.stdout.split():
|
||||
if "=" in token and token.startswith("OLLAMA"):
|
||||
k, _, v = token.partition("=")
|
||||
env[k] = v
|
||||
except Exception as e:
|
||||
logger.debug(f"could not read ollama unit environment: {e}")
|
||||
return env
|
||||
|
||||
|
||||
async def get_engine_config() -> Dict[str, Any]:
|
||||
"""Real, live configuration of both engines, with arbitration implications."""
|
||||
ollama_env = _ollama_unit_environment()
|
||||
ollama_settings = [
|
||||
{"key": k, "value": v, "means": OLLAMA_SETTING_NOTES.get(k, "")}
|
||||
for k, v in sorted(ollama_env.items())
|
||||
]
|
||||
|
||||
# A short, honest summary line to replace the dashboard's hardcoded subtitle.
|
||||
feature_bits: List[str] = []
|
||||
if ollama_env.get("OLLAMA_FLASH_ATTENTION") == "1":
|
||||
feature_bits.append("FlashAttention")
|
||||
kv = ollama_env.get("OLLAMA_KV_CACHE_TYPE")
|
||||
if kv:
|
||||
feature_bits.append(f"{kv} KV cache")
|
||||
host = ollama_env.get("OLLAMA_HOST", "")
|
||||
port = host.rsplit(":", 1)[-1] if ":" in host else "11434"
|
||||
|
||||
ollama = {
|
||||
"port": port,
|
||||
"settings": ollama_settings,
|
||||
"summary": " + ".join(feature_bits) if feature_bits else "default configuration",
|
||||
"max_loaded_models": ollama_env.get("OLLAMA_MAX_LOADED_MODELS"),
|
||||
"num_parallel": ollama_env.get("OLLAMA_NUM_PARALLEL"),
|
||||
"keep_alive": ollama_env.get("OLLAMA_KEEP_ALIVE"),
|
||||
"config_source": "systemd unit environment" if ollama_env else "unavailable",
|
||||
}
|
||||
|
||||
comfy: Dict[str, Any] = {"online": False}
|
||||
try:
|
||||
async with httpx.AsyncClient(timeout=4.0) as c:
|
||||
r = await c.get(f"{vram_arbitrator.COMFY_API_BASE}/system_stats")
|
||||
if r.status_code == 200:
|
||||
data = r.json()
|
||||
system = data.get("system", {})
|
||||
argv = system.get("argv") or []
|
||||
devices = data.get("devices") or []
|
||||
dev = devices[0] if devices else {}
|
||||
# The allocator is named in the device string; it is the closest thing
|
||||
# ComfyUI reports to the "async offload" the old subtitle asserted.
|
||||
dev_name = dev.get("name", "")
|
||||
allocator = ("cudaMallocAsync" if "cudaMallocAsync" in dev_name
|
||||
else "cudaMalloc" if "cudaMalloc" in dev_name else "unknown")
|
||||
vram_flags = [a for a in argv
|
||||
if a in ("--lowvram", "--novram", "--highvram", "--normalvram",
|
||||
"--gpu-only", "--cpu")]
|
||||
comfy = {
|
||||
"online": True,
|
||||
"version": system.get("comfyui_version"),
|
||||
"pytorch": system.get("pytorch_version"),
|
||||
# split()[0] on an absent version raises IndexError, which would have
|
||||
# been swallowed by the except below and reported ComfyUI as offline.
|
||||
"python": ((system.get("python_version") or "").split() or [None])[0],
|
||||
"argv": argv,
|
||||
"vram_mode": vram_flags[0] if vram_flags else "default (auto)",
|
||||
"allocator": allocator,
|
||||
"device": dev_name,
|
||||
"summary": f"{allocator}, {vram_flags[0] if vram_flags else 'auto VRAM'}",
|
||||
}
|
||||
except Exception as e:
|
||||
comfy = {"online": False, "error": str(e)[:120]}
|
||||
|
||||
return {"ollama": ollama, "comfyui": comfy}
|
||||
@@ -8,6 +8,7 @@ from typing import Dict, List, Any, Optional
|
||||
|
||||
from mcp.server import MCPServer
|
||||
import autotune
|
||||
import engines
|
||||
import health
|
||||
import overclock_manager
|
||||
import ram_optimizer
|
||||
@@ -137,6 +138,14 @@ def set_gpu_fan_speed(mode: str = "auto", percent: Optional[int] = None) -> str:
|
||||
res = overclock_manager.set_fan_auto()
|
||||
return json.dumps(res, indent=2)
|
||||
|
||||
@mcp.tool()
|
||||
async def get_engine_config() -> str:
|
||||
"""Live configuration of Ollama and ComfyUI (parallelism, max loaded models,
|
||||
keep-alive, KV cache type, ComfyUI VRAM mode and allocator), with what each setting
|
||||
implies for VRAM arbitration."""
|
||||
return json.dumps(await engines.get_engine_config(), indent=2, default=str)
|
||||
|
||||
|
||||
@mcp.tool()
|
||||
async def check_system_health() -> str:
|
||||
"""Check every dependency HyperSwap needs (NVML, sudo nvidia-smi, fan control via the
|
||||
|
||||
@@ -16,6 +16,7 @@ from fastapi.middleware.cors import CORSMiddleware
|
||||
from pydantic import BaseModel, Field
|
||||
|
||||
import autotune
|
||||
import engines
|
||||
import health
|
||||
import overclock_manager
|
||||
import ram_optimizer
|
||||
@@ -321,6 +322,13 @@ async def api_health():
|
||||
return await health.run_health_checks()
|
||||
|
||||
|
||||
@app.get("/api/engines", summary="Live Engine Configuration", tags=["Telemetry"])
|
||||
async def api_engines():
|
||||
"""Real configuration of Ollama and ComfyUI, with what each setting implies for
|
||||
arbitration. These live outside this codebase but dictate how it must behave."""
|
||||
return await engines.get_engine_config()
|
||||
|
||||
|
||||
@app.get("/api/gpu", summary="GPU Sensors and VRAM Breakdown", tags=["Telemetry"])
|
||||
async def get_gpu_metrics() -> Dict[str, Any]:
|
||||
"""Detailed NVML sensors (utilization, temp, power, fan, clocks, throttle reasons, per-process VRAM)."""
|
||||
|
||||
@@ -1069,3 +1069,35 @@ async function fetchLastSwapFromStore() {
|
||||
document.addEventListener('DOMContentLoaded', () => {
|
||||
setTimeout(fetchLastSwapFromStore, 1500);
|
||||
});
|
||||
|
||||
|
||||
// ---------------------------------------------------------------- engine config
|
||||
|
||||
async function fetchEngineConfig() {
|
||||
// These subtitles used to be hardcoded. They happened to be accurate, which is worse
|
||||
// than being wrong: they would have stayed accurate-looking after the settings changed.
|
||||
try {
|
||||
const d = await (await fetch('/api/engines')).json();
|
||||
const o = document.getElementById('ollama-engine-sub');
|
||||
if (o && d.ollama) {
|
||||
const bits = [`Port :${d.ollama.port}`, d.ollama.summary];
|
||||
if (d.ollama.max_loaded_models) bits.push(`${d.ollama.max_loaded_models} model resident`);
|
||||
if (d.ollama.keep_alive) bits.push(`keep-alive ${d.ollama.keep_alive}`);
|
||||
o.textContent = bits.join(' // ');
|
||||
o.title = (d.ollama.settings || [])
|
||||
.filter(s => s.means)
|
||||
.map(s => `${s.key}=${s.value} — ${s.means}`)
|
||||
.join('\n');
|
||||
}
|
||||
const c = document.getElementById('comfy-engine-sub');
|
||||
if (c && d.comfyui && d.comfyui.online) {
|
||||
c.textContent = `Port :8188 // v${d.comfyui.version} // ${d.comfyui.summary}`;
|
||||
c.title = `torch ${d.comfyui.pytorch}\n${d.comfyui.device || ''}`;
|
||||
}
|
||||
} catch (e) { /* subtitles are cosmetic; never break the page over them */ }
|
||||
}
|
||||
|
||||
document.addEventListener('DOMContentLoaded', () => {
|
||||
fetchEngineConfig();
|
||||
setInterval(fetchEngineConfig, 120000);
|
||||
});
|
||||
|
||||
@@ -194,7 +194,7 @@
|
||||
</div>
|
||||
<div>
|
||||
<h3 class="font-bold text-slate-100 text-sm">Ollama LLM Engine</h3>
|
||||
<p class="text-xs text-slate-400">Port :11434 // FlashAttention + Q4 KV Cache</p>
|
||||
<p class="text-xs text-slate-400" id="ollama-engine-sub" title="Read live from the ollama service environment">Port :11434</p>
|
||||
</div>
|
||||
</div>
|
||||
<button onclick="freeOllamaVRAM()" class="px-2.5 py-1 text-xs font-semibold rounded-lg bg-rose-950/70 border border-rose-800 text-rose-300 hover:bg-rose-900 transition flex items-center space-x-1">
|
||||
@@ -265,7 +265,7 @@
|
||||
</div>
|
||||
<div>
|
||||
<h3 class="font-bold text-slate-100 text-sm">ComfyUI Diffusion Engine</h3>
|
||||
<p class="text-xs text-slate-400">Port :8188 // DynamicVRAM + Pinned Async Offload</p>
|
||||
<p class="text-xs text-slate-400" id="comfy-engine-sub" title="Read live from ComfyUI's /system_stats">Port :8188</p>
|
||||
</div>
|
||||
</div>
|
||||
<button onclick="freeComfyVRAM()" class="px-2.5 py-1 text-xs font-semibold rounded-lg bg-rose-950/70 border border-rose-800 text-rose-300 hover:bg-rose-900 transition flex items-center space-x-1">
|
||||
|
||||
133
tests/test_engines.py
Normal file
133
tests/test_engines.py
Normal file
@@ -0,0 +1,133 @@
|
||||
"""Tests for live engine-configuration reporting.
|
||||
|
||||
These settings live outside this codebase but dictate how arbitration must behave, and
|
||||
working out why a yield behaved a certain way once meant reading journald by hand. The
|
||||
dashboard previously asserted them as hardcoded text, which happened to be accurate --
|
||||
worse than being wrong, because it would have stayed accurate-looking after the settings
|
||||
changed.
|
||||
"""
|
||||
import asyncio
|
||||
|
||||
import pytest
|
||||
|
||||
import engines
|
||||
|
||||
|
||||
class _Proc:
|
||||
def __init__(self, stdout=""):
|
||||
self.stdout = stdout
|
||||
self.returncode = 0
|
||||
|
||||
|
||||
class TestOllamaEnvironmentParsing:
|
||||
def test_parses_the_real_unit_environment(self, monkeypatch):
|
||||
# Verbatim from `systemctl show ollama -p Environment --value` on this machine.
|
||||
raw = ("OLLAMA_HOST=0.0.0.0:11434 OLLAMA_FLASH_ATTENTION=1 "
|
||||
"OLLAMA_KV_CACHE_TYPE=q4_0 OLLAMA_KEEP_ALIVE=30m "
|
||||
"OLLAMA_MAX_LOADED_MODELS=1 OLLAMA_NUM_PARALLEL=1 OLLAMA_NUM_BATCH=2048")
|
||||
monkeypatch.setattr(engines.subprocess, "run", lambda *a, **k: _Proc(raw))
|
||||
env = engines._ollama_unit_environment()
|
||||
assert env["OLLAMA_NUM_PARALLEL"] == "1"
|
||||
assert env["OLLAMA_MAX_LOADED_MODELS"] == "1"
|
||||
assert env["OLLAMA_KV_CACHE_TYPE"] == "q4_0"
|
||||
|
||||
def test_ignores_non_ollama_variables(self, monkeypatch):
|
||||
monkeypatch.setattr(engines.subprocess, "run",
|
||||
lambda *a, **k: _Proc("PATH=/usr/bin OLLAMA_HOST=x:1 HOME=/root"))
|
||||
env = engines._ollama_unit_environment()
|
||||
assert set(env) == {"OLLAMA_HOST"}
|
||||
|
||||
def test_returns_empty_rather_than_raising_when_systemctl_fails(self, monkeypatch):
|
||||
def boom(*a, **k):
|
||||
raise FileNotFoundError("systemctl")
|
||||
monkeypatch.setattr(engines.subprocess, "run", boom)
|
||||
assert engines._ollama_unit_environment() == {}
|
||||
|
||||
|
||||
class TestEngineConfigReport:
|
||||
def _run(self, monkeypatch, env, comfy_ok=True):
|
||||
monkeypatch.setattr(engines, "_ollama_unit_environment", lambda: env)
|
||||
|
||||
class _Resp:
|
||||
status_code = 200 if comfy_ok else 500
|
||||
def json(self):
|
||||
return {"system": {"comfyui_version": "0.33.1",
|
||||
"pytorch_version": "2.11.0+cu128",
|
||||
"python_version": "3.14.4 (main)",
|
||||
"argv": ["main.py", "--listen", "0.0.0.0"]},
|
||||
"devices": [{"name": "cuda:0 NVIDIA GeForce RTX 4080 SUPER "
|
||||
": cudaMallocAsync"}]}
|
||||
|
||||
class _Client:
|
||||
async def __aenter__(self): return self
|
||||
async def __aexit__(self, *a): return False
|
||||
async def get(self, url): return _Resp()
|
||||
|
||||
monkeypatch.setattr(engines.httpx, "AsyncClient", lambda **k: _Client())
|
||||
return asyncio.run(engines.get_engine_config())
|
||||
|
||||
def test_surfaces_the_settings_that_drive_arbitration(self, monkeypatch):
|
||||
d = self._run(monkeypatch, {"OLLAMA_NUM_PARALLEL": "1",
|
||||
"OLLAMA_MAX_LOADED_MODELS": "1",
|
||||
"OLLAMA_KEEP_ALIVE": "30m"})
|
||||
assert d["ollama"]["num_parallel"] == "1"
|
||||
assert d["ollama"]["max_loaded_models"] == "1"
|
||||
assert d["ollama"]["keep_alive"] == "30m"
|
||||
|
||||
def test_num_parallel_explains_the_deferred_yield_behaviour(self, monkeypatch):
|
||||
d = self._run(monkeypatch, {"OLLAMA_NUM_PARALLEL": "1"})
|
||||
note = next(s["means"] for s in d["ollama"]["settings"]
|
||||
if s["key"] == "OLLAMA_NUM_PARALLEL")
|
||||
# The explanation is the point: it is why a busy model is deferred, not failed.
|
||||
assert "queue" in note.lower()
|
||||
|
||||
def test_summary_reflects_actual_flags_not_a_fixed_string(self, monkeypatch):
|
||||
on = self._run(monkeypatch, {"OLLAMA_FLASH_ATTENTION": "1",
|
||||
"OLLAMA_KV_CACHE_TYPE": "q4_0"})
|
||||
assert "FlashAttention" in on["ollama"]["summary"]
|
||||
assert "q4_0" in on["ollama"]["summary"]
|
||||
off = self._run(monkeypatch, {})
|
||||
assert "FlashAttention" not in off["ollama"]["summary"]
|
||||
assert off["ollama"]["config_source"] == "unavailable"
|
||||
|
||||
def test_port_comes_from_ollama_host(self, monkeypatch):
|
||||
d = self._run(monkeypatch, {"OLLAMA_HOST": "0.0.0.0:11500"})
|
||||
assert d["ollama"]["port"] == "11500"
|
||||
|
||||
def test_comfy_allocator_and_vram_mode_are_read_not_asserted(self, monkeypatch):
|
||||
d = self._run(monkeypatch, {})
|
||||
assert d["comfyui"]["allocator"] == "cudaMallocAsync"
|
||||
assert d["comfyui"]["vram_mode"] == "default (auto)"
|
||||
assert d["comfyui"]["version"] == "0.33.1"
|
||||
|
||||
def test_comfy_vram_flag_is_detected_when_present(self, monkeypatch):
|
||||
monkeypatch.setattr(engines, "_ollama_unit_environment", lambda: {})
|
||||
|
||||
class _Resp:
|
||||
status_code = 200
|
||||
def json(self):
|
||||
return {"system": {"argv": ["main.py", "--lowvram"]},
|
||||
"devices": [{"name": "cuda:0 X : cudaMalloc"}]}
|
||||
|
||||
class _Client:
|
||||
async def __aenter__(self): return self
|
||||
async def __aexit__(self, *a): return False
|
||||
async def get(self, url): return _Resp()
|
||||
|
||||
monkeypatch.setattr(engines.httpx, "AsyncClient", lambda **k: _Client())
|
||||
d = asyncio.run(engines.get_engine_config())
|
||||
assert d["comfyui"]["vram_mode"] == "--lowvram"
|
||||
assert d["comfyui"]["allocator"] == "cudaMalloc"
|
||||
|
||||
def test_offline_comfy_is_reported_not_raised(self, monkeypatch):
|
||||
monkeypatch.setattr(engines, "_ollama_unit_environment", lambda: {})
|
||||
|
||||
class _Client:
|
||||
async def __aenter__(self): return self
|
||||
async def __aexit__(self, *a): return False
|
||||
async def get(self, url): raise ConnectionError("refused")
|
||||
|
||||
monkeypatch.setattr(engines.httpx, "AsyncClient", lambda **k: _Client())
|
||||
d = asyncio.run(engines.get_engine_config())
|
||||
assert d["comfyui"]["online"] is False
|
||||
assert "error" in d["comfyui"]
|
||||
Reference in New Issue
Block a user