From 043d61722b1b0891f89a7c4dd5fdb94269418f40 Mon Sep 17 00:00:00 2001 From: drjones Date: Sun, 6 Sep 2026 10:56:04 -0700 Subject: [PATCH] Read engine configuration live instead of asserting it in the dashboard The engine subtitles were hardcoded: "FlashAttention + Q4 KV Cache" and "DynamicVRAM + Pinned Async Offload". The first turned out to be accurate -- OLLAMA_FLASH_ATTENTION and OLLAMA_KV_CACHE_TYPE really are set -- which is worse than being wrong, because it would have gone on looking accurate after the settings changed. engines.py reads both engines' real configuration: the ollama service environment via systemd, and ComfyUI's own /system_stats for version, allocator, VRAM mode and argv. Exposed at GET /api/engines, as an MCP tool, and in the dashboard subtitles with the full settings list as a tooltip. The settings worth surfacing are the ones that dictate how this service must behave and that previously had to be discovered by reading journald: OLLAMA_NUM_PARALLEL=1 is why an unload queues behind a running generation and is reported as deferred rather than failed, and OLLAMA_MAX_LOADED_MODELS=1 is why every swap evicts the previous model. Each is reported with that explanation attached. Writing the tests found a bug in the new code: (system.get("python_version") or "").split()[0] raises IndexError when ComfyUI omits the field, and the surrounding except would have swallowed it and reported ComfyUI as entirely offline. Tests: 192 (was 182). Co-Authored-By: Claude Opus 5 --- README.md | 15 ++++- engines.py | 117 +++++++++++++++++++++++++++++++++++++ mcp_server.py | 9 +++ server.py | 8 +++ static/app.js | 32 ++++++++++ static/index.html | 4 +- tests/test_engines.py | 133 ++++++++++++++++++++++++++++++++++++++++++ 7 files changed, 315 insertions(+), 3 deletions(-) create mode 100644 engines.py create mode 100644 tests/test_engines.py diff --git a/README.md b/README.md index 9c7649e..ef5de6f 100644 --- a/README.md +++ b/README.md @@ -111,6 +111,18 @@ only if the card actually needs it. * **Honest gain reporting**: gain against the profile's *current* setting is reported separately from the spread across values tried. Conflating them turns a flat result into a headline "+102%". * **Safety**: refuses to start while ComfyUI is executing, suspends the arbitrator's automatic profile switching for the duration (otherwise a diffusion benchmark trips `trigger_comfy_priority`, which reapplies the whole profile and overwrites the clock being measured), and restores the original profile in a `finally` block — including on exception or cancellation. +### 🔧 Live Engine Configuration (`engines.py`) +* `GET /api/engines` reports the **real, current** configuration of both engines and what + each setting implies for arbitration — because the settings that dictate this service's + behaviour live outside its own codebase. +* `OLLAMA_NUM_PARALLEL=1` is why a `keep_alive: 0` unload queues behind a running + generation and is reported as *deferred* rather than failed. `OLLAMA_MAX_LOADED_MODELS=1` + is why every swap evicts the previous model. Working these out originally meant reading + journald and the systemd unit by hand. +* The dashboard's engine subtitles now come from this endpoint. They were previously + hardcoded — and happened to be accurate, which is worse than being wrong, since they + would have kept looking accurate after the configuration changed. + ### 🩺 Dependency Self-Check (`health.py`) * `GET /api/health` verifies **everything this service depends on**: NVML, passwordless sudo for `nvidia-smi`, fan control through the headless X server, overclock drift, the @@ -156,7 +168,7 @@ only if the card actually needs it. ## 1a. Tests ```bash -/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/ -q # 182 passed in ~3.7s +/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/ -q # 192 passed in ~3.7s ``` Hermetic: no GPU, no network, no sleeps. An autouse fixture stubs `overclock_manager._sh` @@ -264,6 +276,7 @@ The HyperSwap server runs on port `9090` by default. Interactive OpenAPI/Swagger | `/api/analytics/models` | `GET` | Recency/frequency model ranking used to prioritise the warm budget. | | `/api/history?durable=true` | `GET` | Swap history from the persistent store rather than the in-memory ring. | | `/api/db` | `GET` | Store location, row counts and how many hours of history are held. | +| `/api/engines` | `GET` | Live Ollama and ComfyUI configuration, with what each setting implies for arbitration. | | `/api/health` | `GET` | Dependency self-check: NVML, sudo, fan control, drift, store, upstreams — each with impact and remediation. | ### Governor & Autotune Endpoints diff --git a/engines.py b/engines.py new file mode 100644 index 0000000..888ae51 --- /dev/null +++ b/engines.py @@ -0,0 +1,117 @@ +"""Live configuration of the two engines HyperSwap arbitrates between. + +Arbitration behaviour is largely dictated by settings that live outside this codebase. +Working out why a yield behaved the way it did meant reading journald and the ollama +unit by hand: OLLAMA_NUM_PARALLEL decides whether an unload queues behind a running +generation, OLLAMA_MAX_LOADED_MODELS decides whether more than one model can be +resident, and a pinned n_gpu_layers decides whether a model that will not fit spills to +the CPU or fails outright. Those are worth reading and explaining rather than hardcoding +into a dashboard subtitle that silently goes stale. +""" +import json +import logging +import subprocess +from typing import Any, Dict, List, Optional + +import httpx + +import vram_arbitrator + +logger = logging.getLogger("engines") + +# Settings that change how the arbitrator must behave, with what they imply. +OLLAMA_SETTING_NOTES = { + "OLLAMA_NUM_PARALLEL": ( + "Requests per model. At 1, a keep_alive:0 unload queues behind any running " + "generation and applies when it finishes — which is why a busy model is " + "reported as deferred rather than failed."), + "OLLAMA_MAX_LOADED_MODELS": ( + "How many models may be resident at once. At 1, Ollama evicts the previous " + "model on every swap."), + "OLLAMA_KEEP_ALIVE": ( + "Default residency after a request. Long values keep VRAM occupied and make " + "ComfyUI wait for an explicit yield."), + "OLLAMA_FLASH_ATTENTION": "FlashAttention kernels for attention.", + "OLLAMA_KV_CACHE_TYPE": "KV cache quantisation; smaller types cut VRAM per context.", + "OLLAMA_NUM_BATCH": "Prompt-evaluation batch size.", +} + + +def _ollama_unit_environment() -> Dict[str, str]: + """Read the ollama service's environment. Empty if it is not a systemd unit.""" + env: Dict[str, str] = {} + try: + proc = subprocess.run(["systemctl", "show", "ollama", "-p", "Environment", + "--value"], capture_output=True, text=True, timeout=8) + for token in proc.stdout.split(): + if "=" in token and token.startswith("OLLAMA"): + k, _, v = token.partition("=") + env[k] = v + except Exception as e: + logger.debug(f"could not read ollama unit environment: {e}") + return env + + +async def get_engine_config() -> Dict[str, Any]: + """Real, live configuration of both engines, with arbitration implications.""" + ollama_env = _ollama_unit_environment() + ollama_settings = [ + {"key": k, "value": v, "means": OLLAMA_SETTING_NOTES.get(k, "")} + for k, v in sorted(ollama_env.items()) + ] + + # A short, honest summary line to replace the dashboard's hardcoded subtitle. + feature_bits: List[str] = [] + if ollama_env.get("OLLAMA_FLASH_ATTENTION") == "1": + feature_bits.append("FlashAttention") + kv = ollama_env.get("OLLAMA_KV_CACHE_TYPE") + if kv: + feature_bits.append(f"{kv} KV cache") + host = ollama_env.get("OLLAMA_HOST", "") + port = host.rsplit(":", 1)[-1] if ":" in host else "11434" + + ollama = { + "port": port, + "settings": ollama_settings, + "summary": " + ".join(feature_bits) if feature_bits else "default configuration", + "max_loaded_models": ollama_env.get("OLLAMA_MAX_LOADED_MODELS"), + "num_parallel": ollama_env.get("OLLAMA_NUM_PARALLEL"), + "keep_alive": ollama_env.get("OLLAMA_KEEP_ALIVE"), + "config_source": "systemd unit environment" if ollama_env else "unavailable", + } + + comfy: Dict[str, Any] = {"online": False} + try: + async with httpx.AsyncClient(timeout=4.0) as c: + r = await c.get(f"{vram_arbitrator.COMFY_API_BASE}/system_stats") + if r.status_code == 200: + data = r.json() + system = data.get("system", {}) + argv = system.get("argv") or [] + devices = data.get("devices") or [] + dev = devices[0] if devices else {} + # The allocator is named in the device string; it is the closest thing + # ComfyUI reports to the "async offload" the old subtitle asserted. + dev_name = dev.get("name", "") + allocator = ("cudaMallocAsync" if "cudaMallocAsync" in dev_name + else "cudaMalloc" if "cudaMalloc" in dev_name else "unknown") + vram_flags = [a for a in argv + if a in ("--lowvram", "--novram", "--highvram", "--normalvram", + "--gpu-only", "--cpu")] + comfy = { + "online": True, + "version": system.get("comfyui_version"), + "pytorch": system.get("pytorch_version"), + # split()[0] on an absent version raises IndexError, which would have + # been swallowed by the except below and reported ComfyUI as offline. + "python": ((system.get("python_version") or "").split() or [None])[0], + "argv": argv, + "vram_mode": vram_flags[0] if vram_flags else "default (auto)", + "allocator": allocator, + "device": dev_name, + "summary": f"{allocator}, {vram_flags[0] if vram_flags else 'auto VRAM'}", + } + except Exception as e: + comfy = {"online": False, "error": str(e)[:120]} + + return {"ollama": ollama, "comfyui": comfy} diff --git a/mcp_server.py b/mcp_server.py index e5ac0f6..f5650a3 100644 --- a/mcp_server.py +++ b/mcp_server.py @@ -8,6 +8,7 @@ from typing import Dict, List, Any, Optional from mcp.server import MCPServer import autotune +import engines import health import overclock_manager import ram_optimizer @@ -137,6 +138,14 @@ def set_gpu_fan_speed(mode: str = "auto", percent: Optional[int] = None) -> str: res = overclock_manager.set_fan_auto() return json.dumps(res, indent=2) +@mcp.tool() +async def get_engine_config() -> str: + """Live configuration of Ollama and ComfyUI (parallelism, max loaded models, + keep-alive, KV cache type, ComfyUI VRAM mode and allocator), with what each setting + implies for VRAM arbitration.""" + return json.dumps(await engines.get_engine_config(), indent=2, default=str) + + @mcp.tool() async def check_system_health() -> str: """Check every dependency HyperSwap needs (NVML, sudo nvidia-smi, fan control via the diff --git a/server.py b/server.py index 3066288..db98b9a 100644 --- a/server.py +++ b/server.py @@ -16,6 +16,7 @@ from fastapi.middleware.cors import CORSMiddleware from pydantic import BaseModel, Field import autotune +import engines import health import overclock_manager import ram_optimizer @@ -321,6 +322,13 @@ async def api_health(): return await health.run_health_checks() +@app.get("/api/engines", summary="Live Engine Configuration", tags=["Telemetry"]) +async def api_engines(): + """Real configuration of Ollama and ComfyUI, with what each setting implies for + arbitration. These live outside this codebase but dictate how it must behave.""" + return await engines.get_engine_config() + + @app.get("/api/gpu", summary="GPU Sensors and VRAM Breakdown", tags=["Telemetry"]) async def get_gpu_metrics() -> Dict[str, Any]: """Detailed NVML sensors (utilization, temp, power, fan, clocks, throttle reasons, per-process VRAM).""" diff --git a/static/app.js b/static/app.js index 2ba0faf..6b854ce 100644 --- a/static/app.js +++ b/static/app.js @@ -1069,3 +1069,35 @@ async function fetchLastSwapFromStore() { document.addEventListener('DOMContentLoaded', () => { setTimeout(fetchLastSwapFromStore, 1500); }); + + +// ---------------------------------------------------------------- engine config + +async function fetchEngineConfig() { + // These subtitles used to be hardcoded. They happened to be accurate, which is worse + // than being wrong: they would have stayed accurate-looking after the settings changed. + try { + const d = await (await fetch('/api/engines')).json(); + const o = document.getElementById('ollama-engine-sub'); + if (o && d.ollama) { + const bits = [`Port :${d.ollama.port}`, d.ollama.summary]; + if (d.ollama.max_loaded_models) bits.push(`${d.ollama.max_loaded_models} model resident`); + if (d.ollama.keep_alive) bits.push(`keep-alive ${d.ollama.keep_alive}`); + o.textContent = bits.join(' // '); + o.title = (d.ollama.settings || []) + .filter(s => s.means) + .map(s => `${s.key}=${s.value} — ${s.means}`) + .join('\n'); + } + const c = document.getElementById('comfy-engine-sub'); + if (c && d.comfyui && d.comfyui.online) { + c.textContent = `Port :8188 // v${d.comfyui.version} // ${d.comfyui.summary}`; + c.title = `torch ${d.comfyui.pytorch}\n${d.comfyui.device || ''}`; + } + } catch (e) { /* subtitles are cosmetic; never break the page over them */ } +} + +document.addEventListener('DOMContentLoaded', () => { + fetchEngineConfig(); + setInterval(fetchEngineConfig, 120000); +}); diff --git a/static/index.html b/static/index.html index 97684a5..c639e56 100644 --- a/static/index.html +++ b/static/index.html @@ -194,7 +194,7 @@

Ollama LLM Engine

-

Port :11434 // FlashAttention + Q4 KV Cache

+

Port :11434