Read engine configuration live instead of asserting it in the dashboard

The engine subtitles were hardcoded: "FlashAttention + Q4 KV Cache" and "DynamicVRAM
+ Pinned Async Offload". The first turned out to be accurate -- OLLAMA_FLASH_ATTENTION
and OLLAMA_KV_CACHE_TYPE really are set -- which is worse than being wrong, because it
would have gone on looking accurate after the settings changed.

engines.py reads both engines' real configuration: the ollama service environment via
systemd, and ComfyUI's own /system_stats for version, allocator, VRAM mode and argv.
Exposed at GET /api/engines, as an MCP tool, and in the dashboard subtitles with the
full settings list as a tooltip.

The settings worth surfacing are the ones that dictate how this service must behave
and that previously had to be discovered by reading journald: OLLAMA_NUM_PARALLEL=1
is why an unload queues behind a running generation and is reported as deferred
rather than failed, and OLLAMA_MAX_LOADED_MODELS=1 is why every swap evicts the
previous model. Each is reported with that explanation attached.

Writing the tests found a bug in the new code: (system.get("python_version") or
"").split()[0] raises IndexError when ComfyUI omits the field, and the surrounding
except would have swallowed it and reported ComfyUI as entirely offline.

Tests: 192 (was 182).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
drjones
2026-09-06 10:56:04 -07:00
parent b53b026ced
commit 043d61722b
7 changed files with 315 additions and 3 deletions

View File

@@ -16,6 +16,7 @@ from fastapi.middleware.cors import CORSMiddleware
from pydantic import BaseModel, Field
import autotune
import engines
import health
import overclock_manager
import ram_optimizer
@@ -321,6 +322,13 @@ async def api_health():
return await health.run_health_checks()
@app.get("/api/engines", summary="Live Engine Configuration", tags=["Telemetry"])
async def api_engines():
"""Real configuration of Ollama and ComfyUI, with what each setting implies for
arbitration. These live outside this codebase but dictate how it must behave."""
return await engines.get_engine_config()
@app.get("/api/gpu", summary="GPU Sensors and VRAM Breakdown", tags=["Telemetry"])
async def get_gpu_metrics() -> Dict[str, Any]:
"""Detailed NVML sensors (utilization, temp, power, fan, clocks, throttle reasons, per-process VRAM)."""