Read engine configuration live instead of asserting it in the dashboard

The engine subtitles were hardcoded: "FlashAttention + Q4 KV Cache" and "DynamicVRAM
+ Pinned Async Offload". The first turned out to be accurate -- OLLAMA_FLASH_ATTENTION
and OLLAMA_KV_CACHE_TYPE really are set -- which is worse than being wrong, because it
would have gone on looking accurate after the settings changed.

engines.py reads both engines' real configuration: the ollama service environment via
systemd, and ComfyUI's own /system_stats for version, allocator, VRAM mode and argv.
Exposed at GET /api/engines, as an MCP tool, and in the dashboard subtitles with the
full settings list as a tooltip.

The settings worth surfacing are the ones that dictate how this service must behave
and that previously had to be discovered by reading journald: OLLAMA_NUM_PARALLEL=1
is why an unload queues behind a running generation and is reported as deferred
rather than failed, and OLLAMA_MAX_LOADED_MODELS=1 is why every swap evicts the
previous model. Each is reported with that explanation attached.

Writing the tests found a bug in the new code: (system.get("python_version") or
"").split()[0] raises IndexError when ComfyUI omits the field, and the surrounding
except would have swallowed it and reported ComfyUI as entirely offline.

Tests: 192 (was 182).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
drjones
2026-09-06 10:56:04 -07:00
parent b53b026ced
commit 043d61722b
7 changed files with 315 additions and 3 deletions

View File

@@ -8,6 +8,7 @@ from typing import Dict, List, Any, Optional
from mcp.server import MCPServer
import autotune
import engines
import health
import overclock_manager
import ram_optimizer
@@ -137,6 +138,14 @@ def set_gpu_fan_speed(mode: str = "auto", percent: Optional[int] = None) -> str:
res = overclock_manager.set_fan_auto()
return json.dumps(res, indent=2)
@mcp.tool()
async def get_engine_config() -> str:
"""Live configuration of Ollama and ComfyUI (parallelism, max loaded models,
keep-alive, KV cache type, ComfyUI VRAM mode and allocator), with what each setting
implies for VRAM arbitration."""
return json.dumps(await engines.get_engine_config(), indent=2, default=str)
@mcp.tool()
async def check_system_health() -> str:
"""Check every dependency HyperSwap needs (NVML, sudo nvidia-smi, fan control via the