Read engine configuration live instead of asserting it in the dashboard

The engine subtitles were hardcoded: "FlashAttention + Q4 KV Cache" and "DynamicVRAM
+ Pinned Async Offload". The first turned out to be accurate -- OLLAMA_FLASH_ATTENTION
and OLLAMA_KV_CACHE_TYPE really are set -- which is worse than being wrong, because it
would have gone on looking accurate after the settings changed.

engines.py reads both engines' real configuration: the ollama service environment via
systemd, and ComfyUI's own /system_stats for version, allocator, VRAM mode and argv.
Exposed at GET /api/engines, as an MCP tool, and in the dashboard subtitles with the
full settings list as a tooltip.

The settings worth surfacing are the ones that dictate how this service must behave
and that previously had to be discovered by reading journald: OLLAMA_NUM_PARALLEL=1
is why an unload queues behind a running generation and is reported as deferred
rather than failed, and OLLAMA_MAX_LOADED_MODELS=1 is why every swap evicts the
previous model. Each is reported with that explanation attached.

Writing the tests found a bug in the new code: (system.get("python_version") or
"").split()[0] raises IndexError when ComfyUI omits the field, and the surrounding
except would have swallowed it and reported ComfyUI as entirely offline.

Tests: 192 (was 182).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
drjones
2026-09-06 10:56:04 -07:00
parent b53b026ced
commit 043d61722b
7 changed files with 315 additions and 3 deletions

View File

@@ -1069,3 +1069,35 @@ async function fetchLastSwapFromStore() {
document.addEventListener('DOMContentLoaded', () => {
setTimeout(fetchLastSwapFromStore, 1500);
});
// ---------------------------------------------------------------- engine config
async function fetchEngineConfig() {
// These subtitles used to be hardcoded. They happened to be accurate, which is worse
// than being wrong: they would have stayed accurate-looking after the settings changed.
try {
const d = await (await fetch('/api/engines')).json();
const o = document.getElementById('ollama-engine-sub');
if (o && d.ollama) {
const bits = [`Port :${d.ollama.port}`, d.ollama.summary];
if (d.ollama.max_loaded_models) bits.push(`${d.ollama.max_loaded_models} model resident`);
if (d.ollama.keep_alive) bits.push(`keep-alive ${d.ollama.keep_alive}`);
o.textContent = bits.join(' // ');
o.title = (d.ollama.settings || [])
.filter(s => s.means)
.map(s => `${s.key}=${s.value} — ${s.means}`)
.join('\n');
}
const c = document.getElementById('comfy-engine-sub');
if (c && d.comfyui && d.comfyui.online) {
c.textContent = `Port :8188 // v${d.comfyui.version} // ${d.comfyui.summary}`;
c.title = `torch ${d.comfyui.pytorch}\n${d.comfyui.device || ''}`;
}
} catch (e) { /* subtitles are cosmetic; never break the page over them */ }
}
document.addEventListener('DOMContentLoaded', () => {
fetchEngineConfig();
setInterval(fetchEngineConfig, 120000);
});