Treat a mid-generation LLM as busy, not as a failed yield
The persisted counters showed 19 timeouts in 20 yields. All 19 were one model, ornith-1.5:9b-cron, in two bursts at 06:25 and 06:33. Telemetry for that window shows the GPU pinned at 96-97% with Ollama holding 14.92 GB throughout: the model was mid-generation. Ollama will not unload a model that is inferencing, so every request failed, and with a 1s trigger debounce against a 10s blocking wait the arbitrator simply asked again, four times per burst, blocking the loop for 40s. Ollama's behaviour is correct. Ours was wrong in three ways. Busy is now a distinct outcome. _await_vram_release returns "released", "busy" or "stuck": VRAM that has not moved while the GPU is pinned means a generation is in flight, which is not a failure. With OLLAMA_NUM_PARALLEL=1 our keep_alive:0 request queues behind the running one and applies the moment it finishes, so the correct response is to stop waiting, not to retry. Only "stuck" -- VRAM held with an idle GPU -- is a real fault. The wait is short again (2s, from 10s) because blocking helps nobody: ComfyUI is not gated on our return value, and every blocked second stalls the watchdog and profile switching. Callers who genuinely want to wait out an inference can pass wait_for_generation=true. The unload POST itself now gets a 120s client timeout, since a 5s one could drop the connection before Ollama ever processed a request queued behind a long generation, losing the unload entirely. A busy model gets per-model backoff (5s, 15s, 30s, 60s) instead of being asked again every second, and a detached watcher confirms and logs the release when the generation ends, so the event log tells the whole story rather than stopping at "deferred". Measured: a mid-generation yield now returns busy in 610ms instead of blocking 10s, and the queued unload lands on its own 3s later when the generation completes. Counters are honest: yields (released), yield_deferred_busy, deferred_releases, yield_stalled. The old yield_timeouts conflated a healthy cron job with a fault and implied a 95% failure rate. Also adds a VRAM Arbitration panel to the dashboard. The arbitrator is the core of this application and its state was not displayed anywhere -- there was no way to see whether handoffs were working, which is why this went unnoticed until the persisted counters were read by hand. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -33,8 +33,9 @@ function initSSE() {
|
||||
function updateDashboard(data) {
|
||||
if (!data) return;
|
||||
|
||||
// Governor state rides along in the shared snapshot — no extra polling needed.
|
||||
// Governor and arbitration state ride along in the shared snapshot.
|
||||
if (data.governor) renderGovernor(data.governor);
|
||||
if (data.arbitrator) renderArbitrator(data.arbitrator);
|
||||
|
||||
// 1. GPU VRAM Stats
|
||||
const gpu = data.gpu || {};
|
||||
@@ -881,3 +882,36 @@ document.addEventListener('DOMContentLoaded', () => {
|
||||
if (d.last_result) renderSweep(d.last_result);
|
||||
}).catch(() => {});
|
||||
});
|
||||
|
||||
|
||||
// ---------------------------------------------------------------- arbitration
|
||||
|
||||
function renderArbitrator(arb) {
|
||||
const el = (id) => document.getElementById(id);
|
||||
if (!el('arb-action')) return;
|
||||
const c = arb.counters || {};
|
||||
|
||||
el('arb-action').textContent = arb.last_action || 'Idle';
|
||||
el('arb-yields').textContent = c.yields ?? 0;
|
||||
el('arb-busy').textContent = c.yield_deferred_busy ?? 0;
|
||||
el('arb-later').textContent = c.deferred_releases ?? 0;
|
||||
el('arb-stalled').textContent = c.yield_stalled ?? 0;
|
||||
el('arb-purges').textContent = c.purges ?? 0;
|
||||
el('arb-defpurge').textContent = c.deferred_purges ?? 0;
|
||||
|
||||
const ws = el('arb-ws');
|
||||
ws.textContent = arb.connected_ws ? 'ComfyUI WS live' : 'WS down — polling';
|
||||
ws.className = 'text-xs font-mono ' + (arb.connected_ws ? 'text-emerald-400' : 'text-amber-400');
|
||||
|
||||
// Show why we are holding off, and the idle countdown before ComfyUI is purged.
|
||||
const parts = [];
|
||||
const backoff = arb.yield_backoff || {};
|
||||
for (const [model, secs] of Object.entries(backoff)) {
|
||||
parts.push(`waiting ${secs}s before asking '${model}' again`);
|
||||
}
|
||||
if (arb.pending_purge && arb.comfy_idle_s != null) {
|
||||
const left = Math.max((arb.idle_purge_after_s || 0) - arb.comfy_idle_s, 0).toFixed(0);
|
||||
parts.push(`ComfyUI idle ${arb.comfy_idle_s}s — holding its checkpoints ${left}s longer`);
|
||||
}
|
||||
el('arb-backoff').textContent = parts.join(' · ');
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user