Treat a mid-generation LLM as busy, not as a failed yield

The persisted counters showed 19 timeouts in 20 yields. All 19 were one model,
ornith-1.5:9b-cron, in two bursts at 06:25 and 06:33. Telemetry for that window
shows the GPU pinned at 96-97% with Ollama holding 14.92 GB throughout: the model
was mid-generation. Ollama will not unload a model that is inferencing, so every
request failed, and with a 1s trigger debounce against a 10s blocking wait the
arbitrator simply asked again, four times per burst, blocking the loop for 40s.

Ollama's behaviour is correct. Ours was wrong in three ways.

Busy is now a distinct outcome. _await_vram_release returns "released", "busy" or
"stuck": VRAM that has not moved while the GPU is pinned means a generation is in
flight, which is not a failure. With OLLAMA_NUM_PARALLEL=1 our keep_alive:0 request
queues behind the running one and applies the moment it finishes, so the correct
response is to stop waiting, not to retry. Only "stuck" -- VRAM held with an idle
GPU -- is a real fault.

The wait is short again (2s, from 10s) because blocking helps nobody: ComfyUI is
not gated on our return value, and every blocked second stalls the watchdog and
profile switching. Callers who genuinely want to wait out an inference can pass
wait_for_generation=true. The unload POST itself now gets a 120s client timeout,
since a 5s one could drop the connection before Ollama ever processed a request
queued behind a long generation, losing the unload entirely.

A busy model gets per-model backoff (5s, 15s, 30s, 60s) instead of being asked
again every second, and a detached watcher confirms and logs the release when the
generation ends, so the event log tells the whole story rather than stopping at
"deferred". Measured: a mid-generation yield now returns busy in 610ms instead of
blocking 10s, and the queued unload lands on its own 3s later when the generation
completes.

Counters are honest: yields (released), yield_deferred_busy, deferred_releases,
yield_stalled. The old yield_timeouts conflated a healthy cron job with a fault
and implied a 95% failure rate.

Also adds a VRAM Arbitration panel to the dashboard. The arbitrator is the core of
this application and its state was not displayed anywhere -- there was no way to
see whether handoffs were working, which is why this went unnoticed until the
persisted counters were read by hand.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
drjones
2026-09-01 13:32:17 -07:00
parent 172d812820
commit bacaf50713
4 changed files with 291 additions and 43 deletions

View File

@@ -33,8 +33,9 @@ function initSSE() {
function updateDashboard(data) {
if (!data) return;
// Governor state rides along in the shared snapshot — no extra polling needed.
// Governor and arbitration state ride along in the shared snapshot.
if (data.governor) renderGovernor(data.governor);
if (data.arbitrator) renderArbitrator(data.arbitrator);
// 1. GPU VRAM Stats
const gpu = data.gpu || {};
@@ -881,3 +882,36 @@ document.addEventListener('DOMContentLoaded', () => {
if (d.last_result) renderSweep(d.last_result);
}).catch(() => {});
});
// ---------------------------------------------------------------- arbitration
function renderArbitrator(arb) {
const el = (id) => document.getElementById(id);
if (!el('arb-action')) return;
const c = arb.counters || {};
el('arb-action').textContent = arb.last_action || 'Idle';
el('arb-yields').textContent = c.yields ?? 0;
el('arb-busy').textContent = c.yield_deferred_busy ?? 0;
el('arb-later').textContent = c.deferred_releases ?? 0;
el('arb-stalled').textContent = c.yield_stalled ?? 0;
el('arb-purges').textContent = c.purges ?? 0;
el('arb-defpurge').textContent = c.deferred_purges ?? 0;
const ws = el('arb-ws');
ws.textContent = arb.connected_ws ? 'ComfyUI WS live' : 'WS down — polling';
ws.className = 'text-xs font-mono ' + (arb.connected_ws ? 'text-emerald-400' : 'text-amber-400');
// Show why we are holding off, and the idle countdown before ComfyUI is purged.
const parts = [];
const backoff = arb.yield_backoff || {};
for (const [model, secs] of Object.entries(backoff)) {
parts.push(`waiting ${secs}s before asking '${model}' again`);
}
if (arb.pending_purge && arb.comfy_idle_s != null) {
const left = Math.max((arb.idle_purge_after_s || 0) - arb.comfy_idle_s, 0).toFixed(0);
parts.push(`ComfyUI idle ${arb.comfy_idle_s}s — holding its checkpoints ${left}s longer`);
}
el('arb-backoff').textContent = parts.join(' · ');
}