Treat a mid-generation LLM as busy, not as a failed yield

The persisted counters showed 19 timeouts in 20 yields. All 19 were one model,
ornith-1.5:9b-cron, in two bursts at 06:25 and 06:33. Telemetry for that window
shows the GPU pinned at 96-97% with Ollama holding 14.92 GB throughout: the model
was mid-generation. Ollama will not unload a model that is inferencing, so every
request failed, and with a 1s trigger debounce against a 10s blocking wait the
arbitrator simply asked again, four times per burst, blocking the loop for 40s.

Ollama's behaviour is correct. Ours was wrong in three ways.

Busy is now a distinct outcome. _await_vram_release returns "released", "busy" or
"stuck": VRAM that has not moved while the GPU is pinned means a generation is in
flight, which is not a failure. With OLLAMA_NUM_PARALLEL=1 our keep_alive:0 request
queues behind the running one and applies the moment it finishes, so the correct
response is to stop waiting, not to retry. Only "stuck" -- VRAM held with an idle
GPU -- is a real fault.

The wait is short again (2s, from 10s) because blocking helps nobody: ComfyUI is
not gated on our return value, and every blocked second stalls the watchdog and
profile switching. Callers who genuinely want to wait out an inference can pass
wait_for_generation=true. The unload POST itself now gets a 120s client timeout,
since a 5s one could drop the connection before Ollama ever processed a request
queued behind a long generation, losing the unload entirely.

A busy model gets per-model backoff (5s, 15s, 30s, 60s) instead of being asked
again every second, and a detached watcher confirms and logs the release when the
generation ends, so the event log tells the whole story rather than stopping at
"deferred". Measured: a mid-generation yield now returns busy in 610ms instead of
blocking 10s, and the queued unload lands on its own 3s later when the generation
completes.

Counters are honest: yields (released), yield_deferred_busy, deferred_releases,
yield_stalled. The old yield_timeouts conflated a healthy cron job with a fault
and implied a 95% failure rate.

Also adds a VRAM Arbitration panel to the dashboard. The arbitrator is the core of
this application and its state was not displayed anywhere -- there was no way to
see whether handoffs were working, which is why this went unnoticed until the
persisted counters were read by hand.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
drjones
2026-09-01 13:32:17 -07:00
parent 172d812820
commit bacaf50713
4 changed files with 291 additions and 43 deletions

View File

@@ -668,6 +668,58 @@
<!-- ============ NEXT-LEVEL PANELS: governor / residency / analytics / autotune ============ -->
<div class="grid grid-cols-1 xl:grid-cols-2 gap-5 mt-5">
<!-- VRAM Arbitration — the core handoff, previously invisible -->
<div class="bg-slate-900/80 border border-slate-800 rounded-2xl p-5 xl:col-span-2">
<div class="flex items-center justify-between pb-3 border-b border-slate-800">
<div class="flex items-center space-x-2">
<div class="p-2 rounded-lg bg-emerald-950/80 border border-emerald-800 text-emerald-400">
<i class="fa-solid fa-right-left text-sm"></i>
</div>
<div>
<h3 class="font-bold text-slate-100 text-sm">VRAM Arbitration</h3>
<p class="text-xs text-slate-400">Who holds the GPU, and how handoffs are going</p>
</div>
</div>
<span id="arb-ws" class="text-xs font-mono text-slate-500">—</span>
</div>
<div class="mt-4">
<div id="arb-action" class="text-sm text-slate-200 bg-slate-950/60 border border-slate-800 rounded-lg px-3 py-2 mb-3 font-mono">Idle</div>
<div id="arb-backoff" class="text-[11px] font-mono text-amber-400 mb-3"></div>
<div class="grid grid-cols-3 sm:grid-cols-6 gap-2 text-center">
<div class="bg-slate-950/60 rounded-lg p-2 border border-slate-800">
<div id="arb-yields" class="text-lg font-bold text-emerald-400">0</div>
<div class="text-[10px] text-slate-500 uppercase leading-tight">Released</div>
</div>
<div class="bg-slate-950/60 rounded-lg p-2 border border-slate-800">
<div id="arb-busy" class="text-lg font-bold text-cyan-400">0</div>
<div class="text-[10px] text-slate-500 uppercase leading-tight">Deferred<br>(busy)</div>
</div>
<div class="bg-slate-950/60 rounded-lg p-2 border border-slate-800">
<div id="arb-later" class="text-lg font-bold text-cyan-400">0</div>
<div class="text-[10px] text-slate-500 uppercase leading-tight">Landed<br>later</div>
</div>
<div class="bg-slate-950/60 rounded-lg p-2 border border-slate-800">
<div id="arb-stalled" class="text-lg font-bold text-rose-400">0</div>
<div class="text-[10px] text-slate-500 uppercase leading-tight">Stalled</div>
</div>
<div class="bg-slate-950/60 rounded-lg p-2 border border-slate-800">
<div id="arb-purges" class="text-lg font-bold text-fuchsia-400">0</div>
<div class="text-[10px] text-slate-500 uppercase leading-tight">Comfy<br>purges</div>
</div>
<div class="bg-slate-950/60 rounded-lg p-2 border border-slate-800">
<div id="arb-defpurge" class="text-lg font-bold text-slate-300">0</div>
<div class="text-[10px] text-slate-500 uppercase leading-tight">Purges<br>deferred</div>
</div>
</div>
<p class="text-[11px] text-slate-500 mt-3">
<span class="text-cyan-400">Deferred</span> is healthy — an LLM mid-generation cannot unload, so the
request queues and applies the moment it finishes. Only <span class="text-rose-400">stalled</span>
(VRAM held while the GPU sits idle) indicates a real problem.
</p>
</div>
</div>
<!-- Thermal Governor -->
<div class="bg-slate-900/80 border border-slate-800 rounded-2xl p-5">
<div class="flex items-center justify-between pb-3 border-b border-slate-800">