Treat a mid-generation LLM as busy, not as a failed yield
The persisted counters showed 19 timeouts in 20 yields. All 19 were one model, ornith-1.5:9b-cron, in two bursts at 06:25 and 06:33. Telemetry for that window shows the GPU pinned at 96-97% with Ollama holding 14.92 GB throughout: the model was mid-generation. Ollama will not unload a model that is inferencing, so every request failed, and with a 1s trigger debounce against a 10s blocking wait the arbitrator simply asked again, four times per burst, blocking the loop for 40s. Ollama's behaviour is correct. Ours was wrong in three ways. Busy is now a distinct outcome. _await_vram_release returns "released", "busy" or "stuck": VRAM that has not moved while the GPU is pinned means a generation is in flight, which is not a failure. With OLLAMA_NUM_PARALLEL=1 our keep_alive:0 request queues behind the running one and applies the moment it finishes, so the correct response is to stop waiting, not to retry. Only "stuck" -- VRAM held with an idle GPU -- is a real fault. The wait is short again (2s, from 10s) because blocking helps nobody: ComfyUI is not gated on our return value, and every blocked second stalls the watchdog and profile switching. Callers who genuinely want to wait out an inference can pass wait_for_generation=true. The unload POST itself now gets a 120s client timeout, since a 5s one could drop the connection before Ollama ever processed a request queued behind a long generation, losing the unload entirely. A busy model gets per-model backoff (5s, 15s, 30s, 60s) instead of being asked again every second, and a detached watcher confirms and logs the release when the generation ends, so the event log tells the whole story rather than stopping at "deferred". Measured: a mid-generation yield now returns busy in 610ms instead of blocking 10s, and the queued unload lands on its own 3s later when the generation completes. Counters are honest: yields (released), yield_deferred_busy, deferred_releases, yield_stalled. The old yield_timeouts conflated a healthy cron job with a fault and implied a 95% failure rate. Also adds a VRAM Arbitration panel to the dashboard. The arbitrator is the core of this application and its state was not displayed anywhere -- there was no way to see whether handoffs were working, which is why this went unnoticed until the persisted counters were read by hand. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -668,6 +668,58 @@
|
||||
<!-- ============ NEXT-LEVEL PANELS: governor / residency / analytics / autotune ============ -->
|
||||
<div class="grid grid-cols-1 xl:grid-cols-2 gap-5 mt-5">
|
||||
|
||||
|
||||
<!-- VRAM Arbitration — the core handoff, previously invisible -->
|
||||
<div class="bg-slate-900/80 border border-slate-800 rounded-2xl p-5 xl:col-span-2">
|
||||
<div class="flex items-center justify-between pb-3 border-b border-slate-800">
|
||||
<div class="flex items-center space-x-2">
|
||||
<div class="p-2 rounded-lg bg-emerald-950/80 border border-emerald-800 text-emerald-400">
|
||||
<i class="fa-solid fa-right-left text-sm"></i>
|
||||
</div>
|
||||
<div>
|
||||
<h3 class="font-bold text-slate-100 text-sm">VRAM Arbitration</h3>
|
||||
<p class="text-xs text-slate-400">Who holds the GPU, and how handoffs are going</p>
|
||||
</div>
|
||||
</div>
|
||||
<span id="arb-ws" class="text-xs font-mono text-slate-500">—</span>
|
||||
</div>
|
||||
<div class="mt-4">
|
||||
<div id="arb-action" class="text-sm text-slate-200 bg-slate-950/60 border border-slate-800 rounded-lg px-3 py-2 mb-3 font-mono">Idle</div>
|
||||
<div id="arb-backoff" class="text-[11px] font-mono text-amber-400 mb-3"></div>
|
||||
<div class="grid grid-cols-3 sm:grid-cols-6 gap-2 text-center">
|
||||
<div class="bg-slate-950/60 rounded-lg p-2 border border-slate-800">
|
||||
<div id="arb-yields" class="text-lg font-bold text-emerald-400">0</div>
|
||||
<div class="text-[10px] text-slate-500 uppercase leading-tight">Released</div>
|
||||
</div>
|
||||
<div class="bg-slate-950/60 rounded-lg p-2 border border-slate-800">
|
||||
<div id="arb-busy" class="text-lg font-bold text-cyan-400">0</div>
|
||||
<div class="text-[10px] text-slate-500 uppercase leading-tight">Deferred<br>(busy)</div>
|
||||
</div>
|
||||
<div class="bg-slate-950/60 rounded-lg p-2 border border-slate-800">
|
||||
<div id="arb-later" class="text-lg font-bold text-cyan-400">0</div>
|
||||
<div class="text-[10px] text-slate-500 uppercase leading-tight">Landed<br>later</div>
|
||||
</div>
|
||||
<div class="bg-slate-950/60 rounded-lg p-2 border border-slate-800">
|
||||
<div id="arb-stalled" class="text-lg font-bold text-rose-400">0</div>
|
||||
<div class="text-[10px] text-slate-500 uppercase leading-tight">Stalled</div>
|
||||
</div>
|
||||
<div class="bg-slate-950/60 rounded-lg p-2 border border-slate-800">
|
||||
<div id="arb-purges" class="text-lg font-bold text-fuchsia-400">0</div>
|
||||
<div class="text-[10px] text-slate-500 uppercase leading-tight">Comfy<br>purges</div>
|
||||
</div>
|
||||
<div class="bg-slate-950/60 rounded-lg p-2 border border-slate-800">
|
||||
<div id="arb-defpurge" class="text-lg font-bold text-slate-300">0</div>
|
||||
<div class="text-[10px] text-slate-500 uppercase leading-tight">Purges<br>deferred</div>
|
||||
</div>
|
||||
</div>
|
||||
<p class="text-[11px] text-slate-500 mt-3">
|
||||
<span class="text-cyan-400">Deferred</span> is healthy — an LLM mid-generation cannot unload, so the
|
||||
request queues and applies the moment it finishes. Only <span class="text-rose-400">stalled</span>
|
||||
(VRAM held while the GPU sits idle) indicates a real problem.
|
||||
</p>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
<!-- Thermal Governor -->
|
||||
<div class="bg-slate-900/80 border border-slate-800 rounded-2xl p-5">
|
||||
<div class="flex items-center justify-between pb-3 border-b border-slate-800">
|
||||
|
||||
Reference in New Issue
Block a user