The persisted counters showed 19 timeouts in 20 yields. All 19 were one model,
ornith-1.5:9b-cron, in two bursts at 06:25 and 06:33. Telemetry for that window
shows the GPU pinned at 96-97% with Ollama holding 14.92 GB throughout: the model
was mid-generation. Ollama will not unload a model that is inferencing, so every
request failed, and with a 1s trigger debounce against a 10s blocking wait the
arbitrator simply asked again, four times per burst, blocking the loop for 40s.
Ollama's behaviour is correct. Ours was wrong in three ways.
Busy is now a distinct outcome. _await_vram_release returns "released", "busy" or
"stuck": VRAM that has not moved while the GPU is pinned means a generation is in
flight, which is not a failure. With OLLAMA_NUM_PARALLEL=1 our keep_alive:0 request
queues behind the running one and applies the moment it finishes, so the correct
response is to stop waiting, not to retry. Only "stuck" -- VRAM held with an idle
GPU -- is a real fault.
The wait is short again (2s, from 10s) because blocking helps nobody: ComfyUI is
not gated on our return value, and every blocked second stalls the watchdog and
profile switching. Callers who genuinely want to wait out an inference can pass
wait_for_generation=true. The unload POST itself now gets a 120s client timeout,
since a 5s one could drop the connection before Ollama ever processed a request
queued behind a long generation, losing the unload entirely.
A busy model gets per-model backoff (5s, 15s, 30s, 60s) instead of being asked
again every second, and a detached watcher confirms and logs the release when the
generation ends, so the event log tells the whole story rather than stopping at
"deferred". Measured: a mid-generation yield now returns busy in 610ms instead of
blocking 10s, and the queued unload lands on its own 3s later when the generation
completes.
Counters are honest: yields (released), yield_deferred_busy, deferred_releases,
yield_stalled. The old yield_timeouts conflated a healthy cron job with a fault
and implied a 95% failure rate.
Also adds a VRAM Arbitration panel to the dashboard. The arbitrator is the core of
this application and its state was not displayed anywhere -- there was no way to
see whether handoffs were working, which is why this went unnoticed until the
persisted counters were read by hand.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sweeping a knob the driver ignores measures nothing but benchmark noise, and the
tuner would then confidently report 'best = the highest value tried'. On this box
(driver 595.84) nvidia-settings accepts GPUGraphicsClockOffset/GPUMemoryTransferRate
Offset and silently discards them: assigning 0 reports success and reads back 250.
The ollama profile's core_offset_mhz=35 and mem_offset_mhz=200 have therefore been
doing nothing.
- _knob_effective() applies a probe value and confirms the hardware actually moved
before any sweep starts, choosing the candidate furthest from the current reading
(probing with the maximum fails when the card already sits at its top clock).
- Adds discrete clock-lock knobs (lock_mem_mhz, lock_core_max) driven by the card's
own supported-clock list, since -lmc/-lgc do work where offsets do not.
- Reasoning models return their output in 'thinking' with an empty 'response', which
the degeneracy check was flagging as corruption. Token count is now the primary
signal.
- Separates gain-vs-current-setting from gain-vs-slowest-value-tried. Reporting the
latter as 'gain vs baseline' implied a +102% speedup that nobody would observe.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Nine changes, in rough order of how much they affect real behaviour:
1. VRAM yield is now a barrier. Posting keep_alive:0 only asks Ollama to unload;
measured here, the HTTP call returns in 63ms while the driver takes a further
77ms to release 14.9GB. Returning inside that window is how ComfyUI ends up
allocating into VRAM that is still occupied. instant_free_ollama_vram() polls
NVML until the allocation is actually gone and reports request/confirm split.
2. ComfyUI VRAM is no longer purged 1.5s after every prompt, which forced a full
checkpoint reload on each workflow iteration. It is held for 30s of genuinely
empty queue, with an immediate purge when Ollama actually asks for the memory.
3. Cache-hit classification uses achieved bandwidth (size / load duration) rather
than a fixed `load_duration < 2500ms`. That constant called a 12.9GB model read
at 2.9GB/s a cold load, and a 0.5GB model read from NVMe a cache hit.
4. Page-cache residency is measured, not assumed. mincore(2) reported 128GB
resident on a box with 46GB of page cache: the kernel only permits page-cache
introspection on files you own, and the Ollama blobs are owned by uid ollama,
for which mincore answers "all resident" instead of failing. Uses cachestat(2)
where permitted and a randomised read-rate probe elsewhere, labelling which was
used. Fixed-offset probing was self-fulfilling, so windows are random and cold
ones are returned with FADV_DONTNEED.
5. Warming is budgeted and ranked by recency/frequency instead of reading every
file top-to-bottom, which on 64GB of RAM just evicts whatever was warmed first.
6. Telemetry and events persist to SQLite (~0.38 MB/hour) instead of living in a
50-entry in-memory deque, so /api/analytics/profiles can finally answer whether
an overclock profile actually delivers more tok/s.
7. Thermal governor walks the overclock back on sustained heat or hardware
throttling, with hysteresis, fed from the existing sampler.
8. Autotune sweeps a clock offset, benchmarks decode at each step, watches for Xid
errors and degenerate output, and restores the profile in a finally block.
9. Stock clocks/power/fans are restored on shutdown and via systemd ExecStopPost.
Nothing previously undid a locked clock or a manually pinned fan.
Also: one shared 1Hz telemetry sampler fanned out to SSE subscribers rather than
every client re-running the whole snapshot; wall-clock timestamps in place of the
event loop's monotonic clock; cached nvidia-smi shell-outs; quieter httpx logging.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>