The API claimed the card was at 320W while nvidia-smi reported 370W. Three
separate defects, all introduced by me in this branch.
The readback was stale. get_gpu_state()/get_fan_status() gained a 2s cache so the
dashboard's polling would stop forking sudo every few seconds, but apply_profile
read back through that cache and its invalidation ran afterwards. A profile that
had just moved the card 370W -> 320W therefore returned a payload whose detail
string said "set to 320.00 W from 370.00 W" next to a power_limit_w of 370.0.
Caches are now cleared before the readback, which is forced.
Fan control could fail for an entire session. On boot this unit can start before
the headless X server on :8 that owns the GPU accepts connections, and the fan
assignment fails with "Error resolving target specification 'gpu:0'". Nothing
retried and nothing surfaced it, so the fans were left unconfigured with the
failure visible only inside one log line. apply_fan_control now recognises that
specific error and retries up to 5 times.
Nothing verified the result. ACTIVE_PROFILE defaults to "balanced" at import,
which is indistinguishable from "balanced was successfully applied" -- so a failed
startup apply left the app confidently reporting a profile it had never put on the
hardware. apply_profile now returns a `verified` block comparing intent against
readback and logs a warning on mismatch; profile_drift() exposes the comparison
plus whether any profile has actually been applied since startup; and the 1Hz
sampler calls reconcile_profile() once a minute to re-apply on drift.
Verified by setting 370W externally behind the service's back: the drift was
reported immediately and corrected automatically 40s later.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The profiles were hand-written and had never been checked against the hardware. Adding
a ComfyUI benchmark alongside the existing decode one made the compute side measurable
for the first time, and most of what the profiles configured turned out to do nothing.
Measured on this card (RTX 4080 SUPER, driver 595.84):
- LLM decode is not power-bound: 73.0-73.5 tok/s flat from 222W to 370W, with the card
never drawing more than 224W at any limit. The ollama profile's 370W did nothing.
- Diffusion is power-bound: 5.48 it/s @222W rising to 6.71 @370W, so comfy's 370W is
worth a real +2.8% over the 320W stock default.
- Clock locks did nothing for either workload: 72.6 tok/s locked at 11251MHz vs 72.7
unlocked; 6.77 it/s locked at 3105MHz vs 6.73 unlocked, and 6.78 at 2400MHz.
- Memory bandwidth is still the decode bottleneck (5001MHz halves throughput to 35.9
tok/s), confirming the profile's premise -- the card just gets there unaided.
- Fans: 48,435 samples show 81C all-time max and zero thermal throttle events, while
the ollama profile held 49.6C average by running fans at 87%. All profiles now use
automatic fans and let the thermal governor escalate on demand.
Code changes supporting that:
- _diffusion_benchmark() queues a fixed SDXL graph via ComfyUI's API. The seed must
vary per run: ComfyUI caches by node inputs, so a fixed seed returned in ~1ms without
executing. Implausibly fast results are now rejected as cache hits rather than
recorded as record scores.
- The arbitrator's automatic profile switching is suspended during a sweep. A diffusion
benchmark trips trigger_comfy_priority, which reapplies the whole profile and would
silently overwrite the clock being measured.
- _supported_clocks() queries the mem,gr pair; asking for a single field returned one
column and reading index 1 yielded an empty list rather than an error. Graphics clocks
are subsampled (the card enumerates 194 of them) and lock sweeps include an explicit
unlocked control step.
- offsets_supported() probes once and apply_profile skips inert offset levers with an
explanation instead of pretending they applied.
- Profiles carry a 'measured' field recording the evidence behind each setting.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Nine changes, in rough order of how much they affect real behaviour:
1. VRAM yield is now a barrier. Posting keep_alive:0 only asks Ollama to unload;
measured here, the HTTP call returns in 63ms while the driver takes a further
77ms to release 14.9GB. Returning inside that window is how ComfyUI ends up
allocating into VRAM that is still occupied. instant_free_ollama_vram() polls
NVML until the allocation is actually gone and reports request/confirm split.
2. ComfyUI VRAM is no longer purged 1.5s after every prompt, which forced a full
checkpoint reload on each workflow iteration. It is held for 30s of genuinely
empty queue, with an immediate purge when Ollama actually asks for the memory.
3. Cache-hit classification uses achieved bandwidth (size / load duration) rather
than a fixed `load_duration < 2500ms`. That constant called a 12.9GB model read
at 2.9GB/s a cold load, and a 0.5GB model read from NVMe a cache hit.
4. Page-cache residency is measured, not assumed. mincore(2) reported 128GB
resident on a box with 46GB of page cache: the kernel only permits page-cache
introspection on files you own, and the Ollama blobs are owned by uid ollama,
for which mincore answers "all resident" instead of failing. Uses cachestat(2)
where permitted and a randomised read-rate probe elsewhere, labelling which was
used. Fixed-offset probing was self-fulfilling, so windows are random and cold
ones are returned with FADV_DONTNEED.
5. Warming is budgeted and ranked by recency/frequency instead of reading every
file top-to-bottom, which on 64GB of RAM just evicts whatever was warmed first.
6. Telemetry and events persist to SQLite (~0.38 MB/hour) instead of living in a
50-entry in-memory deque, so /api/analytics/profiles can finally answer whether
an overclock profile actually delivers more tok/s.
7. Thermal governor walks the overclock back on sustained heat or hardware
throttling, with hysteresis, fed from the existing sampler.
8. Autotune sweeps a clock offset, benchmarks decode at each step, watches for Xid
errors and degenerate output, and restores the profile in a finally block.
9. Stock clocks/power/fans are restored on shutdown and via systemd ExecStopPost.
Nothing previously undid a locked clock or a manually pinned fan.
Also: one shared 1Hz telemetry sampler fanned out to SSE subscribers rather than
every client re-running the whole snapshot; wall-clock timestamps in place of the
event loop's monotonic clock; cached nvidia-smi shell-outs; quieter httpx logging.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>