Chasing why the reverse-direction reclaim never fired turned up something worse than the reclaim itself. The starvation check was never running. Instrumenting the watchdog showed busy=6, idle_check=0: every poll took the "ComfyUI is busy" branch. ComfyUI's /queue was reporting a WAN 2.1 i2v job in queue_running while the GPU sat at 0% and ComfyUI held 0.56 GB. The job was dead; ComfyUI had simply never cleared the row. Believing that flag meant this service thought ComfyUI was permanently busy, so it yielded the LLM's VRAM on every poll, never ran the idle purge, and never checked whether the LLM had been squeezed onto the CPU. One stale row disabled half of the arbitration, and it very likely explains the earlier burst of yields against a cron-driven model. A running entry is now corroborated before it is believed. The first attempt used GPU utilisation, which does not work: utilisation is shared with Ollama and with the third-party process on this box, so peak utilisation stayed above any sensible threshold and a stuck entry never looked stale. ComfyUI's own VRAM is the right signal -- a real diffusion job loads gigabytes of checkpoint, a dead one holds only its CUDA context. After the fix the same watchdog reports busy=3, idle_check=32. Every early return in the starvation check now records why it bailed, because with four of them there was no way to tell which had fired. /api/health reports a stale queue entry with its impact and how to clear it. Also confirmed, contradicting an earlier conclusion in this branch: Ollama on this box *does* spill to the CPU. smtek/Qwen3.8-27B:Q2_K_XL held steady at 29.2% on GPU (size=15.59 GB, size_vram=4.56 GB) across twelve seconds of polling -- a stable placement, not a progressive load. Both failure modes are real; which one occurs depends on the model. Tests: 206 (was 199). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
38 lines
1.5 KiB
JSON
38 lines
1.5 KiB
JSON
{
|
|
"ollama": {
|
|
"label": "Ollama \u2014 LLM decode (memory-bandwidth bound; measured insensitive to power and clocks)",
|
|
"power_limit_w": 320,
|
|
"core_offset_mhz": 0,
|
|
"mem_offset_mhz": 0,
|
|
"lock_core_min": 0,
|
|
"lock_core_max": 0,
|
|
"lock_mem_mhz": 0,
|
|
"fan_mode": "auto",
|
|
"fan_speed_pct": 0,
|
|
"measured": "73.0-73.5 tok/s flat from 222W to 370W (qwen3.8long, 2026-08-28). Actual draw never exceeded 224W at any limit. Memory clock lock made no difference (72.6 locked vs 72.7 unlocked)."
|
|
},
|
|
"comfy": {
|
|
"label": "ComfyUI \u2014 diffusion (compute bound; genuinely power-scaling)",
|
|
"power_limit_w": 370,
|
|
"core_offset_mhz": 0,
|
|
"mem_offset_mhz": 0,
|
|
"lock_core_min": 0,
|
|
"lock_core_max": 0,
|
|
"lock_mem_mhz": 0,
|
|
"fan_mode": "auto",
|
|
"fan_speed_pct": 0,
|
|
"measured": "SDXL 1024/20-step: 5.48 it/s @222W, 6.22 @259W, 6.50 @296W, 6.52 @320W, 6.63 @333W, 6.71 @370W (2026-08-28). Worth +2.8% over the 320W stock default. Core clock lock made no difference across 2400-3105 MHz."
|
|
},
|
|
"balanced": {
|
|
"label": "Balanced \u2014 stock power and boost, automatic fans",
|
|
"power_limit_w": 340,
|
|
"core_offset_mhz": 10,
|
|
"mem_offset_mhz": 150,
|
|
"lock_core_min": 0,
|
|
"lock_core_max": 0,
|
|
"lock_mem_mhz": 0,
|
|
"fan_mode": "manual",
|
|
"fan_speed_pct": 95,
|
|
"measured": "Card's own design point. 48k telemetry samples show 67.8C average under load at 39.5% auto fan, 81C all-time max, zero thermal throttle events."
|
|
}
|
|
} |