Stop a stale ComfyUI queue entry from disabling half the arbitration

Chasing why the reverse-direction reclaim never fired turned up something worse than
the reclaim itself.

The starvation check was never running. Instrumenting the watchdog showed busy=6,
idle_check=0: every poll took the "ComfyUI is busy" branch. ComfyUI's /queue was
reporting a WAN 2.1 i2v job in queue_running while the GPU sat at 0% and ComfyUI held
0.56 GB. The job was dead; ComfyUI had simply never cleared the row.

Believing that flag meant this service thought ComfyUI was permanently busy, so it
yielded the LLM's VRAM on every poll, never ran the idle purge, and never checked
whether the LLM had been squeezed onto the CPU. One stale row disabled half of the
arbitration, and it very likely explains the earlier burst of yields against a
cron-driven model.

A running entry is now corroborated before it is believed. The first attempt used GPU
utilisation, which does not work: utilisation is shared with Ollama and with the
third-party process on this box, so peak utilisation stayed above any sensible
threshold and a stuck entry never looked stale. ComfyUI's own VRAM is the right
signal -- a real diffusion job loads gigabytes of checkpoint, a dead one holds only
its CUDA context. After the fix the same watchdog reports busy=3, idle_check=32.

Every early return in the starvation check now records why it bailed, because with
four of them there was no way to tell which had fired. /api/health reports a stale
queue entry with its impact and how to clear it.

Also confirmed, contradicting an earlier conclusion in this branch: Ollama on this box
*does* spill to the CPU. smtek/Qwen3.8-27B:Q2_K_XL held steady at 29.2% on GPU
(size=15.59 GB, size_vram=4.56 GB) across twelve seconds of polling -- a stable
placement, not a progressive load. Both failure modes are real; which one occurs
depends on the model.

Tests: 206 (was 199).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
drjones
2026-09-07 08:48:37 -07:00
parent aeba1b47fd
commit c3d9b36035
6 changed files with 218 additions and 24 deletions

View File

@@ -13,6 +13,7 @@ were observed on real hardware before being encoded here:
out of memory ... unable to allocate CUDA0 buffer"
"""
import asyncio
import time
import pytest
@@ -132,3 +133,70 @@ class TestBusyBackoff:
assert "yield_deferred_busy" in arb.stats
assert "yield_stalled" in arb.stats
assert "yield_timeouts" not in arb.stats
class TestComfyStaleQueueDetection:
"""ComfyUI can leave a dead job in queue_running forever.
Observed on this machine: a WAN 2.1 i2v entry sat in queue_running while the GPU was
idle and ComfyUI held 0.56 GB. Trusting that flag made the watchdog believe ComfyUI
was permanently busy, so it evicted the LLM on every poll, never ran the idle purge,
and never checked whether the LLM had been pushed onto the CPU. Instrumenting the
watchdog showed busy=6, idle_check=0 -- one stale row had disabled half the logic.
"""
def _arb(self, comfy_bytes):
arb = v.AutoArbitrator()
v.get_process_vram_bytes = lambda: {
"ollama_bytes": 0, "comfyui_bytes": int(comfy_bytes), "other_bytes": 0,
"desktop_bytes": 0, "unmanaged_bytes": 0, "free_bytes": 0, "gpu_util_pct": 0}
return arb
def teardown_method(self):
import importlib
importlib.reload(v)
def test_empty_queue_is_not_busy(self):
arb = self._arb(0)
assert arb._comfy_genuinely_busy({"queue_running": [], "queue_pending": []}) is False
def test_pending_work_is_always_busy(self):
arb = self._arb(0)
assert arb._comfy_genuinely_busy(
{"queue_running": [], "queue_pending": [[1, "p"]]}) is True
def test_a_running_job_is_believed_at_first(self):
# It must not be called stale before it has had time to load anything.
arb = self._arb(0.1 * GB)
assert arb._comfy_genuinely_busy(
{"queue_running": [[1, "abc"]], "queue_pending": []}) is True
def test_long_running_job_holding_no_vram_is_stale(self):
arb = self._arb(0.56 * GB) # the observed CUDA-context floor
q = {"queue_running": [[1, "abc"]], "queue_pending": []}
arb._comfy_genuinely_busy(q)
arb._running_since = time.time() - (arb.STALE_RUNNING_S + 5)
assert arb._comfy_genuinely_busy(q) is False
assert arb.comfy_stale_job == "abc"
def test_long_running_job_holding_a_checkpoint_is_real(self):
# 6.8 GB is a loaded SDXL checkpoint; slow is not the same as stuck.
arb = self._arb(6.8 * GB)
q = {"queue_running": [[1, "abc"]], "queue_pending": []}
arb._comfy_genuinely_busy(q)
arb._running_since = time.time() - (arb.STALE_RUNNING_S + 5)
assert arb._comfy_genuinely_busy(q) is True
assert arb.comfy_stale_job is None
def test_a_new_prompt_id_resets_the_staleness_clock(self):
arb = self._arb(0.5 * GB)
arb._comfy_genuinely_busy({"queue_running": [[1, "old"]], "queue_pending": []})
arb._running_since = time.time() - 1000
assert arb._comfy_genuinely_busy(
{"queue_running": [[1, "new"]], "queue_pending": []}) is True
def test_vram_not_utilisation_is_the_signal(self):
# Utilisation is shared with Ollama and any third-party process, so it stayed
# above every sensible threshold and a stuck entry never looked stale.
assert hasattr(v.AutoArbitrator, "STALE_COMFY_BYTES")
assert not hasattr(v.AutoArbitrator, "STALE_UTIL_PCT")