Make unreclaimable VRAM actionable, and account for ComfyUI's CUDA context

The health check now reports what unmanaged VRAM actually costs rather than just
how much of it there is: "0.82 GB held by python (842 MB)" becomes "5 model(s) fit
within 15.42 GB but not the 14.60 GB actually available", naming them.

Getting that arithmetic right took a correction. The first version subtracted only
the desktop and the unmanaged process, and so reported a 14.93 GB model as fitting
against a real ceiling of 14.60 GB -- the same model the service had just refused
with 507. ComfyUI keeps a few hundred MB of CUDA context for as long as the process
lives, which a purge does not free, so it is not available either. The floor is taken
from the minimum ComfyUI VRAM in recent telemetry rather than its current value,
which could be a 7 GB checkpoint mid-generation.

The verifier's reclaim stage now re-runs a graph immediately beforehand to reset the
30 s idle window, since a large model takes longer than that to load and the purge
was freeing ComfyUI mid-load, so the reclaim path was never reached.

Tests: 199 (was 192). The new ones pin the ceiling arithmetic, including that a model
too large to fit on the card at all is not blamed on the third-party process.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
drjones
2026-09-06 17:44:21 -07:00
parent 81e5d88426
commit aeba1b47fd
3 changed files with 165 additions and 3 deletions

View File

@@ -215,6 +215,16 @@ async def stage_reclaim(c: httpx.AsyncClient, model: str) -> bool:
f"HyperSwap cannot free")
return True
# Re-run a graph first. The idle purge fires 30 s after ComfyUI goes quiet, and a
# large model takes longer than that to load -- so without resetting the timer the
# purge frees ComfyUI mid-load and the reclaim path is never reached.
sys.path.insert(0, "/home/drjones/unified-model-manager")
import autotune # noqa: E402
await autotune._diffusion_benchmark()
gpu = await api(c, "GET", "/api/gpu")
print(f" reset the idle window; ComfyUI holds "
f"{gpu['breakdown']['comfyui_gb']} GB, {gpu['vram_free_gb']} GB free")
res = await api(c, "POST", "/api/switch-model", allow_error=True,
json={"model": model, "keep_alive": "2m"}, timeout=600)
if res.get("_status") == 507: