The unit suite covers logic in isolation, but the promise this service exists to make
-- an LLM and a diffusion pipeline sharing one 16 GB card without either failing --
had only ever been checked by hand, piecemeal. verify_arbitration.py walks the whole
cycle against real hardware and reports what happened at each stage: load and its
bandwidth classification, the confirmed yield, a real SDXL graph, the deferred idle
purge, reclaim-and-retry, and finally whether VRAM attribution adds up and the
reported GPU state still matches the card. It restores what it changes and refuses
to start if ComfyUI is busy. Kept out of pytest deliberately: it moves real VRAM and
takes minutes.
Running it immediately found two bugs.
/api/switch-model reported every upstream failure as 500. Asking an embedding model
to generate makes Ollama return 400 -- the request is unusable, the service is fine
-- and calling that an Internal Server Error blames this service for the caller's
mistake. Failures now map to 400 for an upstream client error, 507 for a model that
will not fit (valid request, healthy service, no room), and 502 when Ollama itself
errors.
The verifier also picked the smallest installed model, which here is
nomic-embed-text -- an embedding model with no generate endpoint. It now filters
those out by family and name.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>