The unit suite covers logic in isolation, but the promise this service exists to make -- an LLM and a diffusion pipeline sharing one 16 GB card without either failing -- had only ever been checked by hand, piecemeal. verify_arbitration.py walks the whole cycle against real hardware and reports what happened at each stage: load and its bandwidth classification, the confirmed yield, a real SDXL graph, the deferred idle purge, reclaim-and-retry, and finally whether VRAM attribution adds up and the reported GPU state still matches the card. It restores what it changes and refuses to start if ComfyUI is busy. Kept out of pytest deliberately: it moves real VRAM and takes minutes. Running it immediately found two bugs. /api/switch-model reported every upstream failure as 500. Asking an embedding model to generate makes Ollama return 400 -- the request is unusable, the service is fine -- and calling that an Internal Server Error blames this service for the caller's mistake. Failures now map to 400 for an upstream client error, 507 for a model that will not fit (valid request, healthy service, no room), and 502 when Ollama itself errors. The verifier also picked the smallest installed model, which here is nomic-embed-text -- an embedding model with no generate endpoint. It now filters those out by family and name. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
28 KiB
28 KiB