Add an end-to-end arbitration verifier; return honest HTTP status codes

The unit suite covers logic in isolation, but the promise this service exists to make
-- an LLM and a diffusion pipeline sharing one 16 GB card without either failing --
had only ever been checked by hand, piecemeal. verify_arbitration.py walks the whole
cycle against real hardware and reports what happened at each stage: load and its
bandwidth classification, the confirmed yield, a real SDXL graph, the deferred idle
purge, reclaim-and-retry, and finally whether VRAM attribution adds up and the
reported GPU state still matches the card. It restores what it changes and refuses
to start if ComfyUI is busy. Kept out of pytest deliberately: it moves real VRAM and
takes minutes.

Running it immediately found two bugs.

/api/switch-model reported every upstream failure as 500. Asking an embedding model
to generate makes Ollama return 400 -- the request is unusable, the service is fine
-- and calling that an Internal Server Error blames this service for the caller's
mistake. Failures now map to 400 for an upstream client error, 507 for a model that
will not fit (valid request, healthy service, no room), and 502 when Ollama itself
errors.

The verifier also picked the smallest installed model, which here is
nomic-embed-text -- an embedding model with no generate endpoint. It now filters
those out by family and name.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
drjones
2026-09-06 17:25:42 -07:00
parent 043d61722b
commit aed1c360f0
3 changed files with 264 additions and 1 deletions

View File

@@ -383,7 +383,18 @@ async def api_switch_model(req: SwitchRequest):
await vram_arbitrator.arbitrator.request_vram_for_ollama()
res = await vram_arbitrator.switch_ollama_model(req.model, keep_alive=req.keep_alive or "30m")
if not res.get("success"):
raise HTTPException(status_code=500, detail=res.get("error"))
# Reflect what actually went wrong. Ollama returns 400 for an unusable request --
# asking an embedding model to generate, say -- and reporting that as 500 blames
# this service for the caller's mistake. A model that will not fit is neither:
# the request is valid and the service is healthy, there is simply no room.
upstream = res.get("upstream_status")
if res.get("vram_oom"):
status = 507 # Insufficient Storage
elif isinstance(upstream, int) and 400 <= upstream < 500:
status = 400
else:
status = 502 if upstream else 500
raise HTTPException(status_code=status, detail=res.get("error"))
return res
@app.post("/api/free-vram", summary="Soft-Yield Ollama VRAM", tags=["Orchestration"])