Add an end-to-end arbitration verifier; return honest HTTP status codes
The unit suite covers logic in isolation, but the promise this service exists to make -- an LLM and a diffusion pipeline sharing one 16 GB card without either failing -- had only ever been checked by hand, piecemeal. verify_arbitration.py walks the whole cycle against real hardware and reports what happened at each stage: load and its bandwidth classification, the confirmed yield, a real SDXL graph, the deferred idle purge, reclaim-and-retry, and finally whether VRAM attribution adds up and the reported GPU state still matches the card. It restores what it changes and refuses to start if ComfyUI is busy. Kept out of pytest deliberately: it moves real VRAM and takes minutes. Running it immediately found two bugs. /api/switch-model reported every upstream failure as 500. Asking an embedding model to generate makes Ollama return 400 -- the request is unusable, the service is fine -- and calling that an Internal Server Error blames this service for the caller's mistake. Failures now map to 400 for an upstream client error, 507 for a model that will not fit (valid request, healthy service, no room), and 502 when Ollama itself errors. The verifier also picked the smallest installed model, which here is nomic-embed-text -- an embedding model with no generate endpoint. It now filters those out by family and name. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
13
server.py
13
server.py
@@ -383,7 +383,18 @@ async def api_switch_model(req: SwitchRequest):
|
||||
await vram_arbitrator.arbitrator.request_vram_for_ollama()
|
||||
res = await vram_arbitrator.switch_ollama_model(req.model, keep_alive=req.keep_alive or "30m")
|
||||
if not res.get("success"):
|
||||
raise HTTPException(status_code=500, detail=res.get("error"))
|
||||
# Reflect what actually went wrong. Ollama returns 400 for an unusable request --
|
||||
# asking an embedding model to generate, say -- and reporting that as 500 blames
|
||||
# this service for the caller's mistake. A model that will not fit is neither:
|
||||
# the request is valid and the service is healthy, there is simply no room.
|
||||
upstream = res.get("upstream_status")
|
||||
if res.get("vram_oom"):
|
||||
status = 507 # Insufficient Storage
|
||||
elif isinstance(upstream, int) and 400 <= upstream < 500:
|
||||
status = 400
|
||||
else:
|
||||
status = 502 if upstream else 500
|
||||
raise HTTPException(status_code=status, detail=res.get("error"))
|
||||
return res
|
||||
|
||||
@app.post("/api/free-vram", summary="Soft-Yield Ollama VRAM", tags=["Orchestration"])
|
||||
|
||||
Reference in New Issue
Block a user