_check_ollama_starved and _arbitrate were solving the same problem, one of them hardcoded to two applications. The Ollama-specific version is gone and the watchdog calls only the generic loop. The OOM retry inside switch_ollama_model no longer purges ComfyUI by name either: it asks plan_release which tenant should give up memory, so a third application can be the one that yields, and when the reclaim is not enough the response names the blockers instead of implying ComfyUI was at fault. The dashboard showed exactly two engines, which no longer matched what the service does. A GPU Tenants panel lists every configured application ordered by priority -- VRAM held, whether it is working, whether it can be reclaimed at all, and how much it needs -- along with the last arbitration decision and why it could or could not be satisfied. That panel did not appear at first, and the reason is worth fixing rather than working around: the browser kept serving a cached app.js despite the ETag, so a reload ran the old dashboard against the new API. Assets are now stamped with their mtime, so a changed file is always a different URL. Anyone updating this service would have hit the same thing. Also verified along the way that an apparent horizontal-overflow regression was a measurement artifact from a zero-width browser pane, not a real layout fault. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
53 KiB
53 KiB