Let any tenant declare its GPU profile and event source; fix priority semantics

Two remaining pieces of the two-application coupling are gone.

Overclock profiles were switched by naming 'comfy' and 'ollama' directly, so a third
application could never get tuned clocks. A tenant declares overclock_profile and the
arbitrator applies whichever the highest-priority *working* tenant asks for, falling
back to the idle profile when nothing is running.

The websocket listener parsed ComfyUI's message schema -- status, execution_start,
executing, execution_success -- which tied the fast path to one application. An event
source is now declarative and the messages are not parsed at all: any message means
"look now", and the tenant's own busy probe decides what is true. That gives the same
sub-second reaction to any application that emits anything on state change, with no
knowledge of what it emits.

Generalising this exposed a design error in the priority rule I had introduced.
plan_release excluded candidates ranking above the demander, which broke both
directions in turn. With the LLM at priority 60 and diffusion at 50, ComfyUI could
never reclaim from Ollama -- the premise the whole service is built on, and preserved
until now only by the ComfyUI-specific trigger that was about to be removed. Swapping
the ranks then broke the reverse: a starved Ollama could no longer reclaim from an
idle ComfyUI.

Priority now orders rather than vetoes. Any idle reclaimable tenant is a candidate,
because an idle tenant is not using its VRAM; priority decides who is asked first, and
busy tenants are never interrupted whatever their rank. Diffusion outranks the LLM,
whose weights reload from page cache in seconds. All three cases are pinned by tests,
including that busy work is never interrupted even by a far higher-priority demander.

Tests: 244 (was 242).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
drjones
2026-09-07 15:23:48 -07:00
parent c6455d7c6e
commit ca97f18be6
5 changed files with 161 additions and 12 deletions

View File

@@ -199,13 +199,23 @@ A tenant is now **described as data** in `tenants.json`:
| `match` | Which GPU processes belong to this application (name, cmdline substring, or suffix — ComfyUI is a bare `python main.py`) |
| `busy` | Whether it is *genuinely* working. `http_count` sums queue lists; `vram` needs no API at all. `vram_floor_gb` catches a queue that claims work while nothing is loaded |
| `release` | How to ask for VRAM back — `http_post` with a body, `per_model` for Ollama's per-model unload, or `none` |
| `priority` | Who wins contention |
| `priority` | Who is asked to yield **first** among idle tenants — it never protects idle memory, and never interrupts work |
| `overclock_profile` | GPU profile applied while this tenant is the active workload |
| `events` | Optional stream (e.g. a websocket) used purely as a wake-up, so reaction is sub-second rather than waiting for the next poll |
Two more fields drive the decision loop: **`needs_vram_gb`** (how much free memory the
application needs before it can work) and **`idle_release_after_s`** (how long it may sit
idle holding VRAM before being asked for it back — deliberately not immediate, so
iterating on a ComfyUI workflow does not reload the checkpoint between every run).
**Priority orders, it does not veto.** An idle tenant is not using its VRAM, so
outranking the demander is no reason to keep it; busy tenants are never interrupted
whatever their rank. Getting this wrong broke both directions in turn — with the LLM
ranked above diffusion, ComfyUI could never preempt Ollama (the service's central
behaviour), and once the ranks were swapped, a starved Ollama could no longer reclaim
from an idle ComfyUI. Diffusion now outranks the LLM, whose weights reload from page
cache in seconds.
`plan_release()` then arbitrates generically: a busy tenant that cannot reach
`needs_vram_gb` *even counting what it already holds* is starved, and the memory is taken
from idle reclaimable tenants below it in priority, lowest first, stopping as soon as
@@ -228,7 +238,7 @@ explaining that, rather than silently doing nothing.
## 1a. Tests
```bash
/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/ -q # 242 passed in ~3.8s
/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/ -q # 244 passed in ~3.8s
```
Hermetic: no GPU, no network, no sleeps. An autouse fixture stubs `overclock_manager._sh`