Arbitrate over any number of tenants by priority

Completes the generalisation. Classification and release were already data; the
decision loop was still two hardcoded rules -- yield Ollama when ComfyUI is busy,
purge ComfyUI when Ollama is starved -- which could not express a third participant.

plan_release() works from the registry instead. A busy tenant that cannot reach its
declared needs_vram_gb is starved, and the memory comes from idle reclaimable tenants
below it in priority, lowest first, stopping once enough is freed. Tenants that cannot
be released are named as blockers rather than passed over, so an impossible plan says
which process is in the way. The plan is returned before being acted on, so the
decision is testable and is logged before anything is released. Idle release is now
per-tenant too, replacing the ComfyUI-specific purge timer.

Two bugs found by running it against the live machine rather than only in tests:

Starvation was measured against free VRAM alone, so a tenant working perfectly well on
13 GB was flagged as demanding simply because little was left over -- which is the
normal state of a busy GPU, and would have caused pointless releases from everything
else. A tenant is starved only if it cannot reach what it needs counting what it
already holds.

Fields added to the tenant schema were silently absent from the config already written
to disk, so needs_vram_gb defaulted to 0 and starvation could never trigger for the two
tenants that mattered. Shipped defaults are now merged into an existing config on load,
with explicit user values still winning.

Tests: 242 (was 231), including a three-application contention case -- the property the
hardcoded pair of rules could not express.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
drjones
2026-09-07 14:47:32 -07:00
parent 1bcfbb2335
commit 4a38cd68b3
4 changed files with 319 additions and 2 deletions

View File

@@ -201,8 +201,21 @@ A tenant is now **described as data** in `tenants.json`:
| `release` | How to ask for VRAM back — `http_post` with a body, `per_model` for Ollama's per-model unload, or `none` |
| `priority` | Who wins contention |
Two more fields drive the decision loop: **`needs_vram_gb`** (how much free memory the
application needs before it can work) and **`idle_release_after_s`** (how long it may sit
idle holding VRAM before being asked for it back — deliberately not immediate, so
iterating on a ComfyUI workflow does not reload the checkpoint between every run).
`plan_release()` then arbitrates generically: a busy tenant that cannot reach
`needs_vram_gb` *even counting what it already holds* is starved, and the memory is taken
from idle reclaimable tenants below it in priority, lowest first, stopping as soon as
enough is freed. Tenants that cannot be released are named as blockers rather than
ignored, so `possible: false` comes with the reason. The plan is returned before it is
acted on, which makes the decision testable and loggable.
Ollama, ComfyUI and the desktop compositor ship as defaults, so behaviour is unchanged —
but nothing in the arbitration logic knows their names. Endpoints are generic:
but nothing in the arbitration logic knows their names, and three applications can
contend for the card as easily as two. Endpoints are generic:
`GET /api/tenants`, `GET /api/tenants/{name}`, `POST /api/tenants/{name}/release`.
A tenant with `"release": {"type": "none"}` is still worth declaring. The 842 MB speech
@@ -213,7 +226,7 @@ explaining that, rather than silently doing nothing.
## 1a. Tests
```bash
/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/ -q # 231 passed in ~3.8s
/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/ -q # 242 passed in ~3.8s
```
Hermetic: no GPU, no network, no sleeps. An autouse fixture stubs `overclock_manager._sh`