Describe GPU tenants as data so any application can be arbitrated
The point of this service is fast handoff of one GPU between applications. It grew up
around the two on this box, and their names ended up compiled into process matching,
VRAM attribution, busy detection and release calls alike -- about 385 references
across five modules. That made it a script for Ollama and ComfyUI rather than a GPU
arbitrator.
tenants.py describes an application as data: how to recognise its processes, how to
tell whether it is genuinely working, how to ask it for VRAM back, and how much it
matters when two want the card. Ollama, ComfyUI and the desktop compositor ship as
defaults in tenants.json, so behaviour is unchanged, but the arbitration logic no
longer knows any particular name. Endpoints are generic: GET /api/tenants,
GET /api/tenants/{name}, POST /api/tenants/{name}/release -- the last being the
general form of both the Ollama soft-yield and the ComfyUI purge.
Verified by registering a third application on this machine with no code change: the
speech relay that had been showing up only as anonymous "unmanaged VRAM" is now named,
attributed, and probed by the VRAM it holds rather than by an API it does not have.
Because it declares no release strategy, a release request returns 409 explaining that
its memory cannot be reclaimed, instead of reporting a success that did nothing.
Busy probes deliberately cannot use GPU utilisation. It is shared by every tenant, so
it cannot attribute work to one of them -- the mistake that made a stale ComfyUI queue
entry undetectable earlier in this branch. A tenant's own VRAM is the signal.
Writing the tests exposed that the suite had become non-hermetic: classification is now
configuration, so a test asserting "a third-party process is unmanaged" started failing
the moment the speech relay was registered on this machine. An autouse fixture now
isolates every test from the operator's live tenants.json.
Tests: 231 (was 206).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
38
README.md
38
README.md
@@ -174,10 +174,46 @@ only if the card actually needs it.
|
||||
|
||||
---
|
||||
|
||||
## 1b. Any Application, Not Just These Two
|
||||
|
||||
The purpose is fast handoff of one GPU between applications. It grew up around the two on
|
||||
this box, and their names ended up compiled into process matching, VRAM attribution, busy
|
||||
detection and release calls alike — about 385 references. That made it a script for Ollama
|
||||
and ComfyUI rather than a GPU arbitrator.
|
||||
|
||||
A tenant is now **described as data** in `tenants.json`:
|
||||
|
||||
```json
|
||||
{
|
||||
"name": "trainer",
|
||||
"kind": "other",
|
||||
"priority": 80,
|
||||
"match": { "cmdline": ["train.py"] },
|
||||
"busy": { "type": "vram", "vram_busy_gb": 1.0 },
|
||||
"release": { "type": "http_post", "url": "http://localhost:9999/release" }
|
||||
}
|
||||
```
|
||||
|
||||
| Field | What it answers |
|
||||
| :--- | :--- |
|
||||
| `match` | Which GPU processes belong to this application (name, cmdline substring, or suffix — ComfyUI is a bare `python main.py`) |
|
||||
| `busy` | Whether it is *genuinely* working. `http_count` sums queue lists; `vram` needs no API at all. `vram_floor_gb` catches a queue that claims work while nothing is loaded |
|
||||
| `release` | How to ask for VRAM back — `http_post` with a body, `per_model` for Ollama's per-model unload, or `none` |
|
||||
| `priority` | Who wins contention |
|
||||
|
||||
Ollama, ComfyUI and the desktop compositor ship as defaults, so behaviour is unchanged —
|
||||
but nothing in the arbitration logic knows their names. Endpoints are generic:
|
||||
`GET /api/tenants`, `GET /api/tenants/{name}`, `POST /api/tenants/{name}/release`.
|
||||
|
||||
A tenant with `"release": {"type": "none"}` is still worth declaring. The 842 MB speech
|
||||
relay on this box cannot be reclaimed, and naming it turns anonymous "unmanaged VRAM" into
|
||||
"held by stt-relay, which exposes no release API" — and a release request returns **409**
|
||||
explaining that, rather than silently doing nothing.
|
||||
|
||||
## 1a. Tests
|
||||
|
||||
```bash
|
||||
/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/ -q # 206 passed in ~3.7s
|
||||
/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/ -q # 231 passed in ~3.8s
|
||||
```
|
||||
|
||||
Hermetic: no GPU, no network, no sleeps. An autouse fixture stubs `overclock_manager._sh`
|
||||
|
||||
Reference in New Issue
Block a user