Add a durable cross-tenant job queue with VRAM-aware scheduling
The service only reacted: it noticed an application had started and scrambled to free memory. Nothing could be lined up. Each application has its own queue but they cannot see each other, so work submitted to one had no way to wait for the other. Jobs are stored in SQLite, so the queue is bounded by disk rather than memory and survives a restart. The scheduler takes the highest-priority pending job, arbitrates VRAM for it through the same plan_release, runs it, and moves on -- one at a time, because overlapping jobs would recreate the contention this service exists to resolve. Four bugs found by running it rather than reasoning about it: Dispatching without checking for room destroyed three queued LLM jobs in a row: a CUDA OOM kills llama-server outright, it does not fail gracefully. A job that cannot run yet now waits. The room check used the tenant's needs_vram_gb, which cannot be right for an LLM -- the requirement is a property of the model being loaded. A flat 4 GB passed with 8 GB free and then a 14.9 GB model was dispatched into it. The requirement is now computed per job. Waiting forever is also wrong. Three jobs sat pending indefinitely needing 14.93 GB on a card where at most ~14.8 GB can ever be free, because an unreclaimable process holds 0.82 GB. A job that cannot be satisfied now fails with the ceiling and the blockers named. plan_release assumed releasing a tenant frees everything it holds. ComfyUI keeps its CUDA context for as long as the process lives, so it reported that releasing ComfyUI would free 0.37 GB against a 0.33 GB shortfall; the job was cleared and the memory never arrived. Tenants declare vram_floor_gb and only memory above it counts. An exception during dispatch left the row RUNNING forever while the scheduler moved on. Failures now land on the job, and jobs left running by a previous process are requeued at startup. Verified end to end: five mixed jobs across both applications, queued at once, all completed with no failures. Tests: 250. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
30
README.md
30
README.md
@@ -174,6 +174,34 @@ only if the card actually needs it.
|
||||
|
||||
---
|
||||
|
||||
## 1c. Lining Work Up
|
||||
|
||||
Until now this service only *reacted*: it noticed an application had started and
|
||||
scrambled to free memory. Nothing could be queued. Each application has its own queue,
|
||||
but they cannot see each other, so work submitted to one has no way to wait for the other.
|
||||
|
||||
```bash
|
||||
curl -X POST localhost:9090/api/jobs -H 'Content-Type: application/json' -d '{
|
||||
"tenant": "ollama", "label": "nightly-summary",
|
||||
"payload": {"model": "qwen3.8fast:latest", "prompt": "..."}
|
||||
}'
|
||||
```
|
||||
|
||||
Jobs live in SQLite, so the queue is bounded by disk rather than memory and survives a
|
||||
restart. The scheduler takes the highest-priority pending job, arbitrates VRAM for it with
|
||||
the same `plan_release`, runs it, and moves on. One at a time by design — the GPU is the
|
||||
scarce resource this service exists to hand between applications, and overlapping jobs
|
||||
would just recreate the contention it resolves.
|
||||
|
||||
`GET /api/jobs` · `GET /api/jobs/{id}` · `DELETE /api/jobs/{id}` (pending only — running
|
||||
work is never killed) · `DELETE /api/jobs` to clear the queue.
|
||||
|
||||
**A job that cannot run yet waits; a job that can never run fails with the reason.**
|
||||
Dispatching into insufficient VRAM does not fail gracefully — it kills `llama-server`
|
||||
with a CUDA OOM. The requirement is computed per job (an LLM job needs the size of *its*
|
||||
model, not a tenant-wide figure), and if the memory can never be assembled the job fails
|
||||
naming what stands in the way rather than blocking the queue forever.
|
||||
|
||||
## 1b. Any Application, Not Just These Two
|
||||
|
||||
The purpose is fast handoff of one GPU between applications. It grew up around the two on
|
||||
@@ -238,7 +266,7 @@ explaining that, rather than silently doing nothing.
|
||||
## 1a. Tests
|
||||
|
||||
```bash
|
||||
/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/ -q # 244 passed in ~3.8s
|
||||
/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/ -q # 250 passed in ~3.8s
|
||||
```
|
||||
|
||||
Hermetic: no GPU, no network, no sleeps. An autouse fixture stubs `overclock_manager._sh`
|
||||
|
||||
Reference in New Issue
Block a user