Test the job queue, and surface it in the dashboard and MCP

The queue shipped with no tests despite having produced four bugs during
development, which is the wrong order. 21 tests now cover the parts whose failure
modes are not obvious from reading the code: ordering by priority then FIFO, that
depth is genuinely unbounded and survives a restart, that running work is never
cancelled, that a job abandoned mid-run is requeued rather than left RUNNING
forever, and that an LLM job's VRAM requirement comes from its own model rather
than a tenant-wide figure -- the mistake that dispatched a 14.9 GB model into 8 GB
of free memory and killed llama-server three times.

The dashboard had no view of the queue at all, so a stuck queue was
indistinguishable from an empty one. The Job Queue panel shows what is running and
what it released to get there, pending jobs in execution order with their wait time
and a cancel control, and -- when the scheduler is blocked -- how long it has been
waiting and why.

MCP gains queue_job, get_job_queue and cancel_job, so an agent can line work up
rather than firing a request and hoping the GPU is free.

Tests: 271.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
drjones
2026-09-07 17:41:55 -07:00
parent 0fbc3963b9
commit 48fbde040c
5 changed files with 310 additions and 2 deletions

View File

@@ -194,7 +194,10 @@ scarce resource this service exists to hand between applications, and overlappin
would just recreate the contention it resolves.
`GET /api/jobs` · `GET /api/jobs/{id}` · `DELETE /api/jobs/{id}` (pending only — running
work is never killed) · `DELETE /api/jobs` to clear the queue.
work is never killed) · `DELETE /api/jobs` to clear the queue. Agents get the same through
MCP (`queue_job`, `get_job_queue`, `cancel_job`), and the dashboard's **Job Queue** panel
shows what is running, what it released to get there, and why the scheduler is waiting if
it is.
**A job that cannot run yet waits; a job that can never run fails with the reason.**
Dispatching into insufficient VRAM does not fail gracefully — it kills `llama-server`
@@ -266,7 +269,7 @@ explaining that, rather than silently doing nothing.
## 1a. Tests
```bash
/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/ -q # 250 passed in ~3.8s
/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/ -q # 271 passed in ~4.0s
```
Hermetic: no GPU, no network, no sleeps. An autouse fixture stubs `overclock_manager._sh`