Test the job queue, and surface it in the dashboard and MCP

The queue shipped with no tests despite having produced four bugs during
development, which is the wrong order. 21 tests now cover the parts whose failure
modes are not obvious from reading the code: ordering by priority then FIFO, that
depth is genuinely unbounded and survives a restart, that running work is never
cancelled, that a job abandoned mid-run is requeued rather than left RUNNING
forever, and that an LLM job's VRAM requirement comes from its own model rather
than a tenant-wide figure -- the mistake that dispatched a 14.9 GB model into 8 GB
of free memory and killed llama-server three times.

The dashboard had no view of the queue at all, so a stuck queue was
indistinguishable from an empty one. The Job Queue panel shows what is running and
what it released to get there, pending jobs in execution order with their wait time
and a cancel control, and -- when the scheduler is blocked -- how long it has been
waiting and why.

MCP gains queue_job, get_job_queue and cancel_job, so an agent can line work up
rather than firing a request and hoping the GPU is free.

Tests: 271.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
drjones
2026-09-07 17:41:55 -07:00
parent 0fbc3963b9
commit 48fbde040c
5 changed files with 310 additions and 2 deletions

View File

@@ -10,6 +10,7 @@ from mcp.server import MCPServer
import autotune
import engines
import health
import jobs as jobs_mod
import overclock_manager
import ram_optimizer
import telemetry_store
@@ -146,6 +147,33 @@ async def get_engine_config() -> str:
return json.dumps(await engines.get_engine_config(), indent=2, default=str)
@mcp.tool()
def queue_job(tenant: str, payload: Dict[str, Any], priority: Optional[int] = None,
label: Optional[str] = None) -> str:
"""Queue work for a GPU application without waiting for it.
tenant: 'ollama' (payload: model, prompt, options) or 'comfyui' (payload: {"prompt":
<workflow>}). The queue is on disk, so there is no depth limit; jobs run one at a
time, highest priority first, with VRAM arbitrated before each starts."""
return json.dumps(jobs_mod.submit(tenant, payload, priority, label), indent=2,
default=str)
@mcp.tool()
def get_job_queue(state: Optional[str] = None, limit: int = 50) -> str:
"""Queued and recent jobs, plus what the scheduler is doing and why it may be
waiting. Pending jobs are listed in the order they will run."""
return json.dumps({"jobs": jobs_mod.listing(state, limit),
"scheduler": jobs_mod.scheduler.get_status()},
indent=2, default=str)
@mcp.tool()
def cancel_job(job_id: str) -> str:
"""Cancel a job that has not started. Running work is never killed."""
return json.dumps(jobs_mod.cancel(job_id), indent=2, default=str)
@mcp.tool()
async def check_system_health() -> str:
"""Check every dependency HyperSwap needs (NVML, sudo nvidia-smi, fan control via the