Test the job queue, and surface it in the dashboard and MCP
The queue shipped with no tests despite having produced four bugs during development, which is the wrong order. 21 tests now cover the parts whose failure modes are not obvious from reading the code: ordering by priority then FIFO, that depth is genuinely unbounded and survives a restart, that running work is never cancelled, that a job abandoned mid-run is requeued rather than left RUNNING forever, and that an LLM job's VRAM requirement comes from its own model rather than a tenant-wide figure -- the mistake that dispatched a 14.9 GB model into 8 GB of free memory and killed llama-server three times. The dashboard had no view of the queue at all, so a stuck queue was indistinguishable from an empty one. The Job Queue panel shows what is running and what it released to get there, pending jobs in execution order with their wait time and a cancel control, and -- when the scheduler is blocked -- how long it has been waiting and why. MCP gains queue_job, get_job_queue and cancel_job, so an agent can line work up rather than firing a request and hoping the GPU is free. Tests: 271. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -10,6 +10,7 @@ from mcp.server import MCPServer
|
||||
import autotune
|
||||
import engines
|
||||
import health
|
||||
import jobs as jobs_mod
|
||||
import overclock_manager
|
||||
import ram_optimizer
|
||||
import telemetry_store
|
||||
@@ -146,6 +147,33 @@ async def get_engine_config() -> str:
|
||||
return json.dumps(await engines.get_engine_config(), indent=2, default=str)
|
||||
|
||||
|
||||
@mcp.tool()
|
||||
def queue_job(tenant: str, payload: Dict[str, Any], priority: Optional[int] = None,
|
||||
label: Optional[str] = None) -> str:
|
||||
"""Queue work for a GPU application without waiting for it.
|
||||
|
||||
tenant: 'ollama' (payload: model, prompt, options) or 'comfyui' (payload: {"prompt":
|
||||
<workflow>}). The queue is on disk, so there is no depth limit; jobs run one at a
|
||||
time, highest priority first, with VRAM arbitrated before each starts."""
|
||||
return json.dumps(jobs_mod.submit(tenant, payload, priority, label), indent=2,
|
||||
default=str)
|
||||
|
||||
|
||||
@mcp.tool()
|
||||
def get_job_queue(state: Optional[str] = None, limit: int = 50) -> str:
|
||||
"""Queued and recent jobs, plus what the scheduler is doing and why it may be
|
||||
waiting. Pending jobs are listed in the order they will run."""
|
||||
return json.dumps({"jobs": jobs_mod.listing(state, limit),
|
||||
"scheduler": jobs_mod.scheduler.get_status()},
|
||||
indent=2, default=str)
|
||||
|
||||
|
||||
@mcp.tool()
|
||||
def cancel_job(job_id: str) -> str:
|
||||
"""Cancel a job that has not started. Running work is never killed."""
|
||||
return json.dumps(jobs_mod.cancel(job_id), indent=2, default=str)
|
||||
|
||||
|
||||
@mcp.tool()
|
||||
async def check_system_health() -> str:
|
||||
"""Check every dependency HyperSwap needs (NVML, sudo nvidia-smi, fan control via the
|
||||
|
||||
Reference in New Issue
Block a user