Add a durable cross-tenant job queue with VRAM-aware scheduling

The service only reacted: it noticed an application had started and scrambled to free
memory. Nothing could be lined up. Each application has its own queue but they cannot
see each other, so work submitted to one had no way to wait for the other.

Jobs are stored in SQLite, so the queue is bounded by disk rather than memory and
survives a restart. The scheduler takes the highest-priority pending job, arbitrates
VRAM for it through the same plan_release, runs it, and moves on -- one at a time,
because overlapping jobs would recreate the contention this service exists to resolve.

Four bugs found by running it rather than reasoning about it:

Dispatching without checking for room destroyed three queued LLM jobs in a row: a CUDA
OOM kills llama-server outright, it does not fail gracefully. A job that cannot run yet
now waits.

The room check used the tenant's needs_vram_gb, which cannot be right for an LLM --
the requirement is a property of the model being loaded. A flat 4 GB passed with 8 GB
free and then a 14.9 GB model was dispatched into it. The requirement is now computed
per job.

Waiting forever is also wrong. Three jobs sat pending indefinitely needing 14.93 GB on
a card where at most ~14.8 GB can ever be free, because an unreclaimable process holds
0.82 GB. A job that cannot be satisfied now fails with the ceiling and the blockers
named.

plan_release assumed releasing a tenant frees everything it holds. ComfyUI keeps its
CUDA context for as long as the process lives, so it reported that releasing ComfyUI
would free 0.37 GB against a 0.33 GB shortfall; the job was cleared and the memory
never arrived. Tenants declare vram_floor_gb and only memory above it counts.

An exception during dispatch left the row RUNNING forever while the scheduler moved on.
Failures now land on the job, and jobs left running by a previous process are requeued
at startup.

Verified end to end: five mixed jobs across both applications, queued at once, all
completed with no failures.

Tests: 250.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
drjones
2026-09-07 17:22:34 -07:00
parent 48bd0096ff
commit 0fbc3963b9
6 changed files with 642 additions and 13 deletions

View File

@@ -18,6 +18,7 @@ from pydantic import BaseModel, Field
import autotune
import engines
import health
import jobs as jobs_mod
import overclock_manager
import ram_optimizer
import telemetry_store
@@ -220,10 +221,13 @@ def _install_shutdown_hook() -> None:
@asynccontextmanager
async def lifespan(app: FastAPI):
telemetry_store.start()
jobs_mod.init()
await broker.start()
await jobs_mod.scheduler.start()
await vram_arbitrator.arbitrator.start()
_install_shutdown_hook()
yield
await jobs_mod.scheduler.stop()
await vram_arbitrator.arbitrator.stop()
await broker.stop()
# Never leave the card with locked clocks and pinned fans after we exit.
@@ -328,6 +332,54 @@ class TenantReleaseRequest(BaseModel):
confirm: bool = Field(True, description="Wait for NVML to confirm the VRAM was actually released")
class JobRequest(BaseModel):
tenant: str = Field(..., description="Which application should run this job", example="comfyui")
payload: Dict[str, Any] = Field(..., description="What to run: a ComfyUI workflow under 'prompt', or Ollama generate parameters")
priority: Optional[int] = Field(None, description="Defaults to the tenant's priority; higher runs sooner")
label: Optional[str] = Field(None, description="Human-readable name for the queue view")
@app.post("/api/jobs", summary="Queue a Job", tags=["Jobs"])
async def api_submit_job(req: JobRequest):
"""Queue work for any tenant. The queue is on disk, so there is no depth limit and
it survives a restart. Jobs run one at a time, highest priority first, with VRAM
arbitrated before each one starts."""
res = jobs_mod.submit(req.tenant, req.payload, req.priority, req.label)
if not res.get("success"):
raise HTTPException(status_code=400, detail=res.get("error"))
return res
@app.get("/api/jobs", summary="The Job Queue", tags=["Jobs"])
async def api_jobs(state: Optional[str] = Query(None, description="pending | running | done | failed | cancelled"),
limit: int = Query(100)):
"""Queued and recent jobs. Pending jobs are listed in the order they will run."""
return {"jobs": jobs_mod.listing(state, limit), "scheduler": jobs_mod.scheduler.get_status()}
@app.get("/api/jobs/{job_id}", summary="One Job", tags=["Jobs"])
async def api_job(job_id: str):
job = jobs_mod.get(job_id)
if not job:
raise HTTPException(status_code=404, detail=f"no job '{job_id}'")
return job
@app.delete("/api/jobs/{job_id}", summary="Cancel a Pending Job", tags=["Jobs"])
async def api_cancel_job(job_id: str):
"""Cancel a job that has not started. Running jobs are left alone -- this service
frees VRAM by asking, never by killing work in flight."""
res = jobs_mod.cancel(job_id)
if not res.get("success"):
raise HTTPException(status_code=409, detail=res.get("error"))
return res
@app.delete("/api/jobs", summary="Cancel All Pending Jobs", tags=["Jobs"])
async def api_clear_jobs():
return jobs_mod.clear_pending()
@app.get("/api/tenants", summary="GPU Tenants", tags=["Tenants"])
async def api_tenants():
"""Applications competing for the GPU, as configured.