Add a durable cross-tenant job queue with VRAM-aware scheduling
The service only reacted: it noticed an application had started and scrambled to free memory. Nothing could be lined up. Each application has its own queue but they cannot see each other, so work submitted to one had no way to wait for the other. Jobs are stored in SQLite, so the queue is bounded by disk rather than memory and survives a restart. The scheduler takes the highest-priority pending job, arbitrates VRAM for it through the same plan_release, runs it, and moves on -- one at a time, because overlapping jobs would recreate the contention this service exists to resolve. Four bugs found by running it rather than reasoning about it: Dispatching without checking for room destroyed three queued LLM jobs in a row: a CUDA OOM kills llama-server outright, it does not fail gracefully. A job that cannot run yet now waits. The room check used the tenant's needs_vram_gb, which cannot be right for an LLM -- the requirement is a property of the model being loaded. A flat 4 GB passed with 8 GB free and then a 14.9 GB model was dispatched into it. The requirement is now computed per job. Waiting forever is also wrong. Three jobs sat pending indefinitely needing 14.93 GB on a card where at most ~14.8 GB can ever be free, because an unreclaimable process holds 0.82 GB. A job that cannot be satisfied now fails with the ceiling and the blockers named. plan_release assumed releasing a tenant frees everything it holds. ComfyUI keeps its CUDA context for as long as the process lives, so it reported that releasing ComfyUI would free 0.37 GB against a 0.33 GB shortfall; the job was cleared and the memory never arrived. Tenants declare vram_floor_gb and only memory above it counts. An exception during dispatch left the row RUNNING forever while the scheduler moved on. Failures now land on the job, and jobs left running by a previous process are requeued at startup. Verified end to end: five mixed jobs across both applications, queued at once, all completed with no failures. Tests: 250. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -348,3 +348,58 @@ class TestConfigUpgrade:
|
||||
}]))
|
||||
t = T.get_tenant("ollama")
|
||||
assert t.priority == 5 and t.needs_vram_gb == 99.0
|
||||
|
||||
|
||||
class TestVramFloor:
|
||||
"""VRAM that survives a release must not be promised to anyone else.
|
||||
|
||||
ComfyUI keeps its CUDA context for as long as the process lives, so a purge does not
|
||||
return everything it holds. Ignoring that made plan_release report it would free
|
||||
0.37 GB against a 0.33 GB shortfall; the job was cleared to run and the memory never
|
||||
arrived, so it waited two minutes and then failed.
|
||||
"""
|
||||
|
||||
def test_only_memory_above_the_floor_counts_as_freeable(self):
|
||||
state = [
|
||||
{"name": "llm", "priority": 50, "vram_gb": 0.0, "busy": True,
|
||||
"reclaimable": True, "vram_floor_gb": 0.0},
|
||||
{"name": "diffusion", "priority": 60, "vram_gb": 0.44, "busy": False,
|
||||
"reclaimable": True, "vram_floor_gb": 0.45},
|
||||
]
|
||||
plan = T.plan_release("llm", state, free_gb=14.6, needed_gb=14.93)
|
||||
assert plan["possible"] is False
|
||||
assert plan["release"] == []
|
||||
|
||||
def test_a_loaded_checkpoint_is_still_freeable_above_its_floor(self):
|
||||
state = [
|
||||
{"name": "llm", "priority": 50, "vram_gb": 0.0, "busy": True,
|
||||
"reclaimable": True, "vram_floor_gb": 0.0},
|
||||
{"name": "diffusion", "priority": 60, "vram_gb": 7.0, "busy": False,
|
||||
"reclaimable": True, "vram_floor_gb": 0.45},
|
||||
]
|
||||
plan = T.plan_release("llm", state, free_gb=7.9, needed_gb=14.0)
|
||||
assert plan["release"] == ["diffusion"]
|
||||
# 7.0 held minus a 0.45 floor.
|
||||
assert abs(plan["would_free_gb"] - 6.55) < 0.01
|
||||
|
||||
def test_a_tenant_at_its_floor_is_not_even_listed_for_release(self):
|
||||
state = [
|
||||
{"name": "llm", "priority": 50, "vram_gb": 0.0, "busy": True,
|
||||
"reclaimable": True, "vram_floor_gb": 0.0},
|
||||
{"name": "at_floor", "priority": 10, "vram_gb": 0.3, "busy": False,
|
||||
"reclaimable": True, "vram_floor_gb": 0.45},
|
||||
{"name": "has_room", "priority": 20, "vram_gb": 5.0, "busy": False,
|
||||
"reclaimable": True, "vram_floor_gb": 0.0},
|
||||
]
|
||||
plan = T.plan_release("llm", state, free_gb=0.0, needed_gb=4.0)
|
||||
assert plan["release"] == ["has_room"]
|
||||
|
||||
def test_default_floor_is_zero_so_existing_configs_are_unchanged(self):
|
||||
state = [
|
||||
{"name": "a", "priority": 50, "vram_gb": 0.0, "busy": True,
|
||||
"reclaimable": True},
|
||||
{"name": "b", "priority": 40, "vram_gb": 5.0, "busy": False,
|
||||
"reclaimable": True},
|
||||
]
|
||||
plan = T.plan_release("a", state, free_gb=0.0, needed_gb=5.0)
|
||||
assert plan["release"] == ["b"] and plan["possible"] is True
|
||||
|
||||
Reference in New Issue
Block a user