Describe GPU tenants as data so any application can be arbitrated

The point of this service is fast handoff of one GPU between applications. It grew up
around the two on this box, and their names ended up compiled into process matching,
VRAM attribution, busy detection and release calls alike -- about 385 references
across five modules. That made it a script for Ollama and ComfyUI rather than a GPU
arbitrator.

tenants.py describes an application as data: how to recognise its processes, how to
tell whether it is genuinely working, how to ask it for VRAM back, and how much it
matters when two want the card. Ollama, ComfyUI and the desktop compositor ship as
defaults in tenants.json, so behaviour is unchanged, but the arbitration logic no
longer knows any particular name. Endpoints are generic: GET /api/tenants,
GET /api/tenants/{name}, POST /api/tenants/{name}/release -- the last being the
general form of both the Ollama soft-yield and the ComfyUI purge.

Verified by registering a third application on this machine with no code change: the
speech relay that had been showing up only as anonymous "unmanaged VRAM" is now named,
attributed, and probed by the VRAM it holds rather than by an API it does not have.
Because it declares no release strategy, a release request returns 409 explaining that
its memory cannot be reclaimed, instead of reporting a success that did nothing.

Busy probes deliberately cannot use GPU utilisation. It is shared by every tenant, so
it cannot attribute work to one of them -- the mistake that made a stale ComfyUI queue
entry undetectable earlier in this branch. A tenant's own VRAM is the signal.

Writing the tests exposed that the suite had become non-hermetic: classification is now
configuration, so a test asserting "a third-party process is unmanaged" started failing
the moment the speech relay was registered on this machine. An autouse fixture now
isolates every test from the operator's live tenants.json.

Tests: 231 (was 206).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
drjones
2026-09-07 11:20:44 -07:00
parent c3d9b36035
commit 1bcfbb2335
7 changed files with 764 additions and 20 deletions

View File

@@ -21,6 +21,7 @@ import health
import overclock_manager
import ram_optimizer
import telemetry_store
import tenants as tenants_mod
import thermal_governor
import vram_arbitrator
@@ -322,6 +323,86 @@ async def api_health():
return await health.run_health_checks()
class TenantReleaseRequest(BaseModel):
models: Optional[List[str]] = Field(None, description="For per-model tenants (Ollama), which to unload; defaults to everything resident")
confirm: bool = Field(True, description="Wait for NVML to confirm the VRAM was actually released")
@app.get("/api/tenants", summary="GPU Tenants", tags=["Tenants"])
async def api_tenants():
"""Applications competing for the GPU, as configured.
Each entry declares how its processes are recognised, how to tell whether it is
working, and how to ask it for VRAM back. Adding an application is a config change
in tenants.json, not a code change.
"""
gpu = vram_arbitrator.get_gpu_hardware_stats()
by_tenant = (gpu.get("breakdown", {}) or {}).get("by_tenant_gb", {})
out = []
for t in tenants_mod.describe():
name = t["name"]
# The two original tenants are reported under the bucket names the API has
# always used.
bucket = {"comfyui": "comfy"}.get(name, name)
t["vram_gb"] = by_tenant.get(bucket, 0.0)
out.append(t)
return {"tenants": out, "config_path": tenants_mod.CONFIG_PATH,
"unmanaged_gb": (gpu.get("breakdown", {}) or {}).get("unmanaged_gb", 0.0)}
@app.get("/api/tenants/{name}", summary="One GPU Tenant", tags=["Tenants"])
async def api_tenant(name: str):
"""A single tenant's definition, current VRAM, and whether it is genuinely busy."""
t = tenants_mod.get_tenant(name)
if not t:
raise HTTPException(status_code=404, detail=f"no tenant named '{name}'")
gpu = vram_arbitrator.get_gpu_hardware_stats()
bucket = {"comfyui": "comfy"}.get(name, name)
vram_gb = (gpu.get("breakdown", {}) or {}).get("by_tenant_gb", {}).get(bucket, 0.0)
busy = await tenants_mod.probe_busy(t, vram_gb=vram_gb)
d = t.to_dict()
d.update({"vram_gb": vram_gb, "reclaimable": t.reclaimable, "busy": busy})
return d
@app.post("/api/tenants/{name}/release", summary="Ask a Tenant for its VRAM", tags=["Tenants"])
async def api_tenant_release(name: str, req: Optional[TenantReleaseRequest] = None):
"""Release a tenant's VRAM using whatever mechanism that tenant declares.
This is the generic form of the Ollama soft-yield and the ComfyUI purge: the same
request works for any application in the registry, including ones added later.
"""
t = tenants_mod.get_tenant(name)
if not t:
raise HTTPException(status_code=404, detail=f"no tenant named '{name}'")
if not t.reclaimable:
raise HTTPException(status_code=409,
detail=f"'{name}' declares no way to release VRAM; its "
f"memory cannot be reclaimed by this service")
models = req.models if req else None
if t.release.per_model and not models:
state = await vram_arbitrator.get_ollama_live_state()
models = [m.get("name") for m in state.get("loaded_models", []) if m.get("name")]
before = vram_arbitrator.get_process_vram_bytes()
res = await tenants_mod.release_vram(t, models=models)
if (req is None or req.confirm) and res.get("released"):
bucket = {"comfyui": "comfy"}.get(name, name)
key = {"ollama": "ollama_bytes", "comfy": "comfyui_bytes"}.get(bucket)
if key:
baseline = before[key]
barrier = await vram_arbitrator._await_vram_release(baseline) \
if key == "ollama_bytes" else None
if barrier:
res.update({"outcome": barrier.get("outcome"),
"confirm_ms": barrier.get("confirm_ms")})
after = vram_arbitrator.get_process_vram_bytes()
res["free_vram_gb"] = round(after["free_bytes"] / (1024**3), 2)
return res
@app.get("/api/engines", summary="Live Engine Configuration", tags=["Telemetry"])
async def api_engines():
"""Real configuration of Ollama and ComfyUI, with what each setting implies for