Compare commits
13 Commits
b53b026ced
...
main
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
48fbde040c | ||
|
|
0fbc3963b9 | ||
|
|
48bd0096ff | ||
|
|
ca97f18be6 | ||
|
|
c6455d7c6e | ||
|
|
2b00ab4e12 | ||
|
|
4a38cd68b3 | ||
|
|
1bcfbb2335 | ||
|
|
c3d9b36035 | ||
|
|
aeba1b47fd | ||
|
|
81e5d88426 | ||
|
|
aed1c360f0 | ||
|
|
043d61722b |
143
README.md
143
README.md
@@ -21,6 +21,15 @@
|
||||
|
||||
### ⚡ Bidirectional VRAM Hot-Swapping & Arbitration
|
||||
|
||||
* **A stale ComfyUI queue entry no longer disables arbitration.** ComfyUI can leave a
|
||||
dead job in `queue_running` indefinitely; one was found sitting there with the GPU idle
|
||||
and ComfyUI holding 0.56 GB. Trusting that flag made this service believe ComfyUI was
|
||||
permanently busy — so it evicted the LLM on every poll, never ran the idle purge, and
|
||||
never checked for CPU spill. Instrumenting the watchdog showed `busy=6, idle_check=0`.
|
||||
A running entry is now corroborated against ComfyUI's own VRAM (a real job loads
|
||||
gigabytes; a dead one holds only its CUDA context) before it is believed, and a stale
|
||||
entry is reported by `/api/health`. Utilisation is deliberately *not* the signal — it is
|
||||
shared with Ollama and any third-party process.
|
||||
* **Both directions are now automatic.** Yielding Ollama for ComfyUI always was; the
|
||||
reverse was not, despite "bidirectional" in this heading. Which way an LLM fails when
|
||||
it cannot fit depends on configuration: with `n_gpu_layers` left to Ollama it spills
|
||||
@@ -111,6 +120,18 @@ only if the card actually needs it.
|
||||
* **Honest gain reporting**: gain against the profile's *current* setting is reported separately from the spread across values tried. Conflating them turns a flat result into a headline "+102%".
|
||||
* **Safety**: refuses to start while ComfyUI is executing, suspends the arbitrator's automatic profile switching for the duration (otherwise a diffusion benchmark trips `trigger_comfy_priority`, which reapplies the whole profile and overwrites the clock being measured), and restores the original profile in a `finally` block — including on exception or cancellation.
|
||||
|
||||
### 🔧 Live Engine Configuration (`engines.py`)
|
||||
* `GET /api/engines` reports the **real, current** configuration of both engines and what
|
||||
each setting implies for arbitration — because the settings that dictate this service's
|
||||
behaviour live outside its own codebase.
|
||||
* `OLLAMA_NUM_PARALLEL=1` is why a `keep_alive: 0` unload queues behind a running
|
||||
generation and is reported as *deferred* rather than failed. `OLLAMA_MAX_LOADED_MODELS=1`
|
||||
is why every swap evicts the previous model. Working these out originally meant reading
|
||||
journald and the systemd unit by hand.
|
||||
* The dashboard's engine subtitles now come from this endpoint. They were previously
|
||||
hardcoded — and happened to be accurate, which is worse than being wrong, since they
|
||||
would have kept looking accurate after the configuration changed.
|
||||
|
||||
### 🩺 Dependency Self-Check (`health.py`)
|
||||
* `GET /api/health` verifies **everything this service depends on**: NVML, passwordless
|
||||
sudo for `nvidia-smi`, fan control through the headless X server, overclock drift, the
|
||||
@@ -153,10 +174,102 @@ only if the card actually needs it.
|
||||
|
||||
---
|
||||
|
||||
## 1c. Lining Work Up
|
||||
|
||||
Until now this service only *reacted*: it noticed an application had started and
|
||||
scrambled to free memory. Nothing could be queued. Each application has its own queue,
|
||||
but they cannot see each other, so work submitted to one has no way to wait for the other.
|
||||
|
||||
```bash
|
||||
curl -X POST localhost:9090/api/jobs -H 'Content-Type: application/json' -d '{
|
||||
"tenant": "ollama", "label": "nightly-summary",
|
||||
"payload": {"model": "qwen3.8fast:latest", "prompt": "..."}
|
||||
}'
|
||||
```
|
||||
|
||||
Jobs live in SQLite, so the queue is bounded by disk rather than memory and survives a
|
||||
restart. The scheduler takes the highest-priority pending job, arbitrates VRAM for it with
|
||||
the same `plan_release`, runs it, and moves on. One at a time by design — the GPU is the
|
||||
scarce resource this service exists to hand between applications, and overlapping jobs
|
||||
would just recreate the contention it resolves.
|
||||
|
||||
`GET /api/jobs` · `GET /api/jobs/{id}` · `DELETE /api/jobs/{id}` (pending only — running
|
||||
work is never killed) · `DELETE /api/jobs` to clear the queue. Agents get the same through
|
||||
MCP (`queue_job`, `get_job_queue`, `cancel_job`), and the dashboard's **Job Queue** panel
|
||||
shows what is running, what it released to get there, and why the scheduler is waiting if
|
||||
it is.
|
||||
|
||||
**A job that cannot run yet waits; a job that can never run fails with the reason.**
|
||||
Dispatching into insufficient VRAM does not fail gracefully — it kills `llama-server`
|
||||
with a CUDA OOM. The requirement is computed per job (an LLM job needs the size of *its*
|
||||
model, not a tenant-wide figure), and if the memory can never be assembled the job fails
|
||||
naming what stands in the way rather than blocking the queue forever.
|
||||
|
||||
## 1b. Any Application, Not Just These Two
|
||||
|
||||
The purpose is fast handoff of one GPU between applications. It grew up around the two on
|
||||
this box, and their names ended up compiled into process matching, VRAM attribution, busy
|
||||
detection and release calls alike — about 385 references. That made it a script for Ollama
|
||||
and ComfyUI rather than a GPU arbitrator.
|
||||
|
||||
A tenant is now **described as data** in `tenants.json`:
|
||||
|
||||
```json
|
||||
{
|
||||
"name": "trainer",
|
||||
"kind": "other",
|
||||
"priority": 80,
|
||||
"match": { "cmdline": ["train.py"] },
|
||||
"busy": { "type": "vram", "vram_busy_gb": 1.0 },
|
||||
"release": { "type": "http_post", "url": "http://localhost:9999/release" }
|
||||
}
|
||||
```
|
||||
|
||||
| Field | What it answers |
|
||||
| :--- | :--- |
|
||||
| `match` | Which GPU processes belong to this application (name, cmdline substring, or suffix — ComfyUI is a bare `python main.py`) |
|
||||
| `busy` | Whether it is *genuinely* working. `http_count` sums queue lists; `vram` needs no API at all. `vram_floor_gb` catches a queue that claims work while nothing is loaded |
|
||||
| `release` | How to ask for VRAM back — `http_post` with a body, `per_model` for Ollama's per-model unload, or `none` |
|
||||
| `priority` | Who is asked to yield **first** among idle tenants — it never protects idle memory, and never interrupts work |
|
||||
| `overclock_profile` | GPU profile applied while this tenant is the active workload |
|
||||
| `events` | Optional stream (e.g. a websocket) used purely as a wake-up, so reaction is sub-second rather than waiting for the next poll |
|
||||
|
||||
Two more fields drive the decision loop: **`needs_vram_gb`** (how much free memory the
|
||||
application needs before it can work) and **`idle_release_after_s`** (how long it may sit
|
||||
idle holding VRAM before being asked for it back — deliberately not immediate, so
|
||||
iterating on a ComfyUI workflow does not reload the checkpoint between every run).
|
||||
|
||||
**Priority orders, it does not veto.** An idle tenant is not using its VRAM, so
|
||||
outranking the demander is no reason to keep it; busy tenants are never interrupted
|
||||
whatever their rank. Getting this wrong broke both directions in turn — with the LLM
|
||||
ranked above diffusion, ComfyUI could never preempt Ollama (the service's central
|
||||
behaviour), and once the ranks were swapped, a starved Ollama could no longer reclaim
|
||||
from an idle ComfyUI. Diffusion now outranks the LLM, whose weights reload from page
|
||||
cache in seconds.
|
||||
|
||||
`plan_release()` then arbitrates generically: a busy tenant that cannot reach
|
||||
`needs_vram_gb` *even counting what it already holds* is starved, and the memory is taken
|
||||
from idle reclaimable tenants below it in priority, lowest first, stopping as soon as
|
||||
enough is freed. Tenants that cannot be released are named as blockers rather than
|
||||
ignored, so `possible: false` comes with the reason. The plan is returned before it is
|
||||
acted on, which makes the decision testable and loggable.
|
||||
|
||||
Ollama, ComfyUI and the desktop compositor ship as defaults, so behaviour is unchanged —
|
||||
but nothing in the arbitration logic knows their names, and three applications can
|
||||
contend for the card as easily as two. The dashboard's **GPU Tenants** panel lists all of
|
||||
them ordered by the priority arbitration actually considers, with the last decision and
|
||||
why it could or could not be satisfied. Endpoints are generic:
|
||||
`GET /api/tenants`, `GET /api/tenants/{name}`, `POST /api/tenants/{name}/release`.
|
||||
|
||||
A tenant with `"release": {"type": "none"}` is still worth declaring. The 842 MB speech
|
||||
relay on this box cannot be reclaimed, and naming it turns anonymous "unmanaged VRAM" into
|
||||
"held by stt-relay, which exposes no release API" — and a release request returns **409**
|
||||
explaining that, rather than silently doing nothing.
|
||||
|
||||
## 1a. Tests
|
||||
|
||||
```bash
|
||||
/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/ -q # 182 passed in ~3.7s
|
||||
/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/ -q # 271 passed in ~4.0s
|
||||
```
|
||||
|
||||
Hermetic: no GPU, no network, no sleeps. An autouse fixture stubs `overclock_manager._sh`
|
||||
@@ -173,6 +286,33 @@ which contradicts the hardware fails loudly rather than silently:
|
||||
| Busy yield | VRAM held at ≥50% GPU utilisation | A mid-generation model is finishing, not failing |
|
||||
| Residency confidence | probe trusted only at 100% | A 12-window probe once cleared 90% on a mostly-cold file |
|
||||
|
||||
## 1b. End-to-End Verification
|
||||
|
||||
```bash
|
||||
python verify_arbitration.py # full cycle, a few minutes
|
||||
python verify_arbitration.py --quick # skip the diffusion stages
|
||||
```
|
||||
|
||||
The unit suite covers logic in isolation. This exercises the promise the service exists
|
||||
to make — an LLM and a diffusion pipeline sharing one 16 GB card — against real hardware,
|
||||
and reports what actually happened at each stage. It restores what it changes and refuses
|
||||
to start if ComfyUI is busy.
|
||||
|
||||
A representative run on this machine:
|
||||
|
||||
| Stage | Result |
|
||||
| :--- | :--- |
|
||||
| LLM load, classified by achieved bandwidth | 1.96 GB in 1327 ms → 1.47 GB/s → Partial Cache |
|
||||
| VRAM yield confirmed against NVML | released in 43 ms, 2.39 GB freed |
|
||||
| Diffusion, cold (includes checkpoint load) | 17863 ms → 1.12 it/s |
|
||||
| Diffusion, warm | 3645 ms → **5.49 it/s** |
|
||||
| ComfyUI retains its checkpoint | 7.03 GB held through the idle window |
|
||||
| VRAM attribution adds up | 15.58 GB attributed vs 15.80 GB NVML — Ollama 7.71 + ComfyUI 7.03 coexisting |
|
||||
| Reported GPU state matches hardware | profile asks 320 W, card reports 320 W |
|
||||
|
||||
The stages report warnings rather than passes when they did not actually prove anything —
|
||||
a reclaim that was never needed is not evidence that reclaiming works.
|
||||
|
||||
## 2. Architectural Overview
|
||||
|
||||
```mermaid
|
||||
@@ -264,6 +404,7 @@ The HyperSwap server runs on port `9090` by default. Interactive OpenAPI/Swagger
|
||||
| `/api/analytics/models` | `GET` | Recency/frequency model ranking used to prioritise the warm budget. |
|
||||
| `/api/history?durable=true` | `GET` | Swap history from the persistent store rather than the in-memory ring. |
|
||||
| `/api/db` | `GET` | Store location, row counts and how many hours of history are held. |
|
||||
| `/api/engines` | `GET` | Live Ollama and ComfyUI configuration, with what each setting implies for arbitration. |
|
||||
| `/api/health` | `GET` | Dependency self-check: NVML, sudo, fan control, drift, store, upstreams — each with impact and remediation. |
|
||||
|
||||
### Governor & Autotune Endpoints
|
||||
|
||||
117
engines.py
Normal file
117
engines.py
Normal file
@@ -0,0 +1,117 @@
|
||||
"""Live configuration of the two engines HyperSwap arbitrates between.
|
||||
|
||||
Arbitration behaviour is largely dictated by settings that live outside this codebase.
|
||||
Working out why a yield behaved the way it did meant reading journald and the ollama
|
||||
unit by hand: OLLAMA_NUM_PARALLEL decides whether an unload queues behind a running
|
||||
generation, OLLAMA_MAX_LOADED_MODELS decides whether more than one model can be
|
||||
resident, and a pinned n_gpu_layers decides whether a model that will not fit spills to
|
||||
the CPU or fails outright. Those are worth reading and explaining rather than hardcoding
|
||||
into a dashboard subtitle that silently goes stale.
|
||||
"""
|
||||
import json
|
||||
import logging
|
||||
import subprocess
|
||||
from typing import Any, Dict, List, Optional
|
||||
|
||||
import httpx
|
||||
|
||||
import vram_arbitrator
|
||||
|
||||
logger = logging.getLogger("engines")
|
||||
|
||||
# Settings that change how the arbitrator must behave, with what they imply.
|
||||
OLLAMA_SETTING_NOTES = {
|
||||
"OLLAMA_NUM_PARALLEL": (
|
||||
"Requests per model. At 1, a keep_alive:0 unload queues behind any running "
|
||||
"generation and applies when it finishes — which is why a busy model is "
|
||||
"reported as deferred rather than failed."),
|
||||
"OLLAMA_MAX_LOADED_MODELS": (
|
||||
"How many models may be resident at once. At 1, Ollama evicts the previous "
|
||||
"model on every swap."),
|
||||
"OLLAMA_KEEP_ALIVE": (
|
||||
"Default residency after a request. Long values keep VRAM occupied and make "
|
||||
"ComfyUI wait for an explicit yield."),
|
||||
"OLLAMA_FLASH_ATTENTION": "FlashAttention kernels for attention.",
|
||||
"OLLAMA_KV_CACHE_TYPE": "KV cache quantisation; smaller types cut VRAM per context.",
|
||||
"OLLAMA_NUM_BATCH": "Prompt-evaluation batch size.",
|
||||
}
|
||||
|
||||
|
||||
def _ollama_unit_environment() -> Dict[str, str]:
|
||||
"""Read the ollama service's environment. Empty if it is not a systemd unit."""
|
||||
env: Dict[str, str] = {}
|
||||
try:
|
||||
proc = subprocess.run(["systemctl", "show", "ollama", "-p", "Environment",
|
||||
"--value"], capture_output=True, text=True, timeout=8)
|
||||
for token in proc.stdout.split():
|
||||
if "=" in token and token.startswith("OLLAMA"):
|
||||
k, _, v = token.partition("=")
|
||||
env[k] = v
|
||||
except Exception as e:
|
||||
logger.debug(f"could not read ollama unit environment: {e}")
|
||||
return env
|
||||
|
||||
|
||||
async def get_engine_config() -> Dict[str, Any]:
|
||||
"""Real, live configuration of both engines, with arbitration implications."""
|
||||
ollama_env = _ollama_unit_environment()
|
||||
ollama_settings = [
|
||||
{"key": k, "value": v, "means": OLLAMA_SETTING_NOTES.get(k, "")}
|
||||
for k, v in sorted(ollama_env.items())
|
||||
]
|
||||
|
||||
# A short, honest summary line to replace the dashboard's hardcoded subtitle.
|
||||
feature_bits: List[str] = []
|
||||
if ollama_env.get("OLLAMA_FLASH_ATTENTION") == "1":
|
||||
feature_bits.append("FlashAttention")
|
||||
kv = ollama_env.get("OLLAMA_KV_CACHE_TYPE")
|
||||
if kv:
|
||||
feature_bits.append(f"{kv} KV cache")
|
||||
host = ollama_env.get("OLLAMA_HOST", "")
|
||||
port = host.rsplit(":", 1)[-1] if ":" in host else "11434"
|
||||
|
||||
ollama = {
|
||||
"port": port,
|
||||
"settings": ollama_settings,
|
||||
"summary": " + ".join(feature_bits) if feature_bits else "default configuration",
|
||||
"max_loaded_models": ollama_env.get("OLLAMA_MAX_LOADED_MODELS"),
|
||||
"num_parallel": ollama_env.get("OLLAMA_NUM_PARALLEL"),
|
||||
"keep_alive": ollama_env.get("OLLAMA_KEEP_ALIVE"),
|
||||
"config_source": "systemd unit environment" if ollama_env else "unavailable",
|
||||
}
|
||||
|
||||
comfy: Dict[str, Any] = {"online": False}
|
||||
try:
|
||||
async with httpx.AsyncClient(timeout=4.0) as c:
|
||||
r = await c.get(f"{vram_arbitrator.COMFY_API_BASE}/system_stats")
|
||||
if r.status_code == 200:
|
||||
data = r.json()
|
||||
system = data.get("system", {})
|
||||
argv = system.get("argv") or []
|
||||
devices = data.get("devices") or []
|
||||
dev = devices[0] if devices else {}
|
||||
# The allocator is named in the device string; it is the closest thing
|
||||
# ComfyUI reports to the "async offload" the old subtitle asserted.
|
||||
dev_name = dev.get("name", "")
|
||||
allocator = ("cudaMallocAsync" if "cudaMallocAsync" in dev_name
|
||||
else "cudaMalloc" if "cudaMalloc" in dev_name else "unknown")
|
||||
vram_flags = [a for a in argv
|
||||
if a in ("--lowvram", "--novram", "--highvram", "--normalvram",
|
||||
"--gpu-only", "--cpu")]
|
||||
comfy = {
|
||||
"online": True,
|
||||
"version": system.get("comfyui_version"),
|
||||
"pytorch": system.get("pytorch_version"),
|
||||
# split()[0] on an absent version raises IndexError, which would have
|
||||
# been swallowed by the except below and reported ComfyUI as offline.
|
||||
"python": ((system.get("python_version") or "").split() or [None])[0],
|
||||
"argv": argv,
|
||||
"vram_mode": vram_flags[0] if vram_flags else "default (auto)",
|
||||
"allocator": allocator,
|
||||
"device": dev_name,
|
||||
"summary": f"{allocator}, {vram_flags[0] if vram_flags else 'auto VRAM'}",
|
||||
}
|
||||
except Exception as e:
|
||||
comfy = {"online": False, "error": str(e)[:120]}
|
||||
|
||||
return {"ollama": ollama, "comfyui": comfy}
|
||||
91
health.py
91
health.py
@@ -10,6 +10,7 @@ when it is missing and how to fix it. A degraded dependency should be loud.
|
||||
"""
|
||||
import asyncio
|
||||
import logging
|
||||
import time
|
||||
import os
|
||||
import time
|
||||
from typing import Any, Dict, List
|
||||
@@ -116,6 +117,25 @@ def _check_comfy_ws() -> Dict[str, Any]:
|
||||
return _check("comfyui websocket", OK, "subscribed")
|
||||
|
||||
|
||||
def _check_comfy_queue() -> Dict[str, Any]:
|
||||
"""A stale ComfyUI queue entry disables half of this service's logic."""
|
||||
arb = vram_arbitrator.arbitrator
|
||||
if arb.comfy_stale_job:
|
||||
return _check("comfyui queue", DEGRADED,
|
||||
f"prompt {arb.comfy_stale_job} claims to be running but the GPU is idle",
|
||||
"ComfyUI looks permanently busy, so the LLM is evicted repeatedly, "
|
||||
"the idle purge never runs and CPU-spill is never checked",
|
||||
"Clear it from the ComfyUI queue, or POST /queue with "
|
||||
"{\"clear\": true} to ComfyUI")
|
||||
branches = arb.watchdog_branches
|
||||
if branches.get("idle_check", 0) == 0 and branches.get("busy", 0) > 20:
|
||||
return _check("comfyui queue", DEGRADED,
|
||||
"the watchdog has only ever seen ComfyUI as busy",
|
||||
"The idle purge and starvation check are not running",
|
||||
"Check the ComfyUI queue for a stuck entry")
|
||||
return _check("comfyui queue", OK, "queue state corroborated against GPU activity")
|
||||
|
||||
|
||||
def _check_store() -> Dict[str, Any]:
|
||||
info = telemetry_store.db_info()
|
||||
if not info.get("exists"):
|
||||
@@ -139,6 +159,75 @@ def _check_residency() -> Dict[str, Any]:
|
||||
cap.get("hint", ""))
|
||||
|
||||
|
||||
def _comfy_vram_floor_gb(default: float = 0.0, days: float = 1.0) -> float:
|
||||
"""Lowest VRAM ComfyUI has been observed holding while alive.
|
||||
|
||||
A purge frees checkpoints but not the CUDA context, so ComfyUI keeps a few hundred
|
||||
MB for as long as the process runs. The minimum seen in recent telemetry is a better
|
||||
estimate of that floor than whatever it happens to hold right now, which could be a
|
||||
7 GB checkpoint mid-generation.
|
||||
"""
|
||||
try:
|
||||
rows = telemetry_store._rows(
|
||||
"SELECT MIN(comfy_bytes) AS floor FROM telemetry "
|
||||
"WHERE ts > ? AND comfy_bytes > 0",
|
||||
(time.time() - days * 86400,))
|
||||
if rows and rows[0].get("floor"):
|
||||
return round(rows[0]["floor"] / (1024 ** 3), 2)
|
||||
except Exception as e:
|
||||
logger.debug(f"comfy floor lookup failed: {e}")
|
||||
return default
|
||||
|
||||
|
||||
def _check_unmanaged_vram() -> Dict[str, Any]:
|
||||
"""Report unreclaimable VRAM in terms of what it actually costs.
|
||||
|
||||
"0.82 GB unmanaged" is a number. "0.82 GB unmanaged, which is why three of your
|
||||
models can no longer fit" is something you can act on.
|
||||
"""
|
||||
stats = vram_arbitrator.get_gpu_hardware_stats()
|
||||
if not stats.get("available"):
|
||||
return _check("unmanaged VRAM", DEGRADED, "GPU unavailable")
|
||||
bd = stats.get("breakdown", {})
|
||||
unmanaged_gb = bd.get("unmanaged_gb", 0.0)
|
||||
procs = bd.get("unmanaged", [])
|
||||
if not procs:
|
||||
return _check("unmanaged VRAM", OK, "no third-party GPU processes")
|
||||
|
||||
total_gb = stats.get("vram_total_gb", 0)
|
||||
# What HyperSwap could offer at best. Three things are never available: the desktop,
|
||||
# processes it cannot touch, and ComfyUI's own CUDA context, which survives a purge.
|
||||
# Omitting that last one made this check claim a 14.93 GB model would fit against a
|
||||
# real ceiling of 14.60 GB -- the model that had just returned 507.
|
||||
comfy_floor_gb = _comfy_vram_floor_gb(default=bd.get("comfyui_gb", 0.0))
|
||||
ceiling_gb = total_gb - unmanaged_gb - bd.get("desktop_gb", 0.0) - comfy_floor_gb
|
||||
try:
|
||||
blobs = ram_optimizer.find_ollama_model_files()
|
||||
except Exception:
|
||||
blobs = []
|
||||
# Measured on this box: a 12.87 GB blob occupies 14.9 GB once context and KV cache
|
||||
# are allocated.
|
||||
VRAM_OVERHEAD = 1.16
|
||||
blocked = sorted(
|
||||
{b["model"]: b for b in blobs
|
||||
if b["size_gb"] * VRAM_OVERHEAD > ceiling_gb
|
||||
and b["size_gb"] * VRAM_OVERHEAD <= ceiling_gb + unmanaged_gb}.values(),
|
||||
key=lambda b: -b["size_gb"])
|
||||
|
||||
names = ", ".join(b["model"] for b in blocked[:3])
|
||||
who = ", ".join(f"{p['name']} ({p['vram_mb']} MB)" for p in procs[:2])
|
||||
if blocked:
|
||||
return _check("unmanaged VRAM", DEGRADED,
|
||||
f"{unmanaged_gb} GB held by {who}",
|
||||
f"{len(blocked)} model(s) fit within {ceiling_gb + unmanaged_gb:.2f} GB "
|
||||
f"but not the {ceiling_gb:.2f} GB actually available: {names}",
|
||||
"Stop that process to reclaim the difference, or accept that "
|
||||
"these models cannot load")
|
||||
return _check("unmanaged VRAM", OK,
|
||||
f"{unmanaged_gb} GB held by {who}; no model is blocked by it",
|
||||
"", "")
|
||||
|
||||
|
||||
def _check_model_dirs() -> Dict[str, Any]:
|
||||
comfy_dir = ram_optimizer.COMFY_MODELS_DIR
|
||||
if not os.path.isdir(comfy_dir):
|
||||
@@ -158,7 +247,7 @@ async def run_health_checks() -> Dict[str, Any]:
|
||||
|
||||
sync_checks = [_check_nvml, _check_sudo_smi, _check_fan_control,
|
||||
_check_profile_drift, _check_store, _check_residency,
|
||||
_check_model_dirs, _check_comfy_ws]
|
||||
_check_model_dirs, _check_comfy_ws, _check_unmanaged_vram]
|
||||
results: List[Dict[str, Any]] = []
|
||||
for fn in sync_checks:
|
||||
try:
|
||||
|
||||
445
jobs.py
Normal file
445
jobs.py
Normal file
@@ -0,0 +1,445 @@
|
||||
"""A durable, unbounded job queue across every GPU tenant.
|
||||
|
||||
Until now this service only reacted: it noticed an application had started working and
|
||||
scrambled to free memory. Nothing could be *lined up*. Each application has its own queue
|
||||
(ComfyUI's prompt queue, Ollama's serialised requests), but they cannot see each other, so
|
||||
work submitted to one has no way to wait politely for the other.
|
||||
|
||||
Jobs submitted here are stored in SQLite, so the queue is limited by disk rather than
|
||||
memory and survives a restart. The scheduler takes the highest-priority pending job,
|
||||
makes sure its tenant actually has the VRAM to run it -- reusing the same plan_release
|
||||
arbitration -- dispatches it, and moves on.
|
||||
"""
|
||||
import asyncio
|
||||
import json
|
||||
import logging
|
||||
import sqlite3
|
||||
import time
|
||||
import uuid
|
||||
from typing import Any, Dict, List, Optional
|
||||
|
||||
import httpx
|
||||
|
||||
import telemetry_store
|
||||
import tenants as tenants_mod
|
||||
|
||||
logger = logging.getLogger("jobs")
|
||||
|
||||
PENDING, RUNNING, DONE, FAILED, CANCELLED = (
|
||||
"pending", "running", "done", "failed", "cancelled")
|
||||
|
||||
SCHEMA = """
|
||||
CREATE TABLE IF NOT EXISTS jobs (
|
||||
id TEXT PRIMARY KEY,
|
||||
tenant TEXT NOT NULL,
|
||||
priority INTEGER NOT NULL DEFAULT 50,
|
||||
payload TEXT NOT NULL,
|
||||
state TEXT NOT NULL,
|
||||
submitted_at REAL NOT NULL,
|
||||
started_at REAL,
|
||||
finished_at REAL,
|
||||
error TEXT,
|
||||
result TEXT,
|
||||
label TEXT
|
||||
);
|
||||
CREATE INDEX IF NOT EXISTS idx_jobs_state ON jobs(state, priority DESC, submitted_at);
|
||||
"""
|
||||
|
||||
|
||||
def _conn() -> sqlite3.Connection:
|
||||
c = sqlite3.connect(telemetry_store.DB_PATH, timeout=10.0)
|
||||
c.row_factory = sqlite3.Row
|
||||
return c
|
||||
|
||||
|
||||
def init() -> None:
|
||||
with _conn() as c:
|
||||
c.executescript(SCHEMA)
|
||||
|
||||
|
||||
def submit(tenant: str, payload: Dict[str, Any], priority: Optional[int] = None,
|
||||
label: Optional[str] = None) -> Dict[str, Any]:
|
||||
"""Queue a job. There is no depth limit: the queue lives on disk."""
|
||||
t = tenants_mod.get_tenant(tenant)
|
||||
if not t:
|
||||
return {"success": False, "error": f"no tenant named '{tenant}'"}
|
||||
job_id = uuid.uuid4().hex[:12]
|
||||
row = {
|
||||
"id": job_id, "tenant": tenant,
|
||||
"priority": t.priority if priority is None else int(priority),
|
||||
"payload": json.dumps(payload), "state": PENDING,
|
||||
"submitted_at": time.time(), "label": label,
|
||||
}
|
||||
with _conn() as c:
|
||||
c.execute("INSERT INTO jobs (id, tenant, priority, payload, state, submitted_at,"
|
||||
" label) VALUES (:id,:tenant,:priority,:payload,:state,:submitted_at,"
|
||||
":label)", row)
|
||||
logger.info(f"queued job {job_id} for '{tenant}' at priority {row['priority']}")
|
||||
return {"success": True, "id": job_id, "tenant": tenant,
|
||||
"priority": row["priority"], "state": PENDING}
|
||||
|
||||
|
||||
def cancel(job_id: str) -> Dict[str, Any]:
|
||||
with _conn() as c:
|
||||
cur = c.execute("UPDATE jobs SET state=?, finished_at=? WHERE id=? AND state=?",
|
||||
(CANCELLED, time.time(), job_id, PENDING))
|
||||
if cur.rowcount:
|
||||
return {"success": True, "id": job_id, "state": CANCELLED}
|
||||
return {"success": False, "error": "job is not pending (already running or finished)"}
|
||||
|
||||
|
||||
def clear_pending() -> Dict[str, Any]:
|
||||
with _conn() as c:
|
||||
cur = c.execute("UPDATE jobs SET state=?, finished_at=? WHERE state=?",
|
||||
(CANCELLED, time.time(), PENDING))
|
||||
return {"success": True, "cancelled": cur.rowcount}
|
||||
|
||||
|
||||
def get(job_id: str) -> Optional[Dict[str, Any]]:
|
||||
with _conn() as c:
|
||||
r = c.execute("SELECT * FROM jobs WHERE id=?", (job_id,)).fetchone()
|
||||
return _row(r) if r else None
|
||||
|
||||
|
||||
def _row(r: sqlite3.Row) -> Dict[str, Any]:
|
||||
d = dict(r)
|
||||
for key in ("payload", "result"):
|
||||
if d.get(key):
|
||||
try:
|
||||
d[key] = json.loads(d[key])
|
||||
except Exception:
|
||||
pass
|
||||
if d.get("started_at") and d.get("finished_at"):
|
||||
d["duration_s"] = round(d["finished_at"] - d["started_at"], 2)
|
||||
if d.get("state") == PENDING:
|
||||
d["waiting_s"] = round(time.time() - d["submitted_at"], 1)
|
||||
return d
|
||||
|
||||
|
||||
def listing(state: Optional[str] = None, limit: int = 100) -> List[Dict[str, Any]]:
|
||||
q = "SELECT * FROM jobs"
|
||||
args: List[Any] = []
|
||||
if state:
|
||||
q += " WHERE state=?"
|
||||
args.append(state)
|
||||
# Pending jobs in the order the scheduler will take them; everything else newest first.
|
||||
q += (" ORDER BY priority DESC, submitted_at ASC" if state == PENDING
|
||||
else " ORDER BY submitted_at DESC")
|
||||
q += " LIMIT ?"
|
||||
args.append(limit)
|
||||
with _conn() as c:
|
||||
return [_row(r) for r in c.execute(q, args).fetchall()]
|
||||
|
||||
|
||||
def stats() -> Dict[str, Any]:
|
||||
with _conn() as c:
|
||||
rows = c.execute("SELECT state, COUNT(*) n FROM jobs GROUP BY state").fetchall()
|
||||
by_state = {r["state"]: r["n"] for r in rows}
|
||||
pend = c.execute(
|
||||
"SELECT tenant, COUNT(*) n FROM jobs WHERE state=? GROUP BY tenant",
|
||||
(PENDING,)).fetchall()
|
||||
oldest = c.execute(
|
||||
"SELECT MIN(submitted_at) t FROM jobs WHERE state=?", (PENDING,)).fetchone()
|
||||
return {
|
||||
"by_state": by_state,
|
||||
"pending_by_tenant": {r["tenant"]: r["n"] for r in pend},
|
||||
"queue_depth": by_state.get(PENDING, 0),
|
||||
"oldest_pending_s": (round(time.time() - oldest["t"], 1)
|
||||
if oldest and oldest["t"] else None),
|
||||
}
|
||||
|
||||
|
||||
def requeue_orphans() -> int:
|
||||
"""Return jobs abandoned mid-run to the queue.
|
||||
|
||||
RUNNING means "this process is working on it". If no process is, that is untrue, and
|
||||
the job would otherwise never finish and never retry.
|
||||
"""
|
||||
with _conn() as c:
|
||||
cur = c.execute("UPDATE jobs SET state=?, started_at=NULL WHERE state=?",
|
||||
(PENDING, RUNNING))
|
||||
return cur.rowcount
|
||||
|
||||
|
||||
def _next_job() -> Optional[Dict[str, Any]]:
|
||||
with _conn() as c:
|
||||
r = c.execute(
|
||||
"SELECT * FROM jobs WHERE state=? ORDER BY priority DESC, submitted_at ASC"
|
||||
" LIMIT 1", (PENDING,)).fetchone()
|
||||
return _row(r) if r else None
|
||||
|
||||
|
||||
def _mark(job_id: str, state: str, **fields) -> None:
|
||||
sets = ", ".join(f"{k}=?" for k in fields)
|
||||
args = list(fields.values()) + [state, job_id]
|
||||
with _conn() as c:
|
||||
c.execute(f"UPDATE jobs SET {sets + ', ' if sets else ''}state=? WHERE id=?", args)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------- dispatch
|
||||
|
||||
async def _dispatch_comfy(payload: Dict[str, Any]) -> Dict[str, Any]:
|
||||
"""Hand a workflow to ComfyUI and wait for it to finish."""
|
||||
base = tenants_mod.get_tenant("comfyui").busy.url.rsplit("/", 1)[0]
|
||||
async with httpx.AsyncClient(timeout=30.0) as c:
|
||||
r = await c.post(f"{base}/prompt", json={"prompt": payload.get("prompt", payload),
|
||||
"client_id": "hyperswap-jobs"})
|
||||
if r.status_code != 200:
|
||||
return {"ok": False, "error": f"HTTP {r.status_code}: {r.text[:200]}"}
|
||||
prompt_id = r.json().get("prompt_id")
|
||||
deadline = time.time() + payload.get("timeout_s", 1800)
|
||||
while time.time() < deadline:
|
||||
await asyncio.sleep(0.5)
|
||||
h = await c.get(f"{base}/history/{prompt_id}")
|
||||
entry = (h.json() or {}).get(prompt_id) if h.status_code == 200 else None
|
||||
if not entry:
|
||||
continue
|
||||
status = entry.get("status", {})
|
||||
if status.get("status_str") == "error":
|
||||
return {"ok": False, "error": "ComfyUI reported an execution error"}
|
||||
if status.get("completed"):
|
||||
return {"ok": True, "prompt_id": prompt_id}
|
||||
return {"ok": False, "error": "timed out waiting for ComfyUI"}
|
||||
|
||||
|
||||
async def _dispatch_ollama(payload: Dict[str, Any]) -> Dict[str, Any]:
|
||||
async with httpx.AsyncClient(timeout=payload.get("timeout_s", 1800)) as c:
|
||||
body = {"stream": False, "keep_alive": payload.get("keep_alive", "5m"), **payload}
|
||||
body.pop("timeout_s", None)
|
||||
r = await c.post("http://localhost:11434/api/generate", json=body)
|
||||
if r.status_code != 200:
|
||||
return {"ok": False, "error": f"HTTP {r.status_code}: {r.text[:200]}"}
|
||||
d = r.json()
|
||||
return {"ok": True, "response": (d.get("response") or "")[:2000],
|
||||
"eval_count": d.get("eval_count"),
|
||||
"tokens_per_sec": (round(d.get("eval_count", 0)
|
||||
/ (d.get("eval_duration", 1) / 1e9), 2)
|
||||
if d.get("eval_duration") else None)}
|
||||
|
||||
|
||||
DISPATCHERS = {
|
||||
tenants_mod.KIND_DIFFUSION: _dispatch_comfy,
|
||||
tenants_mod.KIND_LLM: _dispatch_ollama,
|
||||
}
|
||||
|
||||
|
||||
class Scheduler:
|
||||
"""Drains the queue, making room for each job before it runs.
|
||||
|
||||
One job at a time by design. The GPU is the scarce resource this whole service
|
||||
exists to hand between applications; running two jobs concurrently would just
|
||||
recreate the contention it is meant to resolve. Throughput comes from swapping
|
||||
quickly, not from overlapping.
|
||||
"""
|
||||
|
||||
def __init__(self) -> None:
|
||||
self.running = False
|
||||
self.task: Optional[asyncio.Task] = None
|
||||
self.current: Optional[Dict[str, Any]] = None
|
||||
self.last_finished: Optional[Dict[str, Any]] = None
|
||||
self.completed = 0
|
||||
self.failed = 0
|
||||
self.waits = 0
|
||||
self.blocked: Optional[Dict[str, Any]] = None
|
||||
self.idle_poll_s = 1.0
|
||||
self.blocked_poll_s = 2.0
|
||||
# How long a job may wait for VRAM before it is declared impossible. Long enough
|
||||
# to outlast a normal diffusion run, short enough not to wedge the queue.
|
||||
self.max_block_s = 120.0
|
||||
|
||||
async def start(self) -> None:
|
||||
if self.running:
|
||||
return
|
||||
init()
|
||||
# A job left RUNNING by a crash or a hard restart would sit there forever.
|
||||
requeued = requeue_orphans()
|
||||
if requeued:
|
||||
logger.warning(f"requeued {requeued} job(s) left running by a previous process")
|
||||
self.running = True
|
||||
self.task = asyncio.create_task(self._loop())
|
||||
logger.info("job scheduler started")
|
||||
|
||||
async def stop(self) -> None:
|
||||
self.running = False
|
||||
if self.task:
|
||||
self.task.cancel()
|
||||
|
||||
def _job_vram_requirement(self, tenant_name: str, payload: Dict[str, Any]) -> float:
|
||||
"""How much VRAM *this* job needs, not the tenant's generic figure.
|
||||
|
||||
A tenant-wide needs_vram_gb cannot be right for an LLM: the requirement is a
|
||||
property of the model being loaded. Ollama's generic 4 GB passed the room check
|
||||
with 8 GB free, and then a 14.9 GB model was dispatched into it and killed
|
||||
llama-server with a CUDA OOM -- three queued jobs destroyed in a row.
|
||||
"""
|
||||
import vram_arbitrator
|
||||
|
||||
t = tenants_mod.get_tenant(tenant_name)
|
||||
default = t.needs_vram_gb if t else 0.0
|
||||
model = payload.get("model")
|
||||
if t and t.kind == tenants_mod.KIND_LLM and model:
|
||||
size = vram_arbitrator._model_size_bytes(model)
|
||||
if size:
|
||||
# Measured on this box: a 12.87 GB blob occupies 14.9 GB once context
|
||||
# and KV cache are allocated.
|
||||
return round((size / (1024 ** 3)) * 1.16, 2)
|
||||
return default
|
||||
|
||||
async def _make_room(self, tenant_name: str,
|
||||
payload: Optional[Dict[str, Any]] = None) -> Dict[str, Any]:
|
||||
"""Ensure the job's tenant has the VRAM it needs, using the normal arbitration."""
|
||||
import vram_arbitrator # imported late: it imports this module's siblings
|
||||
|
||||
t = tenants_mod.get_tenant(tenant_name)
|
||||
needed = self._job_vram_requirement(tenant_name, payload or {})
|
||||
if not t or not needed:
|
||||
return {"ready": True, "reason": "no VRAM requirement declared"}
|
||||
|
||||
state = await vram_arbitrator.arbitrator._tenant_state()
|
||||
free_gb = vram_arbitrator.arbitrator._last_tenant_state["free_gb"]
|
||||
held = next((s["vram_gb"] for s in state if s["name"] == tenant_name), 0.0)
|
||||
if held + free_gb >= needed:
|
||||
return {"ready": True, "needed_gb": needed,
|
||||
"reason": f"{free_gb:.2f} GB free, job needs {needed:.2f} GB"}
|
||||
|
||||
plan = tenants_mod.plan_release(tenant_name, state, free_gb, needed)
|
||||
for victim in plan["release"]:
|
||||
await vram_arbitrator.arbitrator._release_tenant(
|
||||
victim, f"queued job for '{tenant_name}'")
|
||||
if plan["release"]:
|
||||
# Give the driver a moment to actually hand the memory back.
|
||||
deadline = time.perf_counter() + 30
|
||||
target = int(needed * (1024 ** 3))
|
||||
while time.perf_counter() < deadline:
|
||||
if vram_arbitrator.get_process_vram_bytes()["free_bytes"] >= target:
|
||||
break
|
||||
await asyncio.sleep(0.05)
|
||||
# Only ready once the memory is genuinely there. A plan that *could* work is not
|
||||
# the same as VRAM that *is* free, and dispatching on the former is what OOMs.
|
||||
free_now = (vram_arbitrator.get_process_vram_bytes()["free_bytes"] / (1024 ** 3))
|
||||
# The best this GPU could ever offer this tenant: everything currently free, plus
|
||||
# what it already holds, plus everything that is reclaimable at all.
|
||||
reclaimable_gb = sum(s["vram_gb"] for s in state
|
||||
if s["name"] != tenant_name and s.get("reclaimable"))
|
||||
max_possible = round(held + free_now + reclaimable_gb, 2)
|
||||
return {"ready": plan["possible"] and (held + free_now) >= needed,
|
||||
"max_possible_gb": max_possible,
|
||||
"released": plan["release"], "needed_gb": needed,
|
||||
"free_gb": round(free_now, 2),
|
||||
"reason": f"job needs {needed:.2f} GB; {plan['reason']}",
|
||||
"blockers": plan.get("blockers")}
|
||||
|
||||
async def _loop(self) -> None:
|
||||
while self.running:
|
||||
try:
|
||||
job = _next_job()
|
||||
if not job:
|
||||
await asyncio.sleep(self.idle_poll_s)
|
||||
continue
|
||||
|
||||
tenant = tenants_mod.get_tenant(job["tenant"])
|
||||
dispatcher = DISPATCHERS.get(tenant.kind) if tenant else None
|
||||
if not dispatcher:
|
||||
_mark(job["id"], FAILED, finished_at=time.time(),
|
||||
error=f"no dispatcher for tenant kind "
|
||||
f"'{tenant.kind if tenant else '?'}'")
|
||||
self.failed += 1
|
||||
continue
|
||||
|
||||
# Check for room *before* claiming the job. Dispatching into
|
||||
# insufficient VRAM does not fail gracefully -- it kills llama-server
|
||||
# with a CUDA OOM, which is how three queued LLM jobs were destroyed
|
||||
# while ComfyUI legitimately held the card. A job that cannot run yet
|
||||
# waits; it does not fail.
|
||||
room = await self._make_room(job["tenant"], job.get("payload") or {})
|
||||
if not room.get("ready"):
|
||||
since = (self.blocked.get("since", time.time())
|
||||
if self.blocked and self.blocked.get("id") == job["id"]
|
||||
else time.time())
|
||||
waited = time.time() - since
|
||||
self.blocked = {"id": job["id"], "tenant": job["tenant"],
|
||||
"reason": room.get("reason"),
|
||||
"blockers": room.get("blockers"),
|
||||
"needed_gb": room.get("needed_gb"),
|
||||
"since": since, "waited_s": round(waited, 1)}
|
||||
|
||||
# Waiting is right while the memory might still arrive. It is wrong
|
||||
# when the job can never fit -- three LLM jobs sat pending forever
|
||||
# needing 14.93 GB on a card where only ~14.8 GB can ever be free,
|
||||
# because an unreclaimable process holds 0.82 GB. Say so and move on
|
||||
# rather than blocking the queue behind an impossibility.
|
||||
if waited > self.max_block_s:
|
||||
ceiling = room.get("max_possible_gb")
|
||||
detail = (f"needs {room.get('needed_gb')} GB but at most "
|
||||
f"{ceiling} GB can ever be free on this GPU"
|
||||
if ceiling is not None and room.get("needed_gb", 0) > ceiling
|
||||
else f"waited {int(waited)}s for VRAM: {room.get('reason')}")
|
||||
blockers = ", ".join(
|
||||
f"{b['name']} ({b['vram_gb']} GB, {b['why']})"
|
||||
for b in (room.get("blockers") or []))
|
||||
_mark(job["id"], FAILED, finished_at=time.time(),
|
||||
error=f"{detail}{'; blocked by ' + blockers if blockers else ''}")
|
||||
self.failed += 1
|
||||
self.blocked = None
|
||||
logger.warning(f"job {job['id']} cannot run: {detail}")
|
||||
continue
|
||||
|
||||
self.waits += 1
|
||||
await asyncio.sleep(self.blocked_poll_s)
|
||||
continue
|
||||
self.blocked = None
|
||||
|
||||
t0 = time.time()
|
||||
_mark(job["id"], RUNNING, started_at=t0)
|
||||
self.current = {**job, "state": RUNNING, "room": room,
|
||||
"started_at": t0}
|
||||
logger.info(f"running job {job['id']} for '{job['tenant']}' "
|
||||
f"({room.get('reason')})")
|
||||
|
||||
# Any failure here must land on the job. An exception used to escape to
|
||||
# the loop's handler, leaving the row RUNNING forever while the scheduler
|
||||
# moved on -- an orphan that never completed and never freed its slot.
|
||||
try:
|
||||
res = await dispatcher(job["payload"])
|
||||
except asyncio.CancelledError:
|
||||
_mark(job["id"], PENDING, started_at=None)
|
||||
self.current = None
|
||||
raise
|
||||
except Exception as e:
|
||||
res = {"ok": False, "error": f"dispatch raised: {e}"}
|
||||
|
||||
finished = time.time()
|
||||
if res.get("ok"):
|
||||
_mark(job["id"], DONE, finished_at=finished,
|
||||
result=json.dumps(res))
|
||||
self.completed += 1
|
||||
else:
|
||||
_mark(job["id"], FAILED, finished_at=finished,
|
||||
error=str(res.get("error"))[:500])
|
||||
self.failed += 1
|
||||
self.last_finished = {"id": job["id"], "tenant": job["tenant"],
|
||||
"ok": bool(res.get("ok")),
|
||||
"duration_s": round(finished - t0, 2),
|
||||
"made_room": room.get("released") or []}
|
||||
self.current = None
|
||||
except asyncio.CancelledError:
|
||||
raise
|
||||
except Exception as e:
|
||||
logger.error(f"scheduler error: {e}")
|
||||
self.current = None
|
||||
await asyncio.sleep(1.0)
|
||||
|
||||
def get_status(self) -> Dict[str, Any]:
|
||||
return {
|
||||
"running": self.running,
|
||||
"current": self.current,
|
||||
"last_finished": self.last_finished,
|
||||
"completed": self.completed,
|
||||
"failed": self.failed,
|
||||
"waits": self.waits,
|
||||
"blocked": self.blocked,
|
||||
**stats(),
|
||||
}
|
||||
|
||||
|
||||
scheduler = Scheduler()
|
||||
@@ -8,7 +8,9 @@ from typing import Dict, List, Any, Optional
|
||||
|
||||
from mcp.server import MCPServer
|
||||
import autotune
|
||||
import engines
|
||||
import health
|
||||
import jobs as jobs_mod
|
||||
import overclock_manager
|
||||
import ram_optimizer
|
||||
import telemetry_store
|
||||
@@ -137,6 +139,41 @@ def set_gpu_fan_speed(mode: str = "auto", percent: Optional[int] = None) -> str:
|
||||
res = overclock_manager.set_fan_auto()
|
||||
return json.dumps(res, indent=2)
|
||||
|
||||
@mcp.tool()
|
||||
async def get_engine_config() -> str:
|
||||
"""Live configuration of Ollama and ComfyUI (parallelism, max loaded models,
|
||||
keep-alive, KV cache type, ComfyUI VRAM mode and allocator), with what each setting
|
||||
implies for VRAM arbitration."""
|
||||
return json.dumps(await engines.get_engine_config(), indent=2, default=str)
|
||||
|
||||
|
||||
@mcp.tool()
|
||||
def queue_job(tenant: str, payload: Dict[str, Any], priority: Optional[int] = None,
|
||||
label: Optional[str] = None) -> str:
|
||||
"""Queue work for a GPU application without waiting for it.
|
||||
|
||||
tenant: 'ollama' (payload: model, prompt, options) or 'comfyui' (payload: {"prompt":
|
||||
<workflow>}). The queue is on disk, so there is no depth limit; jobs run one at a
|
||||
time, highest priority first, with VRAM arbitrated before each starts."""
|
||||
return json.dumps(jobs_mod.submit(tenant, payload, priority, label), indent=2,
|
||||
default=str)
|
||||
|
||||
|
||||
@mcp.tool()
|
||||
def get_job_queue(state: Optional[str] = None, limit: int = 50) -> str:
|
||||
"""Queued and recent jobs, plus what the scheduler is doing and why it may be
|
||||
waiting. Pending jobs are listed in the order they will run."""
|
||||
return json.dumps({"jobs": jobs_mod.listing(state, limit),
|
||||
"scheduler": jobs_mod.scheduler.get_status()},
|
||||
indent=2, default=str)
|
||||
|
||||
|
||||
@mcp.tool()
|
||||
def cancel_job(job_id: str) -> str:
|
||||
"""Cancel a job that has not started. Running work is never killed."""
|
||||
return json.dumps(jobs_mod.cancel(job_id), indent=2, default=str)
|
||||
|
||||
|
||||
@mcp.tool()
|
||||
async def check_system_health() -> str:
|
||||
"""Check every dependency HyperSwap needs (NVML, sudo nvidia-smi, fan control via the
|
||||
|
||||
@@ -1,7 +1,6 @@
|
||||
{
|
||||
"ollama": {
|
||||
"label": "Ollama — LLM decode (memory-bandwidth bound; measured insensitive to power and clocks)",
|
||||
"measured": "73.0-73.5 tok/s flat from 222W to 370W (qwen3.8long, 2026-08-28). Actual draw never exceeded 224W at any limit. Memory clock lock made no difference (72.6 locked vs 72.7 unlocked).",
|
||||
"label": "Ollama \u2014 LLM decode (memory-bandwidth bound; measured insensitive to power and clocks)",
|
||||
"power_limit_w": 320,
|
||||
"core_offset_mhz": 0,
|
||||
"mem_offset_mhz": 0,
|
||||
@@ -9,11 +8,11 @@
|
||||
"lock_core_max": 0,
|
||||
"lock_mem_mhz": 0,
|
||||
"fan_mode": "auto",
|
||||
"fan_speed_pct": 0
|
||||
"fan_speed_pct": 0,
|
||||
"measured": "73.0-73.5 tok/s flat from 222W to 370W (qwen3.8long, 2026-08-28). Actual draw never exceeded 224W at any limit. Memory clock lock made no difference (72.6 locked vs 72.7 unlocked)."
|
||||
},
|
||||
"comfy": {
|
||||
"label": "ComfyUI — diffusion (compute bound; genuinely power-scaling)",
|
||||
"measured": "SDXL 1024/20-step: 5.48 it/s @222W, 6.22 @259W, 6.50 @296W, 6.52 @320W, 6.63 @333W, 6.71 @370W (2026-08-28). Worth +2.8% over the 320W stock default. Core clock lock made no difference across 2400-3105 MHz.",
|
||||
"label": "ComfyUI \u2014 diffusion (compute bound; genuinely power-scaling)",
|
||||
"power_limit_w": 370,
|
||||
"core_offset_mhz": 0,
|
||||
"mem_offset_mhz": 0,
|
||||
@@ -21,18 +20,19 @@
|
||||
"lock_core_max": 0,
|
||||
"lock_mem_mhz": 0,
|
||||
"fan_mode": "auto",
|
||||
"fan_speed_pct": 0
|
||||
"fan_speed_pct": 0,
|
||||
"measured": "SDXL 1024/20-step: 5.48 it/s @222W, 6.22 @259W, 6.50 @296W, 6.52 @320W, 6.63 @333W, 6.71 @370W (2026-08-28). Worth +2.8% over the 320W stock default. Core clock lock made no difference across 2400-3105 MHz."
|
||||
},
|
||||
"balanced": {
|
||||
"label": "Balanced — stock power and boost, automatic fans",
|
||||
"measured": "Card's own design point. 48k telemetry samples show 67.8C average under load at 39.5% auto fan, 81C all-time max, zero thermal throttle events.",
|
||||
"power_limit_w": 320,
|
||||
"core_offset_mhz": 0,
|
||||
"mem_offset_mhz": 0,
|
||||
"label": "Balanced \u2014 stock power and boost, automatic fans",
|
||||
"power_limit_w": 340,
|
||||
"core_offset_mhz": 10,
|
||||
"mem_offset_mhz": 150,
|
||||
"lock_core_min": 0,
|
||||
"lock_core_max": 0,
|
||||
"lock_mem_mhz": 0,
|
||||
"fan_mode": "auto",
|
||||
"fan_speed_pct": 0
|
||||
"fan_mode": "manual",
|
||||
"fan_speed_pct": 95,
|
||||
"measured": "Card's own design point. 48k telemetry samples show 67.8C average under load at 39.5% auto fan, 81C all-time max, zero thermal throttle events."
|
||||
}
|
||||
}
|
||||
170
server.py
170
server.py
@@ -16,10 +16,13 @@ from fastapi.middleware.cors import CORSMiddleware
|
||||
from pydantic import BaseModel, Field
|
||||
|
||||
import autotune
|
||||
import engines
|
||||
import health
|
||||
import jobs as jobs_mod
|
||||
import overclock_manager
|
||||
import ram_optimizer
|
||||
import telemetry_store
|
||||
import tenants as tenants_mod
|
||||
import thermal_governor
|
||||
import vram_arbitrator
|
||||
|
||||
@@ -218,10 +221,13 @@ def _install_shutdown_hook() -> None:
|
||||
@asynccontextmanager
|
||||
async def lifespan(app: FastAPI):
|
||||
telemetry_store.start()
|
||||
jobs_mod.init()
|
||||
await broker.start()
|
||||
await jobs_mod.scheduler.start()
|
||||
await vram_arbitrator.arbitrator.start()
|
||||
_install_shutdown_hook()
|
||||
yield
|
||||
await jobs_mod.scheduler.stop()
|
||||
await vram_arbitrator.arbitrator.stop()
|
||||
await broker.stop()
|
||||
# Never leave the card with locked clocks and pinned fans after we exit.
|
||||
@@ -321,6 +327,141 @@ async def api_health():
|
||||
return await health.run_health_checks()
|
||||
|
||||
|
||||
class TenantReleaseRequest(BaseModel):
|
||||
models: Optional[List[str]] = Field(None, description="For per-model tenants (Ollama), which to unload; defaults to everything resident")
|
||||
confirm: bool = Field(True, description="Wait for NVML to confirm the VRAM was actually released")
|
||||
|
||||
|
||||
class JobRequest(BaseModel):
|
||||
tenant: str = Field(..., description="Which application should run this job", example="comfyui")
|
||||
payload: Dict[str, Any] = Field(..., description="What to run: a ComfyUI workflow under 'prompt', or Ollama generate parameters")
|
||||
priority: Optional[int] = Field(None, description="Defaults to the tenant's priority; higher runs sooner")
|
||||
label: Optional[str] = Field(None, description="Human-readable name for the queue view")
|
||||
|
||||
|
||||
@app.post("/api/jobs", summary="Queue a Job", tags=["Jobs"])
|
||||
async def api_submit_job(req: JobRequest):
|
||||
"""Queue work for any tenant. The queue is on disk, so there is no depth limit and
|
||||
it survives a restart. Jobs run one at a time, highest priority first, with VRAM
|
||||
arbitrated before each one starts."""
|
||||
res = jobs_mod.submit(req.tenant, req.payload, req.priority, req.label)
|
||||
if not res.get("success"):
|
||||
raise HTTPException(status_code=400, detail=res.get("error"))
|
||||
return res
|
||||
|
||||
|
||||
@app.get("/api/jobs", summary="The Job Queue", tags=["Jobs"])
|
||||
async def api_jobs(state: Optional[str] = Query(None, description="pending | running | done | failed | cancelled"),
|
||||
limit: int = Query(100)):
|
||||
"""Queued and recent jobs. Pending jobs are listed in the order they will run."""
|
||||
return {"jobs": jobs_mod.listing(state, limit), "scheduler": jobs_mod.scheduler.get_status()}
|
||||
|
||||
|
||||
@app.get("/api/jobs/{job_id}", summary="One Job", tags=["Jobs"])
|
||||
async def api_job(job_id: str):
|
||||
job = jobs_mod.get(job_id)
|
||||
if not job:
|
||||
raise HTTPException(status_code=404, detail=f"no job '{job_id}'")
|
||||
return job
|
||||
|
||||
|
||||
@app.delete("/api/jobs/{job_id}", summary="Cancel a Pending Job", tags=["Jobs"])
|
||||
async def api_cancel_job(job_id: str):
|
||||
"""Cancel a job that has not started. Running jobs are left alone -- this service
|
||||
frees VRAM by asking, never by killing work in flight."""
|
||||
res = jobs_mod.cancel(job_id)
|
||||
if not res.get("success"):
|
||||
raise HTTPException(status_code=409, detail=res.get("error"))
|
||||
return res
|
||||
|
||||
|
||||
@app.delete("/api/jobs", summary="Cancel All Pending Jobs", tags=["Jobs"])
|
||||
async def api_clear_jobs():
|
||||
return jobs_mod.clear_pending()
|
||||
|
||||
|
||||
@app.get("/api/tenants", summary="GPU Tenants", tags=["Tenants"])
|
||||
async def api_tenants():
|
||||
"""Applications competing for the GPU, as configured.
|
||||
|
||||
Each entry declares how its processes are recognised, how to tell whether it is
|
||||
working, and how to ask it for VRAM back. Adding an application is a config change
|
||||
in tenants.json, not a code change.
|
||||
"""
|
||||
gpu = vram_arbitrator.get_gpu_hardware_stats()
|
||||
by_tenant = (gpu.get("breakdown", {}) or {}).get("by_tenant_gb", {})
|
||||
out = []
|
||||
for t in tenants_mod.describe():
|
||||
name = t["name"]
|
||||
# The two original tenants are reported under the bucket names the API has
|
||||
# always used.
|
||||
bucket = {"comfyui": "comfy"}.get(name, name)
|
||||
t["vram_gb"] = by_tenant.get(bucket, 0.0)
|
||||
out.append(t)
|
||||
return {"tenants": out, "config_path": tenants_mod.CONFIG_PATH,
|
||||
"unmanaged_gb": (gpu.get("breakdown", {}) or {}).get("unmanaged_gb", 0.0)}
|
||||
|
||||
|
||||
@app.get("/api/tenants/{name}", summary="One GPU Tenant", tags=["Tenants"])
|
||||
async def api_tenant(name: str):
|
||||
"""A single tenant's definition, current VRAM, and whether it is genuinely busy."""
|
||||
t = tenants_mod.get_tenant(name)
|
||||
if not t:
|
||||
raise HTTPException(status_code=404, detail=f"no tenant named '{name}'")
|
||||
gpu = vram_arbitrator.get_gpu_hardware_stats()
|
||||
bucket = {"comfyui": "comfy"}.get(name, name)
|
||||
vram_gb = (gpu.get("breakdown", {}) or {}).get("by_tenant_gb", {}).get(bucket, 0.0)
|
||||
busy = await tenants_mod.probe_busy(t, vram_gb=vram_gb)
|
||||
d = t.to_dict()
|
||||
d.update({"vram_gb": vram_gb, "reclaimable": t.reclaimable, "busy": busy})
|
||||
return d
|
||||
|
||||
|
||||
@app.post("/api/tenants/{name}/release", summary="Ask a Tenant for its VRAM", tags=["Tenants"])
|
||||
async def api_tenant_release(name: str, req: Optional[TenantReleaseRequest] = None):
|
||||
"""Release a tenant's VRAM using whatever mechanism that tenant declares.
|
||||
|
||||
This is the generic form of the Ollama soft-yield and the ComfyUI purge: the same
|
||||
request works for any application in the registry, including ones added later.
|
||||
"""
|
||||
t = tenants_mod.get_tenant(name)
|
||||
if not t:
|
||||
raise HTTPException(status_code=404, detail=f"no tenant named '{name}'")
|
||||
if not t.reclaimable:
|
||||
raise HTTPException(status_code=409,
|
||||
detail=f"'{name}' declares no way to release VRAM; its "
|
||||
f"memory cannot be reclaimed by this service")
|
||||
|
||||
models = req.models if req else None
|
||||
if t.release.per_model and not models:
|
||||
state = await vram_arbitrator.get_ollama_live_state()
|
||||
models = [m.get("name") for m in state.get("loaded_models", []) if m.get("name")]
|
||||
|
||||
before = vram_arbitrator.get_process_vram_bytes()
|
||||
res = await tenants_mod.release_vram(t, models=models)
|
||||
|
||||
if (req is None or req.confirm) and res.get("released"):
|
||||
bucket = {"comfyui": "comfy"}.get(name, name)
|
||||
key = {"ollama": "ollama_bytes", "comfy": "comfyui_bytes"}.get(bucket)
|
||||
if key:
|
||||
baseline = before[key]
|
||||
barrier = await vram_arbitrator._await_vram_release(baseline) \
|
||||
if key == "ollama_bytes" else None
|
||||
if barrier:
|
||||
res.update({"outcome": barrier.get("outcome"),
|
||||
"confirm_ms": barrier.get("confirm_ms")})
|
||||
after = vram_arbitrator.get_process_vram_bytes()
|
||||
res["free_vram_gb"] = round(after["free_bytes"] / (1024**3), 2)
|
||||
return res
|
||||
|
||||
|
||||
@app.get("/api/engines", summary="Live Engine Configuration", tags=["Telemetry"])
|
||||
async def api_engines():
|
||||
"""Real configuration of Ollama and ComfyUI, with what each setting implies for
|
||||
arbitration. These live outside this codebase but dictate how it must behave."""
|
||||
return await engines.get_engine_config()
|
||||
|
||||
|
||||
@app.get("/api/gpu", summary="GPU Sensors and VRAM Breakdown", tags=["Telemetry"])
|
||||
async def get_gpu_metrics() -> Dict[str, Any]:
|
||||
"""Detailed NVML sensors (utilization, temp, power, fan, clocks, throttle reasons, per-process VRAM)."""
|
||||
@@ -375,7 +516,18 @@ async def api_switch_model(req: SwitchRequest):
|
||||
await vram_arbitrator.arbitrator.request_vram_for_ollama()
|
||||
res = await vram_arbitrator.switch_ollama_model(req.model, keep_alive=req.keep_alive or "30m")
|
||||
if not res.get("success"):
|
||||
raise HTTPException(status_code=500, detail=res.get("error"))
|
||||
# Reflect what actually went wrong. Ollama returns 400 for an unusable request --
|
||||
# asking an embedding model to generate, say -- and reporting that as 500 blames
|
||||
# this service for the caller's mistake. A model that will not fit is neither:
|
||||
# the request is valid and the service is healthy, there is simply no room.
|
||||
upstream = res.get("upstream_status")
|
||||
if res.get("vram_oom"):
|
||||
status = 507 # Insufficient Storage
|
||||
elif isinstance(upstream, int) and 400 <= upstream < 500:
|
||||
status = 400
|
||||
else:
|
||||
status = 502 if upstream else 500
|
||||
raise HTTPException(status_code=status, detail=res.get("error"))
|
||||
return res
|
||||
|
||||
@app.post("/api/free-vram", summary="Soft-Yield Ollama VRAM", tags=["Orchestration"])
|
||||
@@ -595,9 +747,23 @@ app.mount("/static", StaticFiles(directory=f"{BASE_DIR}/static"), name="static")
|
||||
|
||||
@app.get("/", summary="Dashboard Web UI", tags=["UI"])
|
||||
async def root_index():
|
||||
"""Serve the dashboard with cache-busted asset URLs.
|
||||
|
||||
StaticFiles sends an ETag, but browsers were still serving app.js from cache after
|
||||
it changed, so a reload showed the old dashboard against the new API -- a panel that
|
||||
had just been added simply never appeared. Stamping each asset with its mtime means
|
||||
a changed file is always a different URL.
|
||||
"""
|
||||
with open(f"{BASE_DIR}/static/index.html", "r") as f:
|
||||
content = f.read()
|
||||
return HTMLResponse(content=content)
|
||||
for asset in ("app.js", "styles.css"):
|
||||
try:
|
||||
stamp = int(os.path.getmtime(f"{BASE_DIR}/static/{asset}"))
|
||||
except OSError:
|
||||
continue
|
||||
content = content.replace(f"/static/{asset}", f"/static/{asset}?v={stamp}")
|
||||
return HTMLResponse(content=content,
|
||||
headers={"Cache-Control": "no-cache, must-revalidate"})
|
||||
|
||||
if __name__ == "__main__":
|
||||
import uvicorn
|
||||
|
||||
191
static/app.js
191
static/app.js
@@ -36,6 +36,7 @@ function updateDashboard(data) {
|
||||
// Governor and arbitration state ride along in the shared snapshot.
|
||||
if (data.governor) renderGovernor(data.governor);
|
||||
if (data.arbitrator) renderArbitrator(data.arbitrator, data.gpu);
|
||||
if (data.arbitrator) renderTenants(data.arbitrator, data.gpu);
|
||||
|
||||
// 1. GPU VRAM Stats
|
||||
const gpu = data.gpu || {};
|
||||
@@ -1069,3 +1070,193 @@ async function fetchLastSwapFromStore() {
|
||||
document.addEventListener('DOMContentLoaded', () => {
|
||||
setTimeout(fetchLastSwapFromStore, 1500);
|
||||
});
|
||||
|
||||
|
||||
// ---------------------------------------------------------------- engine config
|
||||
|
||||
async function fetchEngineConfig() {
|
||||
// These subtitles used to be hardcoded. They happened to be accurate, which is worse
|
||||
// than being wrong: they would have stayed accurate-looking after the settings changed.
|
||||
try {
|
||||
const d = await (await fetch('/api/engines')).json();
|
||||
const o = document.getElementById('ollama-engine-sub');
|
||||
if (o && d.ollama) {
|
||||
const bits = [`Port :${d.ollama.port}`, d.ollama.summary];
|
||||
if (d.ollama.max_loaded_models) bits.push(`${d.ollama.max_loaded_models} model resident`);
|
||||
if (d.ollama.keep_alive) bits.push(`keep-alive ${d.ollama.keep_alive}`);
|
||||
o.textContent = bits.join(' // ');
|
||||
o.title = (d.ollama.settings || [])
|
||||
.filter(s => s.means)
|
||||
.map(s => `${s.key}=${s.value} — ${s.means}`)
|
||||
.join('\n');
|
||||
}
|
||||
const c = document.getElementById('comfy-engine-sub');
|
||||
if (c && d.comfyui && d.comfyui.online) {
|
||||
c.textContent = `Port :8188 // v${d.comfyui.version} // ${d.comfyui.summary}`;
|
||||
c.title = `torch ${d.comfyui.pytorch}\n${d.comfyui.device || ''}`;
|
||||
}
|
||||
} catch (e) { /* subtitles are cosmetic; never break the page over them */ }
|
||||
}
|
||||
|
||||
document.addEventListener('DOMContentLoaded', () => {
|
||||
fetchEngineConfig();
|
||||
setInterval(fetchEngineConfig, 120000);
|
||||
});
|
||||
|
||||
|
||||
// ---------------------------------------------------------------- chart sizing
|
||||
|
||||
// Chart.js is configured responsive with maintainAspectRatio:false, so it should track
|
||||
// its container on its own. In practice it latched onto a stale size -- the canvas sat
|
||||
// at width:0px, height:288px (the old fixed h-64) while its container had grown to
|
||||
// 988x648 on a large window. An explicit observer makes the chart follow the container
|
||||
// whatever the window does.
|
||||
function watchChartSize() {
|
||||
const canvas = document.getElementById('oc-chart');
|
||||
if (!canvas || !canvas.parentElement) return;
|
||||
const container = canvas.parentElement;
|
||||
|
||||
const resize = () => {
|
||||
if (typeof ocChart === 'undefined' || !ocChart) return;
|
||||
// No explicit dimensions: with responsive + maintainAspectRatio:false, Chart.js
|
||||
// measures the container itself. Passing width/height instead made the canvas grow
|
||||
// but never shrink -- it ended up 988px wide inside a 435px container, overflowing
|
||||
// it, which is also why the canvas is absolutely positioned now.
|
||||
ocChart.resize();
|
||||
};
|
||||
|
||||
if (typeof ResizeObserver !== 'undefined') {
|
||||
new ResizeObserver(resize).observe(container);
|
||||
}
|
||||
window.addEventListener('resize', resize);
|
||||
// Run once after layout settles, in case the chart was built before the container had
|
||||
// a width (which is how it ended up at 0 in the first place).
|
||||
requestAnimationFrame(resize);
|
||||
setTimeout(resize, 300);
|
||||
}
|
||||
|
||||
document.addEventListener('DOMContentLoaded', () => setTimeout(watchChartSize, 200));
|
||||
|
||||
|
||||
// ---------------------------------------------------------------- tenants
|
||||
|
||||
function renderTenants(arb, gpu) {
|
||||
const body = document.getElementById('tenants-body');
|
||||
if (!body) return;
|
||||
const state = arb && arb.tenant_state;
|
||||
if (!state || !state.tenants) return;
|
||||
|
||||
const total = (gpu && gpu.vram_total_gb) || 16;
|
||||
document.getElementById('tenants-free').textContent = `${state.free_gb} GB free`;
|
||||
|
||||
// Sorted by priority, the order arbitration actually considers them in.
|
||||
const rows = [...state.tenants].sort((a, b) => b.priority - a.priority);
|
||||
body.innerHTML = rows.map(t => {
|
||||
const pct = Math.min((t.vram_gb / total) * 100, 100);
|
||||
const bar = t.busy ? 'bg-emerald-500'
|
||||
: t.reclaimable ? 'bg-cyan-600' : 'bg-amber-600';
|
||||
const badge = t.busy
|
||||
? '<span class="text-emerald-400">working</span>'
|
||||
: t.reclaimable
|
||||
? '<span class="text-slate-500">idle · reclaimable</span>'
|
||||
: '<span class="text-amber-400">cannot be reclaimed</span>';
|
||||
return `<div>
|
||||
<div class="flex justify-between text-[11px] font-mono">
|
||||
<span class="text-slate-200">${t.name}
|
||||
<span class="text-slate-600">p${t.priority}</span></span>
|
||||
<span class="text-slate-400">${t.vram_gb.toFixed(2)} GB · ${badge}</span>
|
||||
</div>
|
||||
<div class="w-full bg-slate-950 rounded-full h-1.5 mt-1 overflow-hidden border border-slate-800/60">
|
||||
<div class="${bar} h-full transition-all duration-500" style="width:${pct}%"></div>
|
||||
</div>
|
||||
<div class="text-[10px] text-slate-600 mt-0.5">${t.reason || ''}${
|
||||
t.needs_vram_gb ? ` · needs ${t.needs_vram_gb} GB to work` : ''}</div>
|
||||
</div>`;
|
||||
}).join('');
|
||||
|
||||
// The most recent arbitration decision, including why it could not be satisfied.
|
||||
const dec = document.getElementById('tenants-decision');
|
||||
const a = arb.last_arbitration;
|
||||
if (!a) {
|
||||
dec.innerHTML = '<span class="text-slate-600">No contention — nothing has needed to be released.</span>';
|
||||
return;
|
||||
}
|
||||
const when = new Date(a.ts * 1000).toLocaleTimeString();
|
||||
const blockers = (a.blockers || [])
|
||||
.map(b => `${b.name} (${b.vram_gb} GB, ${b.why})`).join(', ');
|
||||
dec.innerHTML =
|
||||
`<span class="${a.possible ? 'text-cyan-400' : 'text-amber-400'}">${when} · ` +
|
||||
`${a.demanding} short by ${a.shortfall_gb} GB</span> — ${a.reason}` +
|
||||
(blockers ? `<div class="text-slate-600">blocked by: ${blockers}</div>` : '');
|
||||
}
|
||||
|
||||
|
||||
// ---------------------------------------------------------------- job queue
|
||||
|
||||
async function fetchQueue() {
|
||||
const body = document.getElementById('queue-body');
|
||||
if (!body) return;
|
||||
try {
|
||||
const d = await (await fetch('/api/jobs?limit=40')).json();
|
||||
const s = d.scheduler || {};
|
||||
document.getElementById('queue-summary').textContent =
|
||||
`${s.queue_depth ?? 0} queued · ${s.completed ?? 0} done · ${s.failed ?? 0} failed`;
|
||||
|
||||
// What the scheduler is doing right now, including why it is waiting. A job that
|
||||
// cannot get VRAM used to sit silent, which made a stuck queue indistinguishable
|
||||
// from an empty one.
|
||||
const cur = document.getElementById('queue-current');
|
||||
if (s.current) {
|
||||
cur.innerHTML = `<span class="text-emerald-400">running</span> ` +
|
||||
`<span class="text-slate-200">${s.current.tenant}</span>` +
|
||||
`<span class="text-slate-500"> · ${s.current.label || s.current.id}</span>` +
|
||||
(s.current.room && s.current.room.released && s.current.room.released.length
|
||||
? `<span class="text-cyan-400"> · released ${s.current.room.released.join(', ')}</span>` : '');
|
||||
} else if (s.blocked) {
|
||||
cur.innerHTML = `<span class="text-amber-400">waiting ${s.blocked.waited_s}s</span> ` +
|
||||
`<span class="text-slate-200">${s.blocked.tenant}</span>` +
|
||||
`<span class="text-slate-500"> · ${s.blocked.reason || ''}</span>`;
|
||||
} else {
|
||||
cur.innerHTML = '<span class="text-slate-600">scheduler idle</span>';
|
||||
}
|
||||
|
||||
const colour = {
|
||||
pending: 'text-slate-400', running: 'text-emerald-400', done: 'text-cyan-500',
|
||||
failed: 'text-rose-400', cancelled: 'text-slate-600',
|
||||
};
|
||||
body.innerHTML = (d.jobs || []).map(job => {
|
||||
const when = job.state === 'pending'
|
||||
? `waiting ${job.waiting_s}s`
|
||||
: (job.duration_s != null ? `${job.duration_s}s` : '');
|
||||
return `<div class="flex justify-between text-[11px] font-mono gap-2">
|
||||
<span class="truncate">
|
||||
<span class="${colour[job.state] || 'text-slate-400'}">${job.state}</span>
|
||||
<span class="text-slate-600"> p${job.priority}</span>
|
||||
<span class="text-slate-200"> ${job.tenant}</span>
|
||||
<span class="text-slate-500">${job.label ? ' · ' + job.label : ''}</span>
|
||||
</span>
|
||||
<span class="text-slate-500 whitespace-nowrap">${when}${
|
||||
job.state === 'pending'
|
||||
? ` <button onclick="cancelJob('${job.id}')" class="text-rose-500 hover:text-rose-300 ml-1">cancel</button>`
|
||||
: ''}</span>
|
||||
</div>${job.error ? `<div class="text-[10px] text-rose-500/80 pl-2 truncate">${job.error}</div>` : ''}`;
|
||||
}).join('') || '<div class="text-slate-600 text-[11px]">No jobs yet.</div>';
|
||||
} catch (e) {
|
||||
body.innerHTML = `<div class="text-rose-400 text-[11px]">${e}</div>`;
|
||||
}
|
||||
}
|
||||
|
||||
async function cancelJob(id) {
|
||||
await fetch(`/api/jobs/${id}`, { method: 'DELETE' });
|
||||
fetchQueue();
|
||||
}
|
||||
|
||||
async function clearQueue() {
|
||||
await fetch('/api/jobs', { method: 'DELETE' });
|
||||
fetchQueue();
|
||||
}
|
||||
|
||||
document.addEventListener('DOMContentLoaded', () => {
|
||||
fetchQueue();
|
||||
setInterval(fetchQueue, 2000);
|
||||
});
|
||||
|
||||
@@ -5,6 +5,11 @@
|
||||
<meta name="viewport" content="width=device-width, initial-scale=1.0">
|
||||
<title>HYPERSWAP // Dual-Engine Model Orchestrator & Live Telemetry</title>
|
||||
<script src="https://cdn.tailwindcss.com"></script>
|
||||
<script>
|
||||
// The CDN build has no breakpoint above 2xl, so a very wide window kept a
|
||||
// two-column layout with increasingly stretched panels.
|
||||
tailwind.config = { theme: { extend: { screens: { '3xl': '2000px' } } } };
|
||||
</script>
|
||||
<script src="https://cdn.jsdelivr.net/npm/chart.js@4.4.1/dist/chart.umd.min.js"></script>
|
||||
<link rel="stylesheet" href="/static/styles.css">
|
||||
<link rel="stylesheet" href="https://cdnjs.cloudflare.com/ajax/libs/font-awesome/6.4.0/css/all.min.css">
|
||||
@@ -13,7 +18,7 @@
|
||||
|
||||
<!-- TOP HEADER -->
|
||||
<header class="border-b border-slate-800 bg-slate-900/80 backdrop-blur sticky top-0 z-50">
|
||||
<div class="max-w-7xl mx-auto px-4 sm:px-6 lg:px-8 py-3 flex flex-wrap items-center justify-between gap-4">
|
||||
<div class="max-w-[2600px] mx-auto px-4 sm:px-6 lg:px-8 py-3 flex flex-wrap items-center justify-between gap-4">
|
||||
<div class="flex items-center space-x-3">
|
||||
<div class="w-10 h-10 rounded-xl bg-gradient-to-tr from-cyan-500 via-indigo-500 to-purple-500 flex items-center justify-center shadow-lg shadow-cyan-500/20">
|
||||
<i class="fa-solid fa-bolt-lightning text-white text-lg"></i>
|
||||
@@ -60,7 +65,7 @@
|
||||
</header>
|
||||
|
||||
<!-- MAIN CONTAINER -->
|
||||
<main class="max-w-7xl mx-auto px-4 sm:px-6 lg:px-8 py-6 space-y-6">
|
||||
<main class="max-w-[2600px] mx-auto px-4 sm:px-6 lg:px-8 py-6 space-y-6">
|
||||
|
||||
<!-- HERO MEMORY GAUGES -->
|
||||
<div class="grid grid-cols-1 md:grid-cols-2 gap-6">
|
||||
@@ -194,7 +199,7 @@
|
||||
</div>
|
||||
<div>
|
||||
<h3 class="font-bold text-slate-100 text-sm">Ollama LLM Engine</h3>
|
||||
<p class="text-xs text-slate-400">Port :11434 // FlashAttention + Q4 KV Cache</p>
|
||||
<p class="text-xs text-slate-400" id="ollama-engine-sub" title="Read live from the ollama service environment">Port :11434</p>
|
||||
</div>
|
||||
</div>
|
||||
<button onclick="freeOllamaVRAM()" class="px-2.5 py-1 text-xs font-semibold rounded-lg bg-rose-950/70 border border-rose-800 text-rose-300 hover:bg-rose-900 transition flex items-center space-x-1">
|
||||
@@ -265,7 +270,7 @@
|
||||
</div>
|
||||
<div>
|
||||
<h3 class="font-bold text-slate-100 text-sm">ComfyUI Diffusion Engine</h3>
|
||||
<p class="text-xs text-slate-400">Port :8188 // DynamicVRAM + Pinned Async Offload</p>
|
||||
<p class="text-xs text-slate-400" id="comfy-engine-sub" title="Read live from ComfyUI's /system_stats">Port :8188</p>
|
||||
</div>
|
||||
</div>
|
||||
<button onclick="freeComfyVRAM()" class="px-2.5 py-1 text-xs font-semibold rounded-lg bg-rose-950/70 border border-rose-800 text-rose-300 hover:bg-rose-900 transition flex items-center space-x-1">
|
||||
@@ -431,7 +436,7 @@
|
||||
</div>
|
||||
|
||||
<!-- OVERCLOCK CONTROL PANEL -->
|
||||
<div class="bg-slate-900/80 border border-fuchsia-900/50 rounded-2xl p-5 space-y-4 shadow-lg shadow-fuchsia-950/30">
|
||||
<div class="lg:col-span-2 bg-slate-900/80 border border-fuchsia-900/50 rounded-2xl p-5 space-y-4 shadow-lg shadow-fuchsia-950/30">
|
||||
<div class="flex flex-wrap items-center justify-between gap-3 pb-3 border-b border-slate-800">
|
||||
<div class="flex items-center space-x-2">
|
||||
<div class="p-2 rounded-lg bg-fuchsia-950/80 border border-fuchsia-800 text-fuchsia-400">
|
||||
@@ -599,8 +604,8 @@
|
||||
<span class="flex items-center space-x-1"><span class="w-2 h-2 rounded-full bg-fuchsia-400 inline-block"></span>RAM Cache GB</span>
|
||||
</div>
|
||||
</div>
|
||||
<div class="relative h-64">
|
||||
<canvas id="oc-chart"></canvas>
|
||||
<div class="relative h-[clamp(18rem,45vh,52rem)]">
|
||||
<canvas id="oc-chart" class="absolute inset-0 !w-full !h-full"></canvas>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
@@ -672,9 +677,50 @@
|
||||
</div>
|
||||
|
||||
<!-- ============ NEXT-LEVEL PANELS: governor / residency / analytics / autotune ============ -->
|
||||
<div class="grid grid-cols-1 xl:grid-cols-2 gap-5 mt-5">
|
||||
<div class="grid grid-cols-1 xl:grid-cols-2 3xl:grid-cols-3 gap-5 mt-5">
|
||||
|
||||
|
||||
|
||||
|
||||
<!-- The cross-application job queue -->
|
||||
<div class="bg-slate-900/80 border border-slate-800 rounded-2xl p-5 xl:col-span-2">
|
||||
<div class="flex items-center justify-between pb-3 border-b border-slate-800">
|
||||
<div class="flex items-center space-x-2">
|
||||
<div class="p-2 rounded-lg bg-violet-950/80 border border-violet-800 text-violet-400">
|
||||
<i class="fa-solid fa-list-check text-sm"></i>
|
||||
</div>
|
||||
<div>
|
||||
<h3 class="font-bold text-slate-100 text-sm">Job Queue</h3>
|
||||
<p class="text-xs text-slate-400">Work lined up across every application, highest priority first</p>
|
||||
</div>
|
||||
</div>
|
||||
<div class="flex items-center space-x-2">
|
||||
<span id="queue-summary" class="text-xs font-mono text-slate-500">—</span>
|
||||
<button onclick="clearQueue()" class="px-2.5 py-1 text-xs font-semibold rounded-lg bg-slate-800 border border-slate-700 text-slate-300 hover:bg-slate-700 transition">Clear pending</button>
|
||||
</div>
|
||||
</div>
|
||||
<div id="queue-current" class="mt-4 text-[11px] font-mono"></div>
|
||||
<div id="queue-body" class="mt-3 space-y-1.5 max-h-64 overflow-y-auto pr-1"></div>
|
||||
</div>
|
||||
|
||||
<!-- All GPU tenants, however many are configured -->
|
||||
<div class="bg-slate-900/80 border border-slate-800 rounded-2xl p-5 xl:col-span-2">
|
||||
<div class="flex items-center justify-between pb-3 border-b border-slate-800">
|
||||
<div class="flex items-center space-x-2">
|
||||
<div class="p-2 rounded-lg bg-teal-950/80 border border-teal-800 text-teal-400">
|
||||
<i class="fa-solid fa-layer-group text-sm"></i>
|
||||
</div>
|
||||
<div>
|
||||
<h3 class="font-bold text-slate-100 text-sm">GPU Tenants</h3>
|
||||
<p class="text-xs text-slate-400">Every application contending for the card, from <code class="text-teal-400">tenants.json</code></p>
|
||||
</div>
|
||||
</div>
|
||||
<span id="tenants-free" class="text-xs font-mono text-slate-500">—</span>
|
||||
</div>
|
||||
<div id="tenants-body" class="mt-4 space-y-2"></div>
|
||||
<div id="tenants-decision" class="mt-3 text-[11px] font-mono text-slate-400"></div>
|
||||
</div>
|
||||
|
||||
<!-- System health: makes a silently-broken dependency loud -->
|
||||
<div class="bg-slate-900/80 border border-slate-800 rounded-2xl p-5 xl:col-span-2">
|
||||
<div class="flex items-center justify-between pb-3 border-b border-slate-800">
|
||||
|
||||
108
tenants.json
Normal file
108
tenants.json
Normal file
@@ -0,0 +1,108 @@
|
||||
[
|
||||
{
|
||||
"name": "ollama",
|
||||
"kind": "llm",
|
||||
"priority": 50,
|
||||
"match": {
|
||||
"names": [
|
||||
"ollama"
|
||||
],
|
||||
"cmdline": [
|
||||
"llama-server",
|
||||
"ollama"
|
||||
]
|
||||
},
|
||||
"busy": {
|
||||
"type": "http_count",
|
||||
"url": "http://localhost:11434/api/ps",
|
||||
"count_keys": [
|
||||
"models"
|
||||
]
|
||||
},
|
||||
"release": {
|
||||
"type": "http_post",
|
||||
"url": "http://localhost:11434/api/generate",
|
||||
"body": {
|
||||
"keep_alive": 0
|
||||
},
|
||||
"per_model": true,
|
||||
"timeout_s": 120.0
|
||||
},
|
||||
"notes": "Unloads per model. With OLLAMA_NUM_PARALLEL=1 the request queues behind any running generation and applies when it finishes."
|
||||
},
|
||||
{
|
||||
"name": "comfyui",
|
||||
"kind": "diffusion",
|
||||
"priority": 60,
|
||||
"match": {
|
||||
"cmdline": [
|
||||
"comfyui",
|
||||
"comfy"
|
||||
],
|
||||
"cmdline_endswith": [
|
||||
"main.py"
|
||||
]
|
||||
},
|
||||
"busy": {
|
||||
"type": "http_count",
|
||||
"url": "http://127.0.0.1:8188/queue",
|
||||
"count_keys": [
|
||||
"queue_running",
|
||||
"queue_pending"
|
||||
],
|
||||
"vram_floor_gb": 1.5,
|
||||
"stale_after_s": 90.0
|
||||
},
|
||||
"release": {
|
||||
"type": "http_post",
|
||||
"url": "http://127.0.0.1:8188/free",
|
||||
"body": {
|
||||
"unload_models": true,
|
||||
"free_memory": true
|
||||
},
|
||||
"timeout_s": 30.0
|
||||
},
|
||||
"notes": "Leaves dead jobs in queue_running; the queue flag is corroborated against its own VRAM before being believed."
|
||||
},
|
||||
{
|
||||
"name": "stt-relay",
|
||||
"kind": "other",
|
||||
"priority": 70,
|
||||
"match": {
|
||||
"cmdline": [
|
||||
"stt_relay.py"
|
||||
]
|
||||
},
|
||||
"busy": {
|
||||
"type": "vram",
|
||||
"vram_busy_gb": 1.0
|
||||
},
|
||||
"release": {
|
||||
"type": "none"
|
||||
},
|
||||
"notes": "Long-running speech relay. Holds ~0.8 GB permanently and exposes no release API, so its VRAM is headroom this service can never offer. Declared so it is named rather than lumped into 'unmanaged'."
|
||||
},
|
||||
{
|
||||
"name": "desktop",
|
||||
"kind": "desktop",
|
||||
"priority": 90,
|
||||
"match": {
|
||||
"names": [
|
||||
"gnome-shell",
|
||||
"xorg",
|
||||
"mutter",
|
||||
"kwin",
|
||||
"plasmashell",
|
||||
"gnome-remote-desktop",
|
||||
"sddm",
|
||||
"gdm",
|
||||
"picom",
|
||||
"weston"
|
||||
]
|
||||
},
|
||||
"release": {
|
||||
"type": "none"
|
||||
},
|
||||
"notes": "Compositor and display server. Small, permanent, never reclaimable."
|
||||
}
|
||||
]
|
||||
461
tenants.py
Normal file
461
tenants.py
Normal file
@@ -0,0 +1,461 @@
|
||||
"""GPU tenants: the applications competing for the card, described as data.
|
||||
|
||||
The point of this service is fast handoff of a single GPU between applications. It grew
|
||||
up around the two on this box, and their names ended up compiled into process matching,
|
||||
VRAM attribution, busy detection and release calls alike -- roughly 385 references. That
|
||||
makes it a script for Ollama and ComfyUI rather than a GPU arbitrator.
|
||||
|
||||
A tenant is described here instead:
|
||||
|
||||
* how to recognise its processes (match)
|
||||
* how to tell whether it is actually working (busy probe)
|
||||
* how to ask it to give VRAM back (release strategy)
|
||||
* how much it matters when two want the card (priority)
|
||||
|
||||
Ollama and ComfyUI ship as defaults so behaviour is unchanged, but nothing about the
|
||||
arbitration logic knows their names. A third application -- a training run, a speech
|
||||
model, another inference server -- is a config entry, not a code change. A tenant that
|
||||
cannot be released (no API to ask) is still worth declaring, because naming it turns
|
||||
"unmanaged VRAM" into "held by X, which cannot be reclaimed".
|
||||
"""
|
||||
import json
|
||||
import logging
|
||||
import os
|
||||
import time
|
||||
from dataclasses import dataclass, field, asdict
|
||||
from typing import Any, Dict, List, Optional
|
||||
|
||||
import httpx
|
||||
import psutil
|
||||
|
||||
logger = logging.getLogger("tenants")
|
||||
|
||||
_BASE = os.path.dirname(os.path.abspath(__file__))
|
||||
CONFIG_PATH = os.environ.get("HYPERSWAP_TENANTS", os.path.join(_BASE, "tenants.json"))
|
||||
|
||||
# Kinds are advisory: they drive presentation and sensible defaults, never control flow.
|
||||
KIND_LLM, KIND_DIFFUSION, KIND_DESKTOP, KIND_OTHER = "llm", "diffusion", "desktop", "other"
|
||||
|
||||
|
||||
@dataclass
|
||||
class ProcessMatch:
|
||||
"""How to recognise a tenant's processes among those NVML reports."""
|
||||
names: List[str] = field(default_factory=list) # matched against process name
|
||||
cmdline: List[str] = field(default_factory=list) # substrings of the full cmdline
|
||||
cmdline_endswith: List[str] = field(default_factory=list)
|
||||
|
||||
def matches(self, pname: str, cmdline: str) -> bool:
|
||||
pname, cmdline = pname.lower(), cmdline.lower()
|
||||
if any(n.lower() in pname for n in self.names):
|
||||
return True
|
||||
if any(c.lower() in cmdline for c in self.cmdline):
|
||||
return True
|
||||
return any(cmdline.rstrip().endswith(c.lower()) for c in self.cmdline_endswith)
|
||||
|
||||
|
||||
@dataclass
|
||||
class BusyProbe:
|
||||
"""How to tell whether a tenant is genuinely working.
|
||||
|
||||
`vram_floor_gb` exists because a queue flag can lie: ComfyUI leaves dead jobs in
|
||||
queue_running, and only its VRAM reveals that nothing is loaded. GPU utilisation is
|
||||
deliberately unavailable as a signal -- it is shared by every tenant, so it cannot
|
||||
attribute work to one of them.
|
||||
"""
|
||||
type: str = "none" # none | http_count | vram
|
||||
url: Optional[str] = None
|
||||
count_keys: List[str] = field(default_factory=list) # keys whose lists are summed
|
||||
vram_busy_gb: float = 0.0 # busy when its VRAM exceeds this
|
||||
vram_floor_gb: float = 0.0 # below this it holds no real work
|
||||
stale_after_s: float = 90.0
|
||||
|
||||
|
||||
@dataclass
|
||||
class EventSource:
|
||||
"""A stream that tells us *when* to look, not what to think.
|
||||
|
||||
ComfyUI publishes a websocket, and the original listener parsed its message types to
|
||||
decide what was happening -- which meant understanding one application's schema. Any
|
||||
message is instead treated purely as a wake-up: re-run this tenant's busy probe now
|
||||
rather than waiting for the next poll. That gives sub-second reaction to any
|
||||
application with an event stream, with no knowledge of what it emits.
|
||||
"""
|
||||
type: str = "none" # none | websocket
|
||||
url: Optional[str] = None
|
||||
reconnect_backoff_s: float = 2.0
|
||||
max_backoff_s: float = 15.0
|
||||
|
||||
|
||||
@dataclass
|
||||
class ReleaseStrategy:
|
||||
"""How to ask a tenant to give VRAM back."""
|
||||
type: str = "none" # none | http_post
|
||||
url: Optional[str] = None
|
||||
body: Dict[str, Any] = field(default_factory=dict)
|
||||
# Set when the call must name the loaded model (Ollama unloads per model).
|
||||
per_model: bool = False
|
||||
timeout_s: float = 120.0
|
||||
confirm: bool = True # wait for NVML to show the memory released
|
||||
|
||||
|
||||
@dataclass
|
||||
class GpuTenant:
|
||||
name: str
|
||||
kind: str = KIND_OTHER
|
||||
enabled: bool = True
|
||||
# Higher wins contention; a tenant yields to anything above it.
|
||||
priority: int = 50
|
||||
# How much free VRAM this application needs before it can work. Used to decide
|
||||
# whether a busy tenant is actually being starved, rather than merely busy.
|
||||
needs_vram_gb: float = 0.0
|
||||
# How long a reclaimable tenant may sit idle holding VRAM before it is asked for it
|
||||
# back. Iterating on a ComfyUI workflow should not pay a reload between every run,
|
||||
# so this is deliberately not immediate.
|
||||
idle_release_after_s: float = 30.0
|
||||
# GPU profile to apply while this tenant is the active workload. Clock and power
|
||||
# tuning is workload-specific -- diffusion is compute bound, LLM decode is bandwidth
|
||||
# bound -- and that was previously switched by application name in the arbitrator.
|
||||
overclock_profile: Optional[str] = None
|
||||
# VRAM that survives a release. ComfyUI keeps its CUDA context for as long as the
|
||||
# process lives, so purging it does not return everything it holds. Ignoring this
|
||||
# made plan_release over-promise: it reported that releasing ComfyUI would free
|
||||
# 0.37 GB against a 0.33 GB shortfall, the job was cleared to run, and the memory
|
||||
# never actually arrived.
|
||||
vram_floor_gb: float = 0.0
|
||||
match: ProcessMatch = field(default_factory=ProcessMatch)
|
||||
busy: BusyProbe = field(default_factory=BusyProbe)
|
||||
release: ReleaseStrategy = field(default_factory=ReleaseStrategy)
|
||||
events: EventSource = field(default_factory=EventSource)
|
||||
notes: str = ""
|
||||
|
||||
@property
|
||||
def reclaimable(self) -> bool:
|
||||
return self.release.type != "none"
|
||||
|
||||
def to_dict(self) -> Dict[str, Any]:
|
||||
return asdict(self)
|
||||
|
||||
|
||||
def _tenant_from_dict(d: Dict[str, Any]) -> GpuTenant:
|
||||
return GpuTenant(
|
||||
name=d["name"],
|
||||
kind=d.get("kind", KIND_OTHER),
|
||||
enabled=d.get("enabled", True),
|
||||
priority=int(d.get("priority", 50)),
|
||||
needs_vram_gb=float(d.get("needs_vram_gb", 0.0)),
|
||||
overclock_profile=d.get("overclock_profile"),
|
||||
vram_floor_gb=float(d.get("vram_floor_gb", 0.0)),
|
||||
idle_release_after_s=float(d.get("idle_release_after_s", 30.0)),
|
||||
match=ProcessMatch(**(d.get("match") or {})),
|
||||
busy=BusyProbe(**(d.get("busy") or {})),
|
||||
release=ReleaseStrategy(**(d.get("release") or {})),
|
||||
events=EventSource(**(d.get("events") or {})),
|
||||
notes=d.get("notes", ""),
|
||||
)
|
||||
|
||||
|
||||
# Defaults reproduce today's behaviour exactly; they are data, not special cases.
|
||||
DEFAULT_TENANTS: List[Dict[str, Any]] = [
|
||||
{
|
||||
"name": "ollama",
|
||||
"kind": KIND_LLM,
|
||||
# Lower than ComfyUI on purpose: an interactive diffusion job preempts the LLM,
|
||||
# whose weights stay in the page cache and reload in seconds. Getting this the
|
||||
# wrong way round silently disabled the service's central behaviour -- ComfyUI
|
||||
# could never reclaim from Ollama.
|
||||
"priority": 50,
|
||||
"needs_vram_gb": 4.0,
|
||||
"idle_release_after_s": 0.0,
|
||||
"overclock_profile": "ollama",
|
||||
"match": {"names": ["ollama"], "cmdline": ["llama-server", "ollama"]},
|
||||
"busy": {"type": "http_count", "url": "http://localhost:11434/api/ps",
|
||||
"count_keys": ["models"]},
|
||||
"release": {"type": "http_post", "url": "http://localhost:11434/api/generate",
|
||||
"body": {"keep_alive": 0}, "per_model": True, "timeout_s": 120.0},
|
||||
"notes": "Unloads per model. With OLLAMA_NUM_PARALLEL=1 the request queues "
|
||||
"behind any running generation and applies when it finishes.",
|
||||
},
|
||||
{
|
||||
"name": "comfyui",
|
||||
"kind": KIND_DIFFUSION,
|
||||
"priority": 60,
|
||||
"needs_vram_gb": 6.0,
|
||||
"idle_release_after_s": 30.0,
|
||||
"overclock_profile": "comfy",
|
||||
"vram_floor_gb": 0.45,
|
||||
"match": {"cmdline": ["comfyui", "comfy"], "cmdline_endswith": ["main.py"]},
|
||||
"busy": {"type": "http_count", "url": "http://127.0.0.1:8188/queue",
|
||||
"count_keys": ["queue_running", "queue_pending"],
|
||||
"vram_floor_gb": 1.5, "stale_after_s": 90.0},
|
||||
"release": {"type": "http_post", "url": "http://127.0.0.1:8188/free",
|
||||
"body": {"unload_models": True, "free_memory": True},
|
||||
"timeout_s": 30.0},
|
||||
"events": {"type": "websocket", "url": "ws://127.0.0.1:8188/ws?clientId=hyperswap"},
|
||||
"notes": "Leaves dead jobs in queue_running; the queue flag is corroborated "
|
||||
"against its own VRAM before being believed.",
|
||||
},
|
||||
{
|
||||
"name": "desktop",
|
||||
"kind": KIND_DESKTOP,
|
||||
"priority": 90,
|
||||
"match": {"names": ["gnome-shell", "xorg", "mutter", "kwin", "plasmashell",
|
||||
"gnome-remote-desktop", "sddm", "gdm", "picom", "weston"]},
|
||||
"release": {"type": "none"},
|
||||
"notes": "Compositor and display server. Small, permanent, never reclaimable.",
|
||||
},
|
||||
]
|
||||
|
||||
|
||||
_cache: Dict[str, Any] = {"ts": 0.0, "tenants": None, "mtime": None}
|
||||
CACHE_TTL_S = 10.0
|
||||
|
||||
|
||||
def load_tenants(force: bool = False) -> List[GpuTenant]:
|
||||
"""Load tenant definitions, writing the defaults out on first run."""
|
||||
now = time.time()
|
||||
try:
|
||||
mtime = os.path.getmtime(CONFIG_PATH) if os.path.exists(CONFIG_PATH) else None
|
||||
except OSError:
|
||||
mtime = None
|
||||
if (not force and _cache["tenants"] is not None
|
||||
and mtime == _cache["mtime"] and (now - _cache["ts"]) < CACHE_TTL_S):
|
||||
return _cache["tenants"]
|
||||
|
||||
raw: List[Dict[str, Any]]
|
||||
if os.path.exists(CONFIG_PATH):
|
||||
try:
|
||||
with open(CONFIG_PATH) as f:
|
||||
raw = json.load(f)
|
||||
except Exception as e:
|
||||
logger.error(f"could not read {CONFIG_PATH}, using defaults: {e}")
|
||||
raw = DEFAULT_TENANTS
|
||||
else:
|
||||
raw = DEFAULT_TENANTS
|
||||
try:
|
||||
with open(CONFIG_PATH, "w") as f:
|
||||
json.dump(DEFAULT_TENANTS, f, indent=2)
|
||||
logger.info(f"wrote default tenant definitions to {CONFIG_PATH}")
|
||||
except Exception as e:
|
||||
logger.warning(f"could not write {CONFIG_PATH}: {e}")
|
||||
|
||||
# Merge in any fields a shipped default has gained since the config was written.
|
||||
# Without this, adding a field silently disables the behaviour it controls for every
|
||||
# existing install -- needs_vram_gb defaulted to 0, which made starvation
|
||||
# undetectable for the two tenants that had been written out before it existed.
|
||||
defaults_by_name = {d["name"]: d for d in DEFAULT_TENANTS}
|
||||
tenants = []
|
||||
for d in raw:
|
||||
base = defaults_by_name.get(d.get("name"))
|
||||
if base:
|
||||
merged = {**base, **d}
|
||||
for key in ("match", "busy", "release", "events"):
|
||||
if isinstance(base.get(key), dict):
|
||||
merged[key] = {**base[key], **(d.get(key) or {})}
|
||||
d = merged
|
||||
try:
|
||||
tenants.append(_tenant_from_dict(d))
|
||||
except Exception as e:
|
||||
logger.error(f"skipping malformed tenant {d!r}: {e}")
|
||||
_cache.update({"ts": now, "tenants": tenants, "mtime": mtime})
|
||||
return tenants
|
||||
|
||||
|
||||
def get_tenant(name: str) -> Optional[GpuTenant]:
|
||||
return next((t for t in load_tenants() if t.name == name), None)
|
||||
|
||||
|
||||
def save_tenants(tenants: List[Dict[str, Any]]) -> bool:
|
||||
try:
|
||||
with open(CONFIG_PATH, "w") as f:
|
||||
json.dump(tenants, f, indent=2)
|
||||
_cache["tenants"] = None
|
||||
return True
|
||||
except Exception as e:
|
||||
logger.error(f"save_tenants failed: {e}")
|
||||
return False
|
||||
|
||||
|
||||
def classify_process(pname: str, cmdline: str) -> str:
|
||||
"""Return the owning tenant's name, or 'unmanaged'.
|
||||
|
||||
'unmanaged' is meaningful rather than a dumping ground: it is VRAM this service has
|
||||
no way to reclaim, and it is reported as such.
|
||||
"""
|
||||
for t in load_tenants():
|
||||
if t.enabled and t.match.matches(pname, cmdline):
|
||||
return t.name
|
||||
return "unmanaged"
|
||||
|
||||
|
||||
def classify_pid(pid: int) -> str:
|
||||
try:
|
||||
proc = psutil.Process(pid)
|
||||
return classify_process(proc.name(), " ".join(proc.cmdline()))
|
||||
except Exception:
|
||||
return "unmanaged"
|
||||
|
||||
|
||||
async def probe_busy(tenant: GpuTenant, vram_gb: float = 0.0,
|
||||
state: Optional[Dict[str, Any]] = None) -> Dict[str, Any]:
|
||||
"""Is this tenant actually working? Returns {busy, reason, stale}."""
|
||||
probe = tenant.busy
|
||||
if probe.type == "vram":
|
||||
busy = vram_gb > probe.vram_busy_gb
|
||||
return {"busy": busy, "reason": f"{vram_gb:.2f} GB held", "stale": False}
|
||||
if probe.type != "http_count" or not probe.url:
|
||||
return {"busy": False, "reason": "no busy probe configured", "stale": False}
|
||||
|
||||
try:
|
||||
async with httpx.AsyncClient(timeout=3.0) as c:
|
||||
r = await c.get(probe.url)
|
||||
if r.status_code != 200:
|
||||
return {"busy": False, "reason": f"probe HTTP {r.status_code}", "stale": False}
|
||||
data = r.json()
|
||||
count = sum(len(data.get(k) or []) for k in probe.count_keys)
|
||||
except Exception as e:
|
||||
return {"busy": False, "reason": f"probe failed: {str(e)[:60]}", "stale": False}
|
||||
|
||||
if count == 0:
|
||||
return {"busy": False, "reason": "queue empty", "stale": False}
|
||||
# A queue that claims work while the tenant holds no VRAM is not doing work.
|
||||
if probe.vram_floor_gb and vram_gb < probe.vram_floor_gb:
|
||||
return {"busy": True, "reason": f"{count} queued, holding {vram_gb:.2f} GB",
|
||||
"stale": None, "below_floor": True}
|
||||
return {"busy": True, "reason": f"{count} queued/running", "stale": False}
|
||||
|
||||
|
||||
async def release_vram(tenant: GpuTenant, models: Optional[List[str]] = None
|
||||
) -> Dict[str, Any]:
|
||||
"""Ask a tenant to give its VRAM back, however that tenant expects to be asked."""
|
||||
strategy = tenant.release
|
||||
if strategy.type == "none" or not strategy.url:
|
||||
return {"success": False, "tenant": tenant.name, "released": False,
|
||||
"reason": "this tenant exposes no way to release VRAM"}
|
||||
|
||||
t0 = time.perf_counter()
|
||||
payloads: List[Dict[str, Any]] = []
|
||||
if strategy.per_model:
|
||||
for m in (models or []):
|
||||
payloads.append({**strategy.body, "model": m})
|
||||
if not payloads:
|
||||
return {"success": True, "tenant": tenant.name, "released": False,
|
||||
"reason": "nothing loaded to release"}
|
||||
else:
|
||||
payloads.append(dict(strategy.body))
|
||||
|
||||
errors = []
|
||||
try:
|
||||
async with httpx.AsyncClient(timeout=strategy.timeout_s) as c:
|
||||
for body in payloads:
|
||||
try:
|
||||
await c.post(strategy.url, json=body)
|
||||
except Exception as e:
|
||||
errors.append(str(e)[:80])
|
||||
except Exception as e:
|
||||
errors.append(str(e)[:80])
|
||||
|
||||
return {
|
||||
"success": not errors,
|
||||
"tenant": tenant.name,
|
||||
"released": True,
|
||||
"requests": len(payloads),
|
||||
"duration_ms": round((time.perf_counter() - t0) * 1000, 2),
|
||||
"errors": errors or None,
|
||||
}
|
||||
|
||||
|
||||
def describe() -> List[Dict[str, Any]]:
|
||||
"""Tenant definitions for the API, with what each can and cannot do."""
|
||||
out = []
|
||||
for t in sorted(load_tenants(), key=lambda x: -x.priority):
|
||||
d = t.to_dict()
|
||||
d["reclaimable"] = t.reclaimable
|
||||
d["busy_probe"] = t.busy.type
|
||||
d["release_via"] = t.release.type
|
||||
out.append(d)
|
||||
return out
|
||||
|
||||
|
||||
# ---------------------------------------------------------------- arbitration
|
||||
|
||||
def plan_release(demanding: str, tenants_state: List[Dict[str, Any]],
|
||||
free_gb: float, needed_gb: float) -> Dict[str, Any]:
|
||||
"""Decide who should give up VRAM so a starved tenant can work.
|
||||
|
||||
Generic over any number of applications: candidates are every *reclaimable* tenant
|
||||
that is not itself busy and ranks below the demanding one, taken lowest priority
|
||||
first, until enough would be freed. The two-application version of this was a pair
|
||||
of hardcoded rules -- yield Ollama for ComfyUI, purge ComfyUI for Ollama -- which
|
||||
could not express a third participant at all.
|
||||
|
||||
Returns the plan rather than performing it, so the decision is testable and can be
|
||||
logged before anything is actually released.
|
||||
"""
|
||||
by_name = {s["name"]: s for s in tenants_state}
|
||||
demander = by_name.get(demanding)
|
||||
if not demander:
|
||||
return {"possible": False, "reason": f"unknown tenant '{demanding}'", "release": []}
|
||||
|
||||
# The demander keeps what it already holds; only the remainder must be found.
|
||||
shortfall = needed_gb - free_gb - demander.get("vram_gb", 0.0)
|
||||
if shortfall <= 0:
|
||||
return {"possible": True, "reason": "enough VRAM is already free",
|
||||
"release": [], "shortfall_gb": 0.0}
|
||||
|
||||
# Who may be asked for memory:
|
||||
#
|
||||
# * any idle reclaimable tenant, whatever its rank -- idle memory is not in use;
|
||||
# * a *busy* tenant that ranks strictly below the demander.
|
||||
#
|
||||
# That second clause is the point of the whole service and was nearly lost. Refusing
|
||||
# to touch anything busy looks safe and is not: a diffusion job measured here ran for
|
||||
# 46 s instead of 3 s, squeezed into 1.6 GB, because the LLM reloaded straight after
|
||||
# yielding and was then protected as "busy" while ComfyUI starved. Preempting a
|
||||
# lower-priority tenant is safe precisely because releasing is asynchronous -- an
|
||||
# Ollama unload queues behind its running request and applies when that finishes, so
|
||||
# nothing is killed mid-flight.
|
||||
#
|
||||
# Equal or higher priority is never interrupted, so peers cannot fight.
|
||||
demander_priority = demander.get("priority", 0)
|
||||
candidates = [
|
||||
s for s in tenants_state
|
||||
if s["name"] != demanding
|
||||
and s.get("reclaimable")
|
||||
and s.get("vram_gb", 0) > 0
|
||||
and (not s.get("busy") or s.get("priority", 0) < demander_priority)
|
||||
]
|
||||
# Idle tenants first, then lowest priority: never disturb working software while
|
||||
# something idle still has memory to give.
|
||||
candidates.sort(key=lambda s: (bool(s.get("busy")), s.get("priority", 0),
|
||||
-s.get("vram_gb", 0)))
|
||||
|
||||
plan, freed = [], 0.0
|
||||
for c in candidates:
|
||||
if freed >= shortfall:
|
||||
break
|
||||
# Only what the tenant can actually give back, not everything it holds.
|
||||
releasable = max(c.get("vram_gb", 0.0) - c.get("vram_floor_gb", 0.0), 0.0)
|
||||
if releasable <= 0:
|
||||
continue
|
||||
plan.append(c["name"])
|
||||
freed += releasable
|
||||
|
||||
blockers = [
|
||||
{"name": s["name"], "vram_gb": s.get("vram_gb", 0.0),
|
||||
"why": ("busy and ranks at or above the demander" if s.get("busy") else
|
||||
"declares no release mechanism" if not s.get("reclaimable") else
|
||||
"enough was freed without it")}
|
||||
for s in tenants_state
|
||||
if s["name"] != demanding and s.get("vram_gb", 0) > 0 and s["name"] not in plan
|
||||
]
|
||||
|
||||
return {
|
||||
"possible": freed >= shortfall,
|
||||
"shortfall_gb": round(shortfall, 2),
|
||||
"would_free_gb": round(freed, 2),
|
||||
"release": plan,
|
||||
"blockers": blockers,
|
||||
"reason": (f"releasing {', '.join(plan)} frees {freed:.2f} GB of the "
|
||||
f"{shortfall:.2f} GB shortfall" if plan else
|
||||
"no reclaimable idle tenant holds enough VRAM"),
|
||||
}
|
||||
@@ -67,3 +67,24 @@ def temp_db(tmp_path, monkeypatch):
|
||||
telemetry_store.stop()
|
||||
except Exception:
|
||||
pass
|
||||
|
||||
|
||||
@pytest.fixture(autouse=True)
|
||||
def isolated_tenant_registry(tmp_path, monkeypatch):
|
||||
"""Never let tests read the operator's live tenants.json.
|
||||
|
||||
Classification is now configuration, which means a test that reads the real config
|
||||
changes result when someone adds an application to their own machine -- exactly what
|
||||
happened when stt-relay was registered and a "third party is unmanaged" test started
|
||||
seeing it as a named tenant. Every test gets the shipped defaults unless it opts out
|
||||
by pointing CONFIG_PATH somewhere itself.
|
||||
"""
|
||||
import json as _json
|
||||
import tenants as _tenants
|
||||
|
||||
path = tmp_path / "tenants-default.json"
|
||||
path.write_text(_json.dumps(_tenants.DEFAULT_TENANTS))
|
||||
monkeypatch.setattr(_tenants, "CONFIG_PATH", str(path))
|
||||
_tenants._cache.update({"ts": 0.0, "tenants": None, "mtime": None})
|
||||
yield
|
||||
_tenants._cache.update({"ts": 0.0, "tenants": None, "mtime": None})
|
||||
|
||||
133
tests/test_engines.py
Normal file
133
tests/test_engines.py
Normal file
@@ -0,0 +1,133 @@
|
||||
"""Tests for live engine-configuration reporting.
|
||||
|
||||
These settings live outside this codebase but dictate how arbitration must behave, and
|
||||
working out why a yield behaved a certain way once meant reading journald by hand. The
|
||||
dashboard previously asserted them as hardcoded text, which happened to be accurate --
|
||||
worse than being wrong, because it would have stayed accurate-looking after the settings
|
||||
changed.
|
||||
"""
|
||||
import asyncio
|
||||
|
||||
import pytest
|
||||
|
||||
import engines
|
||||
|
||||
|
||||
class _Proc:
|
||||
def __init__(self, stdout=""):
|
||||
self.stdout = stdout
|
||||
self.returncode = 0
|
||||
|
||||
|
||||
class TestOllamaEnvironmentParsing:
|
||||
def test_parses_the_real_unit_environment(self, monkeypatch):
|
||||
# Verbatim from `systemctl show ollama -p Environment --value` on this machine.
|
||||
raw = ("OLLAMA_HOST=0.0.0.0:11434 OLLAMA_FLASH_ATTENTION=1 "
|
||||
"OLLAMA_KV_CACHE_TYPE=q4_0 OLLAMA_KEEP_ALIVE=30m "
|
||||
"OLLAMA_MAX_LOADED_MODELS=1 OLLAMA_NUM_PARALLEL=1 OLLAMA_NUM_BATCH=2048")
|
||||
monkeypatch.setattr(engines.subprocess, "run", lambda *a, **k: _Proc(raw))
|
||||
env = engines._ollama_unit_environment()
|
||||
assert env["OLLAMA_NUM_PARALLEL"] == "1"
|
||||
assert env["OLLAMA_MAX_LOADED_MODELS"] == "1"
|
||||
assert env["OLLAMA_KV_CACHE_TYPE"] == "q4_0"
|
||||
|
||||
def test_ignores_non_ollama_variables(self, monkeypatch):
|
||||
monkeypatch.setattr(engines.subprocess, "run",
|
||||
lambda *a, **k: _Proc("PATH=/usr/bin OLLAMA_HOST=x:1 HOME=/root"))
|
||||
env = engines._ollama_unit_environment()
|
||||
assert set(env) == {"OLLAMA_HOST"}
|
||||
|
||||
def test_returns_empty_rather_than_raising_when_systemctl_fails(self, monkeypatch):
|
||||
def boom(*a, **k):
|
||||
raise FileNotFoundError("systemctl")
|
||||
monkeypatch.setattr(engines.subprocess, "run", boom)
|
||||
assert engines._ollama_unit_environment() == {}
|
||||
|
||||
|
||||
class TestEngineConfigReport:
|
||||
def _run(self, monkeypatch, env, comfy_ok=True):
|
||||
monkeypatch.setattr(engines, "_ollama_unit_environment", lambda: env)
|
||||
|
||||
class _Resp:
|
||||
status_code = 200 if comfy_ok else 500
|
||||
def json(self):
|
||||
return {"system": {"comfyui_version": "0.33.1",
|
||||
"pytorch_version": "2.11.0+cu128",
|
||||
"python_version": "3.14.4 (main)",
|
||||
"argv": ["main.py", "--listen", "0.0.0.0"]},
|
||||
"devices": [{"name": "cuda:0 NVIDIA GeForce RTX 4080 SUPER "
|
||||
": cudaMallocAsync"}]}
|
||||
|
||||
class _Client:
|
||||
async def __aenter__(self): return self
|
||||
async def __aexit__(self, *a): return False
|
||||
async def get(self, url): return _Resp()
|
||||
|
||||
monkeypatch.setattr(engines.httpx, "AsyncClient", lambda **k: _Client())
|
||||
return asyncio.run(engines.get_engine_config())
|
||||
|
||||
def test_surfaces_the_settings_that_drive_arbitration(self, monkeypatch):
|
||||
d = self._run(monkeypatch, {"OLLAMA_NUM_PARALLEL": "1",
|
||||
"OLLAMA_MAX_LOADED_MODELS": "1",
|
||||
"OLLAMA_KEEP_ALIVE": "30m"})
|
||||
assert d["ollama"]["num_parallel"] == "1"
|
||||
assert d["ollama"]["max_loaded_models"] == "1"
|
||||
assert d["ollama"]["keep_alive"] == "30m"
|
||||
|
||||
def test_num_parallel_explains_the_deferred_yield_behaviour(self, monkeypatch):
|
||||
d = self._run(monkeypatch, {"OLLAMA_NUM_PARALLEL": "1"})
|
||||
note = next(s["means"] for s in d["ollama"]["settings"]
|
||||
if s["key"] == "OLLAMA_NUM_PARALLEL")
|
||||
# The explanation is the point: it is why a busy model is deferred, not failed.
|
||||
assert "queue" in note.lower()
|
||||
|
||||
def test_summary_reflects_actual_flags_not_a_fixed_string(self, monkeypatch):
|
||||
on = self._run(monkeypatch, {"OLLAMA_FLASH_ATTENTION": "1",
|
||||
"OLLAMA_KV_CACHE_TYPE": "q4_0"})
|
||||
assert "FlashAttention" in on["ollama"]["summary"]
|
||||
assert "q4_0" in on["ollama"]["summary"]
|
||||
off = self._run(monkeypatch, {})
|
||||
assert "FlashAttention" not in off["ollama"]["summary"]
|
||||
assert off["ollama"]["config_source"] == "unavailable"
|
||||
|
||||
def test_port_comes_from_ollama_host(self, monkeypatch):
|
||||
d = self._run(monkeypatch, {"OLLAMA_HOST": "0.0.0.0:11500"})
|
||||
assert d["ollama"]["port"] == "11500"
|
||||
|
||||
def test_comfy_allocator_and_vram_mode_are_read_not_asserted(self, monkeypatch):
|
||||
d = self._run(monkeypatch, {})
|
||||
assert d["comfyui"]["allocator"] == "cudaMallocAsync"
|
||||
assert d["comfyui"]["vram_mode"] == "default (auto)"
|
||||
assert d["comfyui"]["version"] == "0.33.1"
|
||||
|
||||
def test_comfy_vram_flag_is_detected_when_present(self, monkeypatch):
|
||||
monkeypatch.setattr(engines, "_ollama_unit_environment", lambda: {})
|
||||
|
||||
class _Resp:
|
||||
status_code = 200
|
||||
def json(self):
|
||||
return {"system": {"argv": ["main.py", "--lowvram"]},
|
||||
"devices": [{"name": "cuda:0 X : cudaMalloc"}]}
|
||||
|
||||
class _Client:
|
||||
async def __aenter__(self): return self
|
||||
async def __aexit__(self, *a): return False
|
||||
async def get(self, url): return _Resp()
|
||||
|
||||
monkeypatch.setattr(engines.httpx, "AsyncClient", lambda **k: _Client())
|
||||
d = asyncio.run(engines.get_engine_config())
|
||||
assert d["comfyui"]["vram_mode"] == "--lowvram"
|
||||
assert d["comfyui"]["allocator"] == "cudaMalloc"
|
||||
|
||||
def test_offline_comfy_is_reported_not_raised(self, monkeypatch):
|
||||
monkeypatch.setattr(engines, "_ollama_unit_environment", lambda: {})
|
||||
|
||||
class _Client:
|
||||
async def __aenter__(self): return self
|
||||
async def __aexit__(self, *a): return False
|
||||
async def get(self, url): raise ConnectionError("refused")
|
||||
|
||||
monkeypatch.setattr(engines.httpx, "AsyncClient", lambda **k: _Client())
|
||||
d = asyncio.run(engines.get_engine_config())
|
||||
assert d["comfyui"]["online"] is False
|
||||
assert "error" in d["comfyui"]
|
||||
@@ -105,7 +105,8 @@ class TestAggregation:
|
||||
monkeypatch.setattr(health, "_check_nvml", lambda: checks[0])
|
||||
monkeypatch.setattr(health, "_check_sudo_smi", lambda: checks[1])
|
||||
for fn in ("_check_fan_control", "_check_profile_drift", "_check_store",
|
||||
"_check_residency", "_check_model_dirs", "_check_comfy_ws"):
|
||||
"_check_residency", "_check_model_dirs", "_check_comfy_ws",
|
||||
"_check_unmanaged_vram", "_check_comfy_queue"):
|
||||
monkeypatch.setattr(health, fn, lambda: health._check("x", health.OK, "d"))
|
||||
|
||||
async def fake_http(name, url, impact, fix):
|
||||
@@ -121,7 +122,7 @@ class TestAggregation:
|
||||
monkeypatch.setattr(health, "_check_nvml", boom)
|
||||
for fn in ("_check_sudo_smi", "_check_fan_control", "_check_profile_drift",
|
||||
"_check_store", "_check_residency", "_check_model_dirs",
|
||||
"_check_comfy_ws"):
|
||||
"_check_comfy_ws", "_check_unmanaged_vram", "_check_comfy_queue"):
|
||||
monkeypatch.setattr(health, fn, lambda: health._check("x", health.OK, "d"))
|
||||
|
||||
async def fake_http(name, url, impact, fix):
|
||||
@@ -132,3 +133,84 @@ class TestAggregation:
|
||||
# A broken check must surface as failed, not take down the endpoint.
|
||||
assert res["status"] == health.FAILED
|
||||
assert any("exploded" in c["detail"] for c in res["checks"])
|
||||
|
||||
|
||||
class TestUnmanagedVramCheck:
|
||||
"""Turning an unreclaimable-VRAM number into something actionable.
|
||||
|
||||
The arithmetic here has to be right or the check is worse than useless. A first
|
||||
version omitted ComfyUI's CUDA context -- which survives a purge -- and so reported
|
||||
a 14.93 GB model as fitting against a real ceiling of 14.60 GB. That was the very
|
||||
model the service had just refused with 507 Insufficient Storage.
|
||||
"""
|
||||
|
||||
def _gpu(self, unmanaged_gb=0.82, desktop_gb=0.01, comfy_gb=0.56, total=15.99,
|
||||
procs=None):
|
||||
return {
|
||||
"available": True,
|
||||
"vram_total_gb": total,
|
||||
"breakdown": {
|
||||
"unmanaged_gb": unmanaged_gb, "desktop_gb": desktop_gb,
|
||||
"comfyui_gb": comfy_gb,
|
||||
"unmanaged": procs if procs is not None else
|
||||
[{"pid": 1, "name": "python", "vram_mb": unmanaged_gb * 1024,
|
||||
"cmdline": "stt_relay.py"}],
|
||||
},
|
||||
}
|
||||
|
||||
def _blobs(self, sizes):
|
||||
return [{"model": f"m{i}", "size_gb": s} for i, s in enumerate(sizes)]
|
||||
|
||||
def test_ok_when_nothing_holds_unreclaimable_vram(self, monkeypatch):
|
||||
monkeypatch.setattr(health.vram_arbitrator, "get_gpu_hardware_stats",
|
||||
lambda: self._gpu(unmanaged_gb=0.0, procs=[]))
|
||||
assert health._check_unmanaged_vram()["status"] == health.OK
|
||||
|
||||
def test_comfy_cuda_context_counts_against_the_ceiling(self, monkeypatch):
|
||||
# 15.99 - 0.82 unmanaged - 0.01 desktop - 0.56 comfy floor = 14.60 GB available.
|
||||
# A 12.87 GB blob needs 12.87 * 1.16 = 14.93 GB, so it does not fit -- matching
|
||||
# the observed 507.
|
||||
monkeypatch.setattr(health.vram_arbitrator, "get_gpu_hardware_stats",
|
||||
lambda: self._gpu())
|
||||
monkeypatch.setattr(health, "_comfy_vram_floor_gb", lambda default=0, days=1: 0.56)
|
||||
monkeypatch.setattr(health.ram_optimizer, "find_ollama_model_files",
|
||||
lambda: self._blobs([12.87]))
|
||||
res = health._check_unmanaged_vram()
|
||||
assert res["status"] == health.DEGRADED
|
||||
assert "1 model(s)" in res["impact"]
|
||||
|
||||
def test_model_that_fits_even_without_the_unmanaged_process_is_not_flagged(self, monkeypatch):
|
||||
# A tiny model fits either way, so the unmanaged process is not what blocks it.
|
||||
monkeypatch.setattr(health.vram_arbitrator, "get_gpu_hardware_stats",
|
||||
lambda: self._gpu())
|
||||
monkeypatch.setattr(health, "_comfy_vram_floor_gb", lambda default=0, days=1: 0.56)
|
||||
monkeypatch.setattr(health.ram_optimizer, "find_ollama_model_files",
|
||||
lambda: self._blobs([2.0]))
|
||||
assert health._check_unmanaged_vram()["status"] == health.OK
|
||||
|
||||
def test_model_too_big_to_ever_fit_is_not_blamed_on_the_process(self, monkeypatch):
|
||||
# A 23.7 GB model does not fit on a 16 GB card regardless; saying the 842 MB
|
||||
# process is why would send the user after the wrong thing.
|
||||
monkeypatch.setattr(health.vram_arbitrator, "get_gpu_hardware_stats",
|
||||
lambda: self._gpu())
|
||||
monkeypatch.setattr(health, "_comfy_vram_floor_gb", lambda default=0, days=1: 0.56)
|
||||
monkeypatch.setattr(health.ram_optimizer, "find_ollama_model_files",
|
||||
lambda: self._blobs([23.7]))
|
||||
assert health._check_unmanaged_vram()["status"] == health.OK
|
||||
|
||||
def test_floor_uses_the_minimum_observed_not_the_current_value(self, monkeypatch):
|
||||
# Current VRAM could be a 7 GB checkpoint mid-generation; the floor is what
|
||||
# survives a purge.
|
||||
monkeypatch.setattr(health.telemetry_store, "_rows",
|
||||
lambda *a, **k: [{"floor": int(0.24 * 1024 ** 3)}])
|
||||
assert health._comfy_vram_floor_gb(default=7.0) == 0.24
|
||||
|
||||
def test_floor_falls_back_when_history_is_empty(self, monkeypatch):
|
||||
monkeypatch.setattr(health.telemetry_store, "_rows", lambda *a, **k: [])
|
||||
assert health._comfy_vram_floor_gb(default=0.56) == 0.56
|
||||
|
||||
def test_floor_falls_back_rather_than_raising(self, monkeypatch):
|
||||
def boom(*a, **k):
|
||||
raise RuntimeError("db gone")
|
||||
monkeypatch.setattr(health.telemetry_store, "_rows", boom)
|
||||
assert health._comfy_vram_floor_gb(default=0.5) == 0.5
|
||||
|
||||
184
tests/test_jobs.py
Normal file
184
tests/test_jobs.py
Normal file
@@ -0,0 +1,184 @@
|
||||
"""Tests for the cross-tenant job queue.
|
||||
|
||||
Every case below corresponds to something that actually went wrong while building this,
|
||||
because the failure modes are not obvious from the code:
|
||||
|
||||
* dispatching without room does not fail gracefully -- a CUDA OOM kills llama-server,
|
||||
and three queued jobs were destroyed in a row;
|
||||
* the VRAM requirement is a property of the job's model, not of the tenant, and a flat
|
||||
4 GB let a 14.9 GB model be dispatched into 8 GB of free memory;
|
||||
* waiting forever is as wrong as failing immediately, when the job can never fit;
|
||||
* an exception during dispatch left the row RUNNING while the scheduler moved on.
|
||||
"""
|
||||
import asyncio
|
||||
import json
|
||||
import time
|
||||
|
||||
import pytest
|
||||
|
||||
import jobs as J
|
||||
import telemetry_store
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def queue(tmp_path, monkeypatch):
|
||||
monkeypatch.setattr(telemetry_store, "DB_PATH", str(tmp_path / "q.db"))
|
||||
J.init()
|
||||
return tmp_path
|
||||
|
||||
|
||||
class TestQueueBasics:
|
||||
def test_submit_and_read_back(self, queue):
|
||||
r = J.submit("ollama", {"model": "m"}, label="first")
|
||||
assert r["success"]
|
||||
job = J.get(r["id"])
|
||||
assert job["state"] == J.PENDING and job["label"] == "first"
|
||||
assert job["payload"] == {"model": "m"}
|
||||
|
||||
def test_unknown_tenant_is_rejected(self, queue):
|
||||
assert J.submit("nope", {})["success"] is False
|
||||
|
||||
def test_priority_defaults_to_the_tenants_own(self, queue):
|
||||
import tenants as T
|
||||
r = J.submit("comfyui", {})
|
||||
assert r["priority"] == T.get_tenant("comfyui").priority
|
||||
|
||||
def test_pending_jobs_are_listed_in_execution_order(self, queue):
|
||||
J.submit("ollama", {}, priority=10, label="low")
|
||||
J.submit("ollama", {}, priority=90, label="high")
|
||||
J.submit("ollama", {}, priority=50, label="mid")
|
||||
order = [j["label"] for j in J.listing(J.PENDING)]
|
||||
assert order == ["high", "mid", "low"]
|
||||
|
||||
def test_equal_priority_is_first_in_first_out(self, queue):
|
||||
a = J.submit("ollama", {}, priority=50, label="a")["id"]
|
||||
time.sleep(0.01)
|
||||
J.submit("ollama", {}, priority=50, label="b")
|
||||
assert J._next_job()["id"] == a
|
||||
|
||||
def test_the_queue_has_no_depth_limit(self, queue):
|
||||
# "Unbounded" is the point; it lives on disk, not in memory.
|
||||
for i in range(500):
|
||||
J.submit("ollama", {}, label=f"j{i}")
|
||||
assert J.stats()["queue_depth"] == 500
|
||||
|
||||
def test_queue_survives_a_restart(self, queue):
|
||||
J.submit("ollama", {}, label="persisted")
|
||||
J._cache = None # nothing in-process is holding it
|
||||
assert [j["label"] for j in J.listing(J.PENDING)] == ["persisted"]
|
||||
|
||||
|
||||
class TestCancellation:
|
||||
def test_pending_jobs_can_be_cancelled(self, queue):
|
||||
jid = J.submit("ollama", {})["id"]
|
||||
assert J.cancel(jid)["success"] is True
|
||||
assert J.get(jid)["state"] == J.CANCELLED
|
||||
|
||||
def test_running_work_is_never_cancelled(self, queue):
|
||||
# This service frees VRAM by asking, never by killing work in flight.
|
||||
jid = J.submit("ollama", {})["id"]
|
||||
J._mark(jid, J.RUNNING, started_at=time.time())
|
||||
assert J.cancel(jid)["success"] is False
|
||||
assert J.get(jid)["state"] == J.RUNNING
|
||||
|
||||
def test_clearing_the_queue_leaves_running_work_alone(self, queue):
|
||||
running = J.submit("ollama", {})["id"]
|
||||
J._mark(running, J.RUNNING, started_at=time.time())
|
||||
J.submit("ollama", {})
|
||||
J.submit("ollama", {})
|
||||
assert J.clear_pending()["cancelled"] == 2
|
||||
assert J.get(running)["state"] == J.RUNNING
|
||||
|
||||
|
||||
class TestOrphanRecovery:
|
||||
def test_jobs_left_running_by_a_dead_process_are_requeued(self, queue):
|
||||
"""RUNNING means "this process is working on it".
|
||||
|
||||
An exception during dispatch left the row RUNNING while the scheduler moved on,
|
||||
so the job never finished and never retried.
|
||||
"""
|
||||
jid = J.submit("ollama", {})["id"]
|
||||
J._mark(jid, J.RUNNING, started_at=time.time())
|
||||
assert J.requeue_orphans() == 1
|
||||
job = J.get(jid)
|
||||
assert job["state"] == J.PENDING and job["started_at"] is None
|
||||
|
||||
def test_finished_jobs_are_untouched_by_recovery(self, queue):
|
||||
done = J.submit("ollama", {})["id"]
|
||||
J._mark(done, J.DONE, finished_at=time.time())
|
||||
assert J.requeue_orphans() == 0
|
||||
assert J.get(done)["state"] == J.DONE
|
||||
|
||||
|
||||
class TestPerJobVramRequirement:
|
||||
"""A tenant-wide figure cannot be right for an LLM."""
|
||||
|
||||
def test_llm_requirement_comes_from_the_model_being_loaded(self, queue, monkeypatch):
|
||||
import vram_arbitrator
|
||||
monkeypatch.setattr(vram_arbitrator, "_model_size_bytes",
|
||||
lambda m: int(12.87 * 1024 ** 3))
|
||||
s = J.Scheduler()
|
||||
# 12.87 GB on disk occupies ~14.9 GB once context and KV cache are allocated.
|
||||
assert 14.5 < s._job_vram_requirement("ollama", {"model": "big"}) < 15.5
|
||||
|
||||
def test_a_small_model_needs_correspondingly_less(self, queue, monkeypatch):
|
||||
import vram_arbitrator
|
||||
monkeypatch.setattr(vram_arbitrator, "_model_size_bytes",
|
||||
lambda m: int(1.96 * 1024 ** 3))
|
||||
s = J.Scheduler()
|
||||
assert s._job_vram_requirement("ollama", {"model": "small"}) < 3.0
|
||||
|
||||
def test_falls_back_to_the_tenant_figure_when_the_model_is_unknown(self, queue, monkeypatch):
|
||||
import vram_arbitrator
|
||||
monkeypatch.setattr(vram_arbitrator, "_model_size_bytes", lambda m: 0)
|
||||
s = J.Scheduler()
|
||||
import tenants as T
|
||||
assert (s._job_vram_requirement("ollama", {"model": "?"})
|
||||
== T.get_tenant("ollama").needs_vram_gb)
|
||||
|
||||
def test_non_llm_tenants_use_their_declared_figure(self, queue):
|
||||
import tenants as T
|
||||
s = J.Scheduler()
|
||||
assert (s._job_vram_requirement("comfyui", {})
|
||||
== T.get_tenant("comfyui").needs_vram_gb)
|
||||
|
||||
|
||||
class TestSchedulerStatus:
|
||||
def test_stats_report_depth_and_the_oldest_wait(self, queue):
|
||||
J.submit("ollama", {})
|
||||
J.submit("comfyui", {})
|
||||
st = J.stats()
|
||||
assert st["queue_depth"] == 2
|
||||
assert st["pending_by_tenant"] == {"ollama": 1, "comfyui": 1}
|
||||
assert st["oldest_pending_s"] is not None
|
||||
|
||||
def test_status_includes_queue_stats(self, queue):
|
||||
J.submit("ollama", {})
|
||||
s = J.Scheduler().get_status()
|
||||
assert s["queue_depth"] == 1 and s["running"] is False
|
||||
|
||||
def test_a_blocked_job_reports_why_it_is_waiting(self, queue):
|
||||
# Silence here is what left three jobs pending indefinitely with no explanation.
|
||||
s = J.Scheduler()
|
||||
s.blocked = {"id": "x", "tenant": "ollama", "reason": "needs 14.93 GB",
|
||||
"waited_s": 42.0, "since": time.time()}
|
||||
assert "14.93" in s.get_status()["blocked"]["reason"]
|
||||
|
||||
|
||||
class TestDispatchers:
|
||||
def test_every_tenant_kind_that_can_run_work_has_a_dispatcher(self):
|
||||
import tenants as T
|
||||
assert T.KIND_LLM in J.DISPATCHERS
|
||||
assert T.KIND_DIFFUSION in J.DISPATCHERS
|
||||
|
||||
def test_ollama_dispatch_reports_failure_rather_than_raising(self, queue, monkeypatch):
|
||||
class _R:
|
||||
status_code = 500
|
||||
text = '{"error":"llama-server process has terminated: cudaMalloc failed"}'
|
||||
class _C:
|
||||
async def __aenter__(self): return self
|
||||
async def __aexit__(self, *a): return False
|
||||
async def post(self, url, json=None): return _R()
|
||||
monkeypatch.setattr(J.httpx, "AsyncClient", lambda **k: _C())
|
||||
res = asyncio.run(J._dispatch_ollama({"model": "m", "prompt": "hi"}))
|
||||
assert res["ok"] is False and "cudaMalloc" in res["error"]
|
||||
405
tests/test_tenants.py
Normal file
405
tests/test_tenants.py
Normal file
@@ -0,0 +1,405 @@
|
||||
"""Tests for the GPU tenant registry.
|
||||
|
||||
The point of this service is fast handoff of one GPU between applications, and it should
|
||||
work for any application -- not only the two it grew up around. Their names had ended up
|
||||
compiled into process matching, VRAM attribution, busy detection and release calls alike.
|
||||
These tests pin the properties that make the registry generic: adding an application is
|
||||
configuration, and nothing in the arbitration logic knows a particular name.
|
||||
"""
|
||||
import asyncio
|
||||
import json
|
||||
|
||||
import pytest
|
||||
|
||||
import tenants as T
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def cfg(tmp_path, monkeypatch):
|
||||
path = tmp_path / "tenants.json"
|
||||
monkeypatch.setattr(T, "CONFIG_PATH", str(path))
|
||||
T._cache.update({"ts": 0.0, "tenants": None, "mtime": None})
|
||||
return path
|
||||
|
||||
|
||||
class TestProcessMatching:
|
||||
def test_matches_by_process_name(self):
|
||||
m = T.ProcessMatch(names=["ollama"])
|
||||
assert m.matches("ollama", "/usr/bin/ollama serve")
|
||||
assert not m.matches("python", "main.py")
|
||||
|
||||
def test_matches_by_cmdline_substring(self):
|
||||
m = T.ProcessMatch(cmdline=["llama-server"])
|
||||
assert m.matches("python", "/usr/local/lib/ollama/llama-server --model x")
|
||||
|
||||
def test_matches_by_cmdline_suffix(self):
|
||||
# ComfyUI is a bare `python main.py`, with nothing else distinguishing it.
|
||||
m = T.ProcessMatch(cmdline_endswith=["main.py"])
|
||||
assert m.matches("python", "/opt/ComfyUI/venv/bin/python main.py")
|
||||
assert not m.matches("python", "/opt/other/main.py --serve")
|
||||
|
||||
def test_matching_is_case_insensitive(self):
|
||||
assert T.ProcessMatch(names=["Xorg"]).matches("XORG", "")
|
||||
|
||||
|
||||
class TestDefaultsPreserveExistingBehaviour:
|
||||
"""The shipped defaults must classify exactly as the hardcoded version did."""
|
||||
|
||||
@pytest.mark.parametrize("pname,cmdline,expected", [
|
||||
("llama-server", "/usr/local/lib/ollama/llama-server --model x", "ollama"),
|
||||
("ollama", "/usr/bin/ollama serve", "ollama"),
|
||||
("python", "/home/u/ComfyUI/venv/bin/python main.py --listen", "comfyui"),
|
||||
("gnome-shell", "/usr/bin/gnome-shell --mode=ubuntu", "desktop"),
|
||||
("Xorg", "/usr/lib/xorg/Xorg :8", "desktop"),
|
||||
("python", "/home/u/robopest-venv/bin/python /home/u/stt_relay.py", "unmanaged"),
|
||||
("trainer", "/opt/ml/bin/trainer --epochs 3", "unmanaged"),
|
||||
])
|
||||
def test_classification(self, cfg, pname, cmdline, expected):
|
||||
assert T.classify_process(pname, cmdline) == expected
|
||||
|
||||
def test_unknown_process_is_unmanaged_not_silently_owned(self, cfg):
|
||||
# Misattributing a third party's VRAM to a tenant would make this service
|
||||
# promise headroom it cannot deliver.
|
||||
assert T.classify_process("weird", "/opt/x/weird --run") == "unmanaged"
|
||||
|
||||
|
||||
class TestAddingAnApplicationIsConfiguration:
|
||||
def test_a_new_tenant_is_recognised_without_code_changes(self, cfg):
|
||||
cfg.write_text(json.dumps(T.DEFAULT_TENANTS + [{
|
||||
"name": "trainer",
|
||||
"kind": "other",
|
||||
"priority": 80,
|
||||
"match": {"cmdline": ["train.py"]},
|
||||
"release": {"type": "http_post", "url": "http://localhost:9999/release"},
|
||||
}]))
|
||||
assert T.classify_process("python", "/opt/ml/train.py --epochs 3") == "trainer"
|
||||
t = T.get_tenant("trainer")
|
||||
assert t.priority == 80 and t.reclaimable
|
||||
|
||||
def test_first_run_writes_the_defaults(self, cfg):
|
||||
assert not cfg.exists()
|
||||
T.load_tenants(force=True)
|
||||
assert cfg.exists()
|
||||
assert {t["name"] for t in json.loads(cfg.read_text())} == {
|
||||
"ollama", "comfyui", "desktop"}
|
||||
|
||||
def test_a_malformed_entry_is_skipped_not_fatal(self, cfg):
|
||||
cfg.write_text(json.dumps([{"name": "ok", "match": {"names": ["a"]}},
|
||||
{"no_name": True}]))
|
||||
names = [t.name for t in T.load_tenants(force=True)]
|
||||
assert names == ["ok"]
|
||||
|
||||
def test_corrupt_config_falls_back_to_defaults(self, cfg):
|
||||
cfg.write_text("{ not json")
|
||||
assert {t.name for t in T.load_tenants(force=True)} >= {"ollama", "comfyui"}
|
||||
|
||||
|
||||
class TestReclaimability:
|
||||
def test_a_tenant_with_no_release_strategy_is_not_reclaimable(self, cfg):
|
||||
t = T.GpuTenant(name="x", release=T.ReleaseStrategy(type="none"))
|
||||
assert t.reclaimable is False
|
||||
|
||||
def test_release_refuses_rather_than_reporting_success(self, cfg):
|
||||
t = T.GpuTenant(name="x", release=T.ReleaseStrategy(type="none"))
|
||||
res = asyncio.run(T.release_vram(t))
|
||||
assert res["success"] is False and res["released"] is False
|
||||
assert "no way to release" in res["reason"]
|
||||
|
||||
def test_per_model_release_with_nothing_loaded_is_a_no_op(self, cfg):
|
||||
t = T.GpuTenant(name="ollama", release=T.ReleaseStrategy(
|
||||
type="http_post", url="http://x/api", per_model=True))
|
||||
res = asyncio.run(T.release_vram(t, models=[]))
|
||||
assert res["success"] is True and res["released"] is False
|
||||
|
||||
|
||||
class TestBusyProbe:
|
||||
def _probe(self, monkeypatch, payload, status=200):
|
||||
class _R:
|
||||
status_code = status
|
||||
def json(self_inner): return payload
|
||||
class _C:
|
||||
async def __aenter__(self): return self
|
||||
async def __aexit__(self, *a): return False
|
||||
async def get(self, url): return _R()
|
||||
monkeypatch.setattr(T.httpx, "AsyncClient", lambda **k: _C())
|
||||
|
||||
def test_empty_queue_is_not_busy(self, monkeypatch):
|
||||
self._probe(monkeypatch, {"queue_running": [], "queue_pending": []})
|
||||
t = T.GpuTenant(name="c", busy=T.BusyProbe(
|
||||
type="http_count", url="http://x/queue",
|
||||
count_keys=["queue_running", "queue_pending"]))
|
||||
assert asyncio.run(T.probe_busy(t))["busy"] is False
|
||||
|
||||
def test_queued_work_while_holding_no_vram_is_flagged_below_floor(self, monkeypatch):
|
||||
# ComfyUI leaves dead jobs in queue_running; only its VRAM reveals that nothing
|
||||
# is loaded.
|
||||
self._probe(monkeypatch, {"queue_running": [[1, "abc"]], "queue_pending": []})
|
||||
t = T.GpuTenant(name="c", busy=T.BusyProbe(
|
||||
type="http_count", url="http://x/queue", count_keys=["queue_running"],
|
||||
vram_floor_gb=1.5))
|
||||
res = asyncio.run(T.probe_busy(t, vram_gb=0.56))
|
||||
assert res["busy"] is True and res.get("below_floor") is True
|
||||
|
||||
def test_queued_work_with_a_checkpoint_loaded_is_plainly_busy(self, monkeypatch):
|
||||
self._probe(monkeypatch, {"queue_running": [[1, "abc"]], "queue_pending": []})
|
||||
t = T.GpuTenant(name="c", busy=T.BusyProbe(
|
||||
type="http_count", url="http://x/queue", count_keys=["queue_running"],
|
||||
vram_floor_gb=1.5))
|
||||
res = asyncio.run(T.probe_busy(t, vram_gb=6.8))
|
||||
assert res["busy"] is True and not res.get("below_floor")
|
||||
|
||||
def test_vram_probe_needs_no_http_endpoint(self):
|
||||
# An application with no API can still be observed by what it holds.
|
||||
t = T.GpuTenant(name="x", busy=T.BusyProbe(type="vram", vram_busy_gb=1.0))
|
||||
assert asyncio.run(T.probe_busy(t, vram_gb=2.0))["busy"] is True
|
||||
assert asyncio.run(T.probe_busy(t, vram_gb=0.5))["busy"] is False
|
||||
|
||||
def test_an_unreachable_probe_reports_not_busy_rather_than_raising(self, monkeypatch):
|
||||
class _C:
|
||||
async def __aenter__(self): return self
|
||||
async def __aexit__(self, *a): return False
|
||||
async def get(self, url): raise ConnectionError("refused")
|
||||
monkeypatch.setattr(T.httpx, "AsyncClient", lambda **k: _C())
|
||||
t = T.GpuTenant(name="c", busy=T.BusyProbe(type="http_count", url="http://x",
|
||||
count_keys=["q"]))
|
||||
res = asyncio.run(T.probe_busy(t))
|
||||
assert res["busy"] is False and "failed" in res["reason"]
|
||||
|
||||
|
||||
class TestPriority:
|
||||
def test_describe_orders_by_priority(self, cfg):
|
||||
rows = T.describe()
|
||||
prios = [r["priority"] for r in rows]
|
||||
assert prios == sorted(prios, reverse=True)
|
||||
assert all("reclaimable" in r for r in rows)
|
||||
|
||||
|
||||
class TestReleasePlanning:
|
||||
"""Deciding who gives up VRAM, generically over any number of applications.
|
||||
|
||||
The two-application version was a pair of hardcoded rules -- yield Ollama when
|
||||
ComfyUI is busy, purge ComfyUI when Ollama is starved -- which could not express a
|
||||
third participant at all.
|
||||
"""
|
||||
|
||||
def _state(self, **overrides):
|
||||
base = [
|
||||
{"name": "desktop", "priority": 90, "vram_gb": 0.01, "busy": False,
|
||||
"reclaimable": False},
|
||||
{"name": "stt-relay", "priority": 70, "vram_gb": 0.8, "busy": False,
|
||||
"reclaimable": False},
|
||||
# Shipped priorities: diffusion outranks the LLM, whose weights reload
|
||||
# from page cache in seconds.
|
||||
{"name": "comfyui", "priority": 60, "vram_gb": 7.0, "busy": False,
|
||||
"reclaimable": True},
|
||||
{"name": "ollama", "priority": 50, "vram_gb": 0.0, "busy": True,
|
||||
"reclaimable": True},
|
||||
]
|
||||
for s in base:
|
||||
s.update(overrides.get(s["name"], {}))
|
||||
return base
|
||||
|
||||
def test_a_tenant_already_holding_what_it_needs_is_not_starved(self):
|
||||
# A busy GPU has little free by definition. Comparing free VRAM alone flagged a
|
||||
# tenant working fine on 13 GB as demanding, which would have caused pointless
|
||||
# releases from everything else.
|
||||
state = self._state(ollama={"vram_gb": 13.0})
|
||||
plan = T.plan_release("ollama", state, free_gb=1.5, needed_gb=4.0)
|
||||
assert plan["release"] == []
|
||||
assert "already free" in plan["reason"]
|
||||
|
||||
def test_starved_tenant_reclaims_from_the_idle_one_below_it(self):
|
||||
plan = T.plan_release("ollama", self._state(), free_gb=1.5, needed_gb=14.9)
|
||||
assert plan["release"] == ["comfyui"]
|
||||
|
||||
def test_a_busy_tenant_ranking_above_the_demander_is_not_a_victim(self):
|
||||
# comfyui outranks ollama, so ollama may not interrupt it.
|
||||
state = self._state(comfyui={"busy": True})
|
||||
plan = T.plan_release("ollama", state, free_gb=1.5, needed_gb=14.9)
|
||||
assert plan["release"] == []
|
||||
assert any(b["name"] == "comfyui" and "busy" in b["why"]
|
||||
for b in plan["blockers"])
|
||||
|
||||
def test_a_higher_priority_demander_preempts_busy_lower_priority_work(self):
|
||||
"""The measured regression that made this rule necessary.
|
||||
|
||||
Refusing to touch anything busy looks safe and is not. With the LLM protected as
|
||||
"busy", a diffusion job ran 46 s instead of 3 s, squeezed into 1.6 GB, because
|
||||
the LLM reloaded immediately after yielding and was then untouchable. Preempting
|
||||
a lower-priority tenant is safe because releasing is asynchronous: an Ollama
|
||||
unload queues behind its running request rather than killing it.
|
||||
"""
|
||||
state = [
|
||||
{"name": "comfyui", "priority": 60, "vram_gb": 1.65, "busy": True,
|
||||
"reclaimable": True},
|
||||
{"name": "ollama", "priority": 50, "vram_gb": 13.03, "busy": True,
|
||||
"reclaimable": True},
|
||||
]
|
||||
plan = T.plan_release("comfyui", state, free_gb=0.28, needed_gb=6.0)
|
||||
assert plan["release"] == ["ollama"]
|
||||
|
||||
def test_an_idle_tenant_is_preferred_over_preempting_a_busy_one(self):
|
||||
state = [
|
||||
{"name": "d", "priority": 60, "vram_gb": 0.0, "busy": True,
|
||||
"reclaimable": True},
|
||||
{"name": "busy_low", "priority": 10, "vram_gb": 8.0, "busy": True,
|
||||
"reclaimable": True},
|
||||
{"name": "idle_high", "priority": 90, "vram_gb": 8.0, "busy": False,
|
||||
"reclaimable": True},
|
||||
]
|
||||
plan = T.plan_release("d", state, free_gb=0.0, needed_gb=8.0)
|
||||
assert plan["release"] == ["idle_high"]
|
||||
|
||||
def test_unreclaimable_tenants_are_named_as_blockers_not_ignored(self):
|
||||
# The user needs to know a third-party process is what stands in the way.
|
||||
plan = T.plan_release("ollama", self._state(), free_gb=0.0, needed_gb=15.5)
|
||||
blockers = {b["name"]: b["why"] for b in plan["blockers"]}
|
||||
assert blockers["stt-relay"] == "declares no release mechanism"
|
||||
assert plan["possible"] is False
|
||||
|
||||
def test_an_idle_tenant_yields_even_if_it_outranks_the_demander(self):
|
||||
"""Priority orders who is asked first; it does not protect idle memory.
|
||||
|
||||
Filtering candidates by priority broke both directions in turn: with the LLM
|
||||
ranked above diffusion, ComfyUI could never preempt Ollama -- the service's
|
||||
central behaviour -- and once the ranks were swapped, a starved Ollama could no
|
||||
longer reclaim from an idle ComfyUI. An idle tenant is not using its VRAM, so
|
||||
outranking the demander is not a reason to keep it.
|
||||
"""
|
||||
state = self._state(comfyui={"priority": 99, "busy": False})
|
||||
plan = T.plan_release("ollama", state, free_gb=1.0, needed_gb=14.9)
|
||||
assert plan["release"] == ["comfyui"]
|
||||
|
||||
def test_priority_decides_who_is_asked_first(self):
|
||||
state = [
|
||||
{"name": "demander", "priority": 50, "vram_gb": 0.0, "busy": True,
|
||||
"reclaimable": True},
|
||||
{"name": "high", "priority": 90, "vram_gb": 4.0, "busy": False,
|
||||
"reclaimable": True},
|
||||
{"name": "low", "priority": 10, "vram_gb": 4.0, "busy": False,
|
||||
"reclaimable": True},
|
||||
]
|
||||
plan = T.plan_release("demander", state, free_gb=0.0, needed_gb=5.0)
|
||||
# The lowest-priority idle tenant gives up memory first.
|
||||
assert plan["release"][0] == "low"
|
||||
|
||||
def test_peers_cannot_interrupt_each_other(self):
|
||||
# Equal priority is never preempted, so two tenants at the same rank cannot
|
||||
# fight over the card.
|
||||
state = [
|
||||
{"name": "a", "priority": 50, "vram_gb": 0.0, "busy": True,
|
||||
"reclaimable": True},
|
||||
{"name": "b", "priority": 50, "vram_gb": 8.0, "busy": True,
|
||||
"reclaimable": True},
|
||||
]
|
||||
plan = T.plan_release("a", state, free_gb=0.0, needed_gb=8.0)
|
||||
assert plan["release"] == []
|
||||
assert "busy" in plan["blockers"][0]["why"]
|
||||
|
||||
def test_lowest_priority_is_released_first(self):
|
||||
state = self._state() + [
|
||||
{"name": "batch", "priority": 10, "vram_gb": 3.0, "busy": False,
|
||||
"reclaimable": True}]
|
||||
plan = T.plan_release("ollama", state, free_gb=0.0, needed_gb=5.0)
|
||||
assert plan["release"][0] == "batch"
|
||||
|
||||
def test_releases_only_as_many_tenants_as_needed(self):
|
||||
state = self._state() + [
|
||||
{"name": "batch", "priority": 10, "vram_gb": 9.0, "busy": False,
|
||||
"reclaimable": True}]
|
||||
plan = T.plan_release("ollama", state, free_gb=0.0, needed_gb=8.0)
|
||||
assert plan["release"] == ["batch"] # 9 GB covers it; comfyui is left alone
|
||||
|
||||
def test_three_applications_can_all_participate(self):
|
||||
# The property the hardcoded pair of rules could not express.
|
||||
state = [
|
||||
{"name": "llm", "priority": 60, "vram_gb": 0.0, "busy": True,
|
||||
"reclaimable": True},
|
||||
{"name": "diffusion", "priority": 50, "vram_gb": 4.0, "busy": False,
|
||||
"reclaimable": True},
|
||||
{"name": "trainer", "priority": 40, "vram_gb": 5.0, "busy": False,
|
||||
"reclaimable": True},
|
||||
]
|
||||
plan = T.plan_release("llm", state, free_gb=0.0, needed_gb=9.0)
|
||||
assert set(plan["release"]) == {"trainer", "diffusion"}
|
||||
assert plan["possible"] is True
|
||||
|
||||
def test_unknown_tenant_is_rejected_cleanly(self):
|
||||
plan = T.plan_release("nope", self._state(), free_gb=0.0, needed_gb=1.0)
|
||||
assert plan["possible"] is False and plan["release"] == []
|
||||
|
||||
|
||||
class TestConfigUpgrade:
|
||||
def test_fields_added_later_are_merged_into_an_existing_config(self, cfg):
|
||||
# A config written before needs_vram_gb existed must not silently lose the
|
||||
# behaviour that field controls.
|
||||
cfg.write_text(json.dumps([{
|
||||
"name": "ollama",
|
||||
"match": {"names": ["ollama"]},
|
||||
}]))
|
||||
t = T.get_tenant("ollama")
|
||||
assert t.needs_vram_gb > 0
|
||||
assert t.release.type == "http_post"
|
||||
|
||||
def test_explicit_user_values_still_win_over_defaults(self, cfg):
|
||||
cfg.write_text(json.dumps([{
|
||||
"name": "ollama", "priority": 5, "needs_vram_gb": 99.0,
|
||||
"match": {"names": ["ollama"]},
|
||||
}]))
|
||||
t = T.get_tenant("ollama")
|
||||
assert t.priority == 5 and t.needs_vram_gb == 99.0
|
||||
|
||||
|
||||
class TestVramFloor:
|
||||
"""VRAM that survives a release must not be promised to anyone else.
|
||||
|
||||
ComfyUI keeps its CUDA context for as long as the process lives, so a purge does not
|
||||
return everything it holds. Ignoring that made plan_release report it would free
|
||||
0.37 GB against a 0.33 GB shortfall; the job was cleared to run and the memory never
|
||||
arrived, so it waited two minutes and then failed.
|
||||
"""
|
||||
|
||||
def test_only_memory_above_the_floor_counts_as_freeable(self):
|
||||
state = [
|
||||
{"name": "llm", "priority": 50, "vram_gb": 0.0, "busy": True,
|
||||
"reclaimable": True, "vram_floor_gb": 0.0},
|
||||
{"name": "diffusion", "priority": 60, "vram_gb": 0.44, "busy": False,
|
||||
"reclaimable": True, "vram_floor_gb": 0.45},
|
||||
]
|
||||
plan = T.plan_release("llm", state, free_gb=14.6, needed_gb=14.93)
|
||||
assert plan["possible"] is False
|
||||
assert plan["release"] == []
|
||||
|
||||
def test_a_loaded_checkpoint_is_still_freeable_above_its_floor(self):
|
||||
state = [
|
||||
{"name": "llm", "priority": 50, "vram_gb": 0.0, "busy": True,
|
||||
"reclaimable": True, "vram_floor_gb": 0.0},
|
||||
{"name": "diffusion", "priority": 60, "vram_gb": 7.0, "busy": False,
|
||||
"reclaimable": True, "vram_floor_gb": 0.45},
|
||||
]
|
||||
plan = T.plan_release("llm", state, free_gb=7.9, needed_gb=14.0)
|
||||
assert plan["release"] == ["diffusion"]
|
||||
# 7.0 held minus a 0.45 floor.
|
||||
assert abs(plan["would_free_gb"] - 6.55) < 0.01
|
||||
|
||||
def test_a_tenant_at_its_floor_is_not_even_listed_for_release(self):
|
||||
state = [
|
||||
{"name": "llm", "priority": 50, "vram_gb": 0.0, "busy": True,
|
||||
"reclaimable": True, "vram_floor_gb": 0.0},
|
||||
{"name": "at_floor", "priority": 10, "vram_gb": 0.3, "busy": False,
|
||||
"reclaimable": True, "vram_floor_gb": 0.45},
|
||||
{"name": "has_room", "priority": 20, "vram_gb": 5.0, "busy": False,
|
||||
"reclaimable": True, "vram_floor_gb": 0.0},
|
||||
]
|
||||
plan = T.plan_release("llm", state, free_gb=0.0, needed_gb=4.0)
|
||||
assert plan["release"] == ["has_room"]
|
||||
|
||||
def test_default_floor_is_zero_so_existing_configs_are_unchanged(self):
|
||||
state = [
|
||||
{"name": "a", "priority": 50, "vram_gb": 0.0, "busy": True,
|
||||
"reclaimable": True},
|
||||
{"name": "b", "priority": 40, "vram_gb": 5.0, "busy": False,
|
||||
"reclaimable": True},
|
||||
]
|
||||
plan = T.plan_release("a", state, free_gb=0.0, needed_gb=5.0)
|
||||
assert plan["release"] == ["b"] and plan["possible"] is True
|
||||
@@ -13,6 +13,7 @@ were observed on real hardware before being encoded here:
|
||||
out of memory ... unable to allocate CUDA0 buffer"
|
||||
"""
|
||||
import asyncio
|
||||
import time
|
||||
|
||||
import pytest
|
||||
|
||||
@@ -132,3 +133,70 @@ class TestBusyBackoff:
|
||||
assert "yield_deferred_busy" in arb.stats
|
||||
assert "yield_stalled" in arb.stats
|
||||
assert "yield_timeouts" not in arb.stats
|
||||
|
||||
|
||||
class TestComfyStaleQueueDetection:
|
||||
"""ComfyUI can leave a dead job in queue_running forever.
|
||||
|
||||
Observed on this machine: a WAN 2.1 i2v entry sat in queue_running while the GPU was
|
||||
idle and ComfyUI held 0.56 GB. Trusting that flag made the watchdog believe ComfyUI
|
||||
was permanently busy, so it evicted the LLM on every poll, never ran the idle purge,
|
||||
and never checked whether the LLM had been pushed onto the CPU. Instrumenting the
|
||||
watchdog showed busy=6, idle_check=0 -- one stale row had disabled half the logic.
|
||||
"""
|
||||
|
||||
def _arb(self, comfy_bytes):
|
||||
arb = v.AutoArbitrator()
|
||||
v.get_process_vram_bytes = lambda: {
|
||||
"ollama_bytes": 0, "comfyui_bytes": int(comfy_bytes), "other_bytes": 0,
|
||||
"desktop_bytes": 0, "unmanaged_bytes": 0, "free_bytes": 0, "gpu_util_pct": 0}
|
||||
return arb
|
||||
|
||||
def teardown_method(self):
|
||||
import importlib
|
||||
importlib.reload(v)
|
||||
|
||||
def test_empty_queue_is_not_busy(self):
|
||||
arb = self._arb(0)
|
||||
assert arb._comfy_genuinely_busy({"queue_running": [], "queue_pending": []}) is False
|
||||
|
||||
def test_pending_work_is_always_busy(self):
|
||||
arb = self._arb(0)
|
||||
assert arb._comfy_genuinely_busy(
|
||||
{"queue_running": [], "queue_pending": [[1, "p"]]}) is True
|
||||
|
||||
def test_a_running_job_is_believed_at_first(self):
|
||||
# It must not be called stale before it has had time to load anything.
|
||||
arb = self._arb(0.1 * GB)
|
||||
assert arb._comfy_genuinely_busy(
|
||||
{"queue_running": [[1, "abc"]], "queue_pending": []}) is True
|
||||
|
||||
def test_long_running_job_holding_no_vram_is_stale(self):
|
||||
arb = self._arb(0.56 * GB) # the observed CUDA-context floor
|
||||
q = {"queue_running": [[1, "abc"]], "queue_pending": []}
|
||||
arb._comfy_genuinely_busy(q)
|
||||
arb._running_since = time.time() - (arb.STALE_RUNNING_S + 5)
|
||||
assert arb._comfy_genuinely_busy(q) is False
|
||||
assert arb.comfy_stale_job == "abc"
|
||||
|
||||
def test_long_running_job_holding_a_checkpoint_is_real(self):
|
||||
# 6.8 GB is a loaded SDXL checkpoint; slow is not the same as stuck.
|
||||
arb = self._arb(6.8 * GB)
|
||||
q = {"queue_running": [[1, "abc"]], "queue_pending": []}
|
||||
arb._comfy_genuinely_busy(q)
|
||||
arb._running_since = time.time() - (arb.STALE_RUNNING_S + 5)
|
||||
assert arb._comfy_genuinely_busy(q) is True
|
||||
assert arb.comfy_stale_job is None
|
||||
|
||||
def test_a_new_prompt_id_resets_the_staleness_clock(self):
|
||||
arb = self._arb(0.5 * GB)
|
||||
arb._comfy_genuinely_busy({"queue_running": [[1, "old"]], "queue_pending": []})
|
||||
arb._running_since = time.time() - 1000
|
||||
assert arb._comfy_genuinely_busy(
|
||||
{"queue_running": [[1, "new"]], "queue_pending": []}) is True
|
||||
|
||||
def test_vram_not_utilisation_is_the_signal(self):
|
||||
# Utilisation is shared with Ollama and any third-party process, so it stayed
|
||||
# above every sensible threshold and a stuck entry never looked stale.
|
||||
assert hasattr(v.AutoArbitrator, "STALE_COMFY_BYTES")
|
||||
assert not hasattr(v.AutoArbitrator, "STALE_UTIL_PCT")
|
||||
|
||||
338
verify_arbitration.py
Executable file
338
verify_arbitration.py
Executable file
@@ -0,0 +1,338 @@
|
||||
#!/usr/bin/env python
|
||||
"""End-to-end verification of the arbitration cycle, against real hardware.
|
||||
|
||||
The unit suite covers logic in isolation; this exercises the promise the whole service
|
||||
exists to make -- that an LLM and a diffusion pipeline can share one 16 GB card without
|
||||
either failing -- and reports what actually happened at each stage.
|
||||
|
||||
It is deliberately not part of `pytest tests/`: it loads real models, runs a real
|
||||
diffusion graph and moves real VRAM, taking a few minutes. Run it when you want proof
|
||||
the system works on this machine:
|
||||
|
||||
python verify_arbitration.py # full cycle
|
||||
python verify_arbitration.py --quick # skip the diffusion stages
|
||||
|
||||
Every stage restores what it changed, and the script refuses to start if ComfyUI is
|
||||
already busy.
|
||||
"""
|
||||
import argparse
|
||||
import asyncio
|
||||
import sys
|
||||
import time
|
||||
from typing import Any, Dict, List, Optional
|
||||
|
||||
import httpx
|
||||
|
||||
BASE = "http://localhost:9090"
|
||||
|
||||
PASS, FAIL, SKIP, WARN = "PASS", "FAIL", "SKIP", "WARN"
|
||||
_results: List[Dict[str, Any]] = []
|
||||
|
||||
|
||||
def record(stage: str, status: str, detail: str, evidence: str = "") -> None:
|
||||
_results.append({"stage": stage, "status": status, "detail": detail,
|
||||
"evidence": evidence})
|
||||
colour = {"PASS": "\033[32m", "FAIL": "\033[31m",
|
||||
"SKIP": "\033[33m", "WARN": "\033[33m"}[status]
|
||||
print(f" {colour}{status:<4}\033[0m {stage:<38} {detail}")
|
||||
if evidence:
|
||||
print(f" {evidence}")
|
||||
|
||||
|
||||
async def api(client: httpx.AsyncClient, method: str, path: str,
|
||||
allow_error: bool = False, **kw) -> Any:
|
||||
r = await client.request(method, f"{BASE}{path}", **kw)
|
||||
if allow_error:
|
||||
# Some stages deliberately provoke a failure and need to read it.
|
||||
body = r.json() if r.headers.get("content-type", "").startswith("application/json") else {}
|
||||
return {"_status": r.status_code, **(body if isinstance(body, dict) else {})}
|
||||
r.raise_for_status()
|
||||
return r.json()
|
||||
|
||||
|
||||
async def stage_preflight(c: httpx.AsyncClient) -> Optional[Dict[str, Any]]:
|
||||
health = await api(c, "GET", "/api/health")
|
||||
failed = health.get("failed") or []
|
||||
if failed:
|
||||
record("preflight: dependencies", FAIL,
|
||||
f"{len(failed)} dependency check(s) failing", ", ".join(failed))
|
||||
return None
|
||||
record("preflight: dependencies", PASS, health["summary"])
|
||||
|
||||
stats = await api(c, "GET", "/api/stats")
|
||||
if stats["comfyui"].get("executing") or stats["comfyui"].get("queue_remaining"):
|
||||
record("preflight: ComfyUI idle", FAIL, "ComfyUI is busy; refusing to interfere")
|
||||
return None
|
||||
record("preflight: ComfyUI idle", PASS, "queue empty")
|
||||
return stats
|
||||
|
||||
|
||||
async def stage_llm_load(c: httpx.AsyncClient, model: str) -> bool:
|
||||
res = await api(c, "POST", "/api/switch-model",
|
||||
json={"model": model, "keep_alive": "5m"}, timeout=300)
|
||||
if not res.get("success"):
|
||||
record("LLM loads", FAIL, res.get("error", "")[:90])
|
||||
return False
|
||||
gbps, status = res.get("load_gbps"), res.get("cache_status")
|
||||
record("LLM loads", PASS, f"{res['model_size_gb']} GB in {res['load_duration_ms']:.0f} ms",
|
||||
f"{gbps} GB/s -> {status}; {res['tokens_per_sec']} tok/s")
|
||||
|
||||
# The classification must follow from the measured bandwidth, not a fixed duration.
|
||||
if gbps is not None:
|
||||
expected = ("RAM Cache Hit" if gbps >= 2.0
|
||||
else "Partial Cache" if gbps >= 0.8 else "Cold Disk Load")
|
||||
ok = expected.split()[0] in (status or "")
|
||||
record("load classified by bandwidth", PASS if ok else FAIL,
|
||||
f"{gbps} GB/s reported as '{status}'",
|
||||
"" if ok else f"expected something matching '{expected}'")
|
||||
return True
|
||||
|
||||
|
||||
async def stage_yield(c: httpx.AsyncClient) -> bool:
|
||||
before = await api(c, "GET", "/api/gpu")
|
||||
held = before["breakdown"]["ollama_gb"]
|
||||
res = await api(c, "POST", "/api/free-vram", timeout=60)
|
||||
outcome = res.get("outcome")
|
||||
|
||||
if outcome == "released":
|
||||
after = await api(c, "GET", "/api/gpu")
|
||||
freed = held - after["breakdown"]["ollama_gb"]
|
||||
# The barrier's promise: on return, the VRAM is genuinely gone.
|
||||
ok = after["breakdown"]["ollama_gb"] < 0.3
|
||||
record("VRAM yield is confirmed", PASS if ok else FAIL,
|
||||
f"released in {res['confirm_ms']} ms",
|
||||
f"{freed:.2f} GB freed; NVML now reports "
|
||||
f"{after['breakdown']['ollama_gb']} GB held by Ollama")
|
||||
return ok
|
||||
if outcome == "busy":
|
||||
record("VRAM yield is confirmed", WARN,
|
||||
"model was mid-generation, so the unload was deferred",
|
||||
res.get("error", ""))
|
||||
return True
|
||||
record("VRAM yield is confirmed", FAIL, f"outcome={outcome}", res.get("error", ""))
|
||||
return False
|
||||
|
||||
|
||||
async def stage_diffusion(c: httpx.AsyncClient) -> bool:
|
||||
sys.path.insert(0, "/home/drjones/unified-model-manager")
|
||||
import autotune, vram_arbitrator # noqa: E402 (imported late; needs the service's deps)
|
||||
|
||||
# The first run loads the checkpoint from disk. Timing that and calling the result
|
||||
# "it/s" understates throughput by roughly 10x -- 0.67 it/s against a steady-state
|
||||
# 6.7 -- so the load is measured separately and reported as what it is.
|
||||
first = await autotune._diffusion_benchmark()
|
||||
if not first.get("ok"):
|
||||
record("diffusion runs", FAIL, first.get("error", "")[:90])
|
||||
return False
|
||||
record("diffusion runs (cold, includes checkpoint load)", PASS,
|
||||
f"{first['exec_ms']} ms", f"{first['it_per_sec']} it/s including load")
|
||||
|
||||
res = await autotune._diffusion_benchmark()
|
||||
if not res.get("ok"):
|
||||
record("diffusion runs (warm)", FAIL, res.get("error", "")[:90])
|
||||
return False
|
||||
record("diffusion throughput (warm)", PASS,
|
||||
f"SDXL 1024/20 steps in {res['exec_ms']} ms", f"{res['it_per_sec']} it/s")
|
||||
|
||||
snap = vram_arbitrator.get_process_vram_bytes()
|
||||
comfy_gb = snap["comfyui_bytes"] / (1024 ** 3)
|
||||
record("ComfyUI holds its checkpoint", PASS if comfy_gb > 0.5 else WARN,
|
||||
f"{comfy_gb:.2f} GB retained",
|
||||
"held for the idle window rather than purged between iterations")
|
||||
return True
|
||||
|
||||
|
||||
async def stage_idle_purge(c: httpx.AsyncClient) -> bool:
|
||||
# The completion event arrives over the ComfyUI websocket, so the flag is set a
|
||||
# moment after the graph returns. Checking instantly raced it.
|
||||
arb = {}
|
||||
for _ in range(12):
|
||||
arb = (await api(c, "GET", "/api/stats"))["arbitrator"]
|
||||
if arb.get("pending_purge"):
|
||||
break
|
||||
await asyncio.sleep(0.5)
|
||||
if not arb.get("pending_purge"):
|
||||
record("purge is deferred, not immediate", WARN,
|
||||
"no purge pending after 6 s (ComfyUI may already be clean)")
|
||||
return True
|
||||
idle_s = arb.get("comfy_idle_s")
|
||||
record("purge is deferred, not immediate", PASS,
|
||||
f"holding checkpoints for {arb.get('idle_purge_after_s')} s",
|
||||
f"idle {idle_s} s so far" if idle_s is not None
|
||||
else "idle timer just started")
|
||||
return True
|
||||
|
||||
|
||||
async def stage_reclaim(c: httpx.AsyncClient, model: str) -> bool:
|
||||
"""The direction that used to fail outright: an LLM that will not fit.
|
||||
|
||||
This only proves anything if the chosen model genuinely cannot fit in what ComfyUI
|
||||
has left free. A small model fits alongside the checkpoint and the stage passes
|
||||
without exercising the reclaim path at all, so pick the largest model that will not
|
||||
fit and say plainly when no such model exists.
|
||||
"""
|
||||
gpu = await api(c, "GET", "/api/gpu")
|
||||
comfy_gb = gpu["breakdown"]["comfyui_gb"]
|
||||
free_gb = gpu["vram_free_gb"]
|
||||
if comfy_gb < 0.5:
|
||||
record("reclaims VRAM for the LLM", SKIP,
|
||||
f"ComfyUI only holds {comfy_gb} GB; nothing to reclaim")
|
||||
return True
|
||||
|
||||
models = (await api(c, "GET", "/api/models"))["ollama_models"]
|
||||
EMBED = {"bert", "nomic-bert", "gte", "jina-bert"}
|
||||
usable = [m for m in models
|
||||
if (m.get("details", {}).get("family") or "").lower() not in EMBED
|
||||
and "embed" not in m["name"].lower()]
|
||||
# On-disk weight size is not the VRAM footprint: measured on this box, a 12.87 GB
|
||||
# blob occupies 14.9 GB once context and KV cache are allocated. Sizing the test off
|
||||
# disk size picks a model that cannot fit even after a successful reclaim.
|
||||
VRAM_OVERHEAD = 1.18
|
||||
HEADROOM_GB = 0.4
|
||||
|
||||
def vram_need(m):
|
||||
return m.get("size", 0) / (1024 ** 3) * VRAM_OVERHEAD
|
||||
|
||||
# Unmanaged VRAM never comes back, so it is not part of what a reclaim can offer.
|
||||
# Ignoring it picked a model that failed even after a correct reclaim -- on this box
|
||||
# an 842 MB third-party process is the difference between a 14.9 GB model fitting
|
||||
# and not.
|
||||
reclaimable_gb = free_gb + comfy_gb - HEADROOM_GB
|
||||
too_big = [m for m in usable
|
||||
if vram_need(m) > free_gb and vram_need(m) < reclaimable_gb]
|
||||
if too_big:
|
||||
target = max(too_big, key=lambda m: m.get("size", 0))
|
||||
model = target["name"]
|
||||
print(f" using {model} ({target['size'] / (1024**3):.1f} GB on disk, "
|
||||
f"~{vram_need(target):.1f} GB in VRAM) — will not fit in "
|
||||
f"{free_gb:.1f} GB free, should fit after reclaiming {comfy_gb:.1f} GB")
|
||||
else:
|
||||
unmanaged = gpu["breakdown"].get("unmanaged_gb", 0)
|
||||
record("reclaims VRAM for the LLM", SKIP,
|
||||
f"no installed model needs between {free_gb:.1f} and "
|
||||
f"{reclaimable_gb:.1f} GB of VRAM",
|
||||
f"reclaimable ceiling excludes {unmanaged} GB held by processes "
|
||||
f"HyperSwap cannot free")
|
||||
return True
|
||||
|
||||
# Re-run a graph first. The idle purge fires 30 s after ComfyUI goes quiet, and a
|
||||
# large model takes longer than that to load -- so without resetting the timer the
|
||||
# purge frees ComfyUI mid-load and the reclaim path is never reached.
|
||||
sys.path.insert(0, "/home/drjones/unified-model-manager")
|
||||
import autotune # noqa: E402
|
||||
await autotune._diffusion_benchmark()
|
||||
gpu = await api(c, "GET", "/api/gpu")
|
||||
print(f" reset the idle window; ComfyUI holds "
|
||||
f"{gpu['breakdown']['comfyui_gb']} GB, {gpu['vram_free_gb']} GB free")
|
||||
|
||||
res = await api(c, "POST", "/api/switch-model", allow_error=True,
|
||||
json={"model": model, "keep_alive": "2m"}, timeout=600)
|
||||
if res.get("_status") == 507:
|
||||
record("reclaims VRAM for the LLM", FAIL,
|
||||
"reclaim ran but the model still did not fit",
|
||||
str(res.get("detail", ""))[:150])
|
||||
return False
|
||||
if res.get("_status", 200) >= 400 or not res.get("success"):
|
||||
record("reclaims VRAM for the LLM", FAIL,
|
||||
f"HTTP {res.get('_status')}", str(res.get("detail", ""))[:120])
|
||||
return False
|
||||
if res.get("_status") and res.get("_status") != 200:
|
||||
pass
|
||||
if res.get("reclaimed_from_comfyui_gb"):
|
||||
record("reclaims VRAM for the LLM", PASS,
|
||||
f"reclaimed {res['reclaimed_from_comfyui_gb']} GB and retried",
|
||||
f"'{model}' then loaded at {res.get('load_gbps')} GB/s")
|
||||
else:
|
||||
# It fit anyway, so nothing was proven; do not report that as a pass. The usual
|
||||
# cause is the idle purge firing during the load and freeing ComfyUI first.
|
||||
record("reclaims VRAM for the LLM", WARN,
|
||||
"model fit without a reclaim, so the path was not exercised",
|
||||
"the idle purge most likely freed ComfyUI during the load; "
|
||||
f"loaded at {res.get('load_gbps')} GB/s")
|
||||
return True
|
||||
|
||||
|
||||
async def stage_accounting(c: httpx.AsyncClient) -> bool:
|
||||
"""Reported VRAM must add up, and reported settings must match the hardware."""
|
||||
gpu = await api(c, "GET", "/api/gpu")
|
||||
bd = gpu["breakdown"]
|
||||
parts = bd["ollama_gb"] + bd["comfyui_gb"] + bd["system_gb"]
|
||||
used = gpu["vram_used_gb"]
|
||||
# Driver overhead means the parts never sum exactly; a large gap means mis-attribution.
|
||||
ok = abs(parts - used) < 1.5
|
||||
record("VRAM attribution adds up", PASS if ok else FAIL,
|
||||
f"parts {parts:.2f} GB vs NVML used {used:.2f} GB",
|
||||
f"ollama {bd['ollama_gb']} + comfy {bd['comfyui_gb']} + system {bd['system_gb']} "
|
||||
f"(desktop {bd.get('desktop_gb')}, unmanaged {bd.get('unmanaged_gb')})")
|
||||
|
||||
oc = await api(c, "GET", "/api/overclock")
|
||||
drift = oc["drift"]
|
||||
ok2 = not drift["drifted"]
|
||||
record("reported GPU state matches hardware", PASS if ok2 else FAIL,
|
||||
f"profile '{drift['profile']}' asks {drift['power_limit_intended_w']} W, "
|
||||
f"card reports {drift['power_limit_actual_w']} W",
|
||||
drift.get("reason") or "")
|
||||
return ok and ok2
|
||||
|
||||
|
||||
async def main() -> int:
|
||||
ap = argparse.ArgumentParser(description=__doc__,
|
||||
formatter_class=argparse.RawDescriptionHelpFormatter)
|
||||
ap.add_argument("--quick", action="store_true",
|
||||
help="skip the diffusion and reclaim stages")
|
||||
ap.add_argument("--model", default=None,
|
||||
help="Ollama model to test with (default: smallest installed)")
|
||||
args = ap.parse_args()
|
||||
|
||||
print("\nHyperSwap arbitration verification\n" + "=" * 62)
|
||||
async with httpx.AsyncClient(timeout=60.0) as c:
|
||||
baseline = await stage_preflight(c)
|
||||
if baseline is None:
|
||||
print("\nPreflight failed; not continuing.\n")
|
||||
return 2
|
||||
|
||||
model = args.model
|
||||
if not model:
|
||||
models = (await api(c, "GET", "/api/models"))["ollama_models"]
|
||||
# Embedding models have no /api/generate endpoint, and the smallest model
|
||||
# installed is often one of them.
|
||||
EMBED_FAMILIES = {"bert", "nomic-bert", "gte", "jina-bert"}
|
||||
usable = [m for m in models
|
||||
if (m.get("details", {}).get("family") or "").lower()
|
||||
not in EMBED_FAMILIES and "embed" not in m["name"].lower()]
|
||||
if not usable:
|
||||
record("choose a test model", FAIL,
|
||||
"no generative Ollama models installed "
|
||||
f"({len(models)} found, all embedding-only)")
|
||||
return 2
|
||||
model = min(usable, key=lambda m: m.get("size", 0))["name"]
|
||||
print(f"\n using model: {model}\n")
|
||||
|
||||
await stage_llm_load(c, model)
|
||||
await stage_yield(c)
|
||||
|
||||
if not args.quick:
|
||||
if await stage_diffusion(c):
|
||||
await stage_idle_purge(c)
|
||||
await stage_reclaim(c, model)
|
||||
else:
|
||||
record("diffusion stages", SKIP, "--quick")
|
||||
|
||||
await stage_accounting(c)
|
||||
|
||||
# Leave the box as we found it.
|
||||
await api(c, "POST", "/api/free-vram", timeout=60)
|
||||
|
||||
failed = [r for r in _results if r["status"] == FAIL]
|
||||
warned = [r for r in _results if r["status"] == WARN]
|
||||
print("=" * 62)
|
||||
print(f" {len(_results) - len(failed) - len(warned)} passed, "
|
||||
f"{len(warned)} warned, {len(failed)} failed")
|
||||
for r in failed:
|
||||
print(f" FAILED: {r['stage']} — {r['detail']}")
|
||||
print()
|
||||
return 1 if failed else 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(asyncio.run(main()))
|
||||
@@ -12,6 +12,7 @@ import websockets
|
||||
|
||||
import overclock_manager
|
||||
import ram_optimizer
|
||||
import tenants as tenants_mod
|
||||
import telemetry_store
|
||||
|
||||
try:
|
||||
@@ -72,7 +73,8 @@ RECLAIM_MIN_COMFY_BYTES = 512 * 1024 ** 2
|
||||
# depends on configuration: with n_gpu_layers left to Ollama it spills layers to the CPU
|
||||
# and reports size_vram < size; with n_gpu_layers pinned (99 on this box) it refuses and
|
||||
# returns a hard CUDA OOM instead. Both are handled -- the spill by
|
||||
# AutoArbitrator._check_ollama_starved, the hard failure by the retry below.
|
||||
# AutoArbitrator._arbitrate (generically, from the tenant registry), the hard failure by
|
||||
# the retry below.
|
||||
OOM_SIGNATURES = ("out of memory", "cudamalloc", "unable to allocate",
|
||||
"failed to allocate", "cuda error")
|
||||
|
||||
@@ -149,7 +151,8 @@ def get_process_vram_bytes() -> Dict[str, int]:
|
||||
to actually drain.
|
||||
"""
|
||||
out = {"ollama_bytes": 0, "comfyui_bytes": 0, "other_bytes": 0, "free_bytes": 0,
|
||||
"desktop_bytes": 0, "unmanaged_bytes": 0, "gpu_util_pct": 0}
|
||||
"desktop_bytes": 0, "unmanaged_bytes": 0, "gpu_util_pct": 0,
|
||||
"by_tenant_bytes": {}}
|
||||
if not NVML_AVAILABLE:
|
||||
return out
|
||||
try:
|
||||
@@ -176,6 +179,7 @@ def get_process_vram_bytes() -> Dict[str, int]:
|
||||
if len(_PID_KIND_CACHE) >= _PID_KIND_CACHE_MAX:
|
||||
_PID_KIND_CACHE.clear()
|
||||
_PID_KIND_CACHE[key] = kind
|
||||
out["by_tenant_bytes"][kind] = out["by_tenant_bytes"].get(kind, 0) + used
|
||||
if kind == "ollama":
|
||||
out["ollama_bytes"] += used
|
||||
elif kind == "comfy":
|
||||
@@ -198,6 +202,19 @@ _PID_KIND_CACHE: Dict[tuple, str] = {}
|
||||
_PID_KIND_CACHE_MAX = 512
|
||||
|
||||
|
||||
def _pid_key(pid: int) -> Optional[tuple]:
|
||||
try:
|
||||
return (pid, psutil.Process(pid).create_time())
|
||||
except Exception:
|
||||
return None
|
||||
|
||||
|
||||
# Tenant names as used by this module's buckets. The tenant registry is the source of
|
||||
# truth for *which* application a process belongs to; these two names are kept because
|
||||
# the REST payloads and the dashboard have used them since the beginning.
|
||||
_BUCKET_ALIASES = {"comfyui": "comfy"}
|
||||
|
||||
|
||||
def _pid_key(pid: int) -> Optional[tuple]:
|
||||
try:
|
||||
return (pid, psutil.Process(pid).create_time())
|
||||
@@ -214,26 +231,16 @@ DESKTOP_PROCESS_HINTS = (
|
||||
|
||||
|
||||
def _classify_pid(pid: int) -> str:
|
||||
"""Bucket a GPU process into ollama | comfy | desktop | unmanaged.
|
||||
"""Which tenant owns this GPU process.
|
||||
|
||||
The old version had one catch-all "other" bucket, which put a 3.9 MB compositor and
|
||||
an 842 MB long-running inference script in the same number. That matters: this
|
||||
service can reclaim VRAM from ComfyUI, but it cannot touch a third-party workload,
|
||||
and pretending otherwise makes it promise headroom it cannot deliver.
|
||||
The matching rules used to be substrings compiled into this function, which made the
|
||||
two applications on this box part of the arbitrator rather than input to it. They now
|
||||
come from the tenant registry, so a third application is a config entry.
|
||||
|
||||
"unmanaged" still means something specific and useful: VRAM held by something with no
|
||||
declared way to release it, and therefore headroom this service can never offer.
|
||||
"""
|
||||
try:
|
||||
proc = psutil.Process(pid)
|
||||
pname = proc.name().lower()
|
||||
cmdline = " ".join(proc.cmdline()).lower()
|
||||
except Exception:
|
||||
return "unmanaged"
|
||||
if "ollama" in pname or "llama-server" in cmdline:
|
||||
return "ollama"
|
||||
if "comfyui" in cmdline or "comfy" in cmdline or cmdline.rstrip().endswith("main.py"):
|
||||
return "comfy"
|
||||
if any(hint in pname or hint in cmdline for hint in DESKTOP_PROCESS_HINTS):
|
||||
return "desktop"
|
||||
return "unmanaged"
|
||||
return _BUCKET_ALIASES.get(tenants_mod.classify_pid(pid), tenants_mod.classify_pid(pid))
|
||||
|
||||
|
||||
def get_gpu_hardware_stats() -> Dict[str, Any]:
|
||||
@@ -336,7 +343,10 @@ def get_gpu_hardware_stats() -> Dict[str, Any]:
|
||||
"desktop_bytes": 0,
|
||||
"unmanaged_bytes": 0,
|
||||
"unmanaged": [],
|
||||
"processes": []
|
||||
"processes": [],
|
||||
# Generic attribution: one entry per tenant, so an application added to the
|
||||
# registry is reported without any change here.
|
||||
"by_tenant": {},
|
||||
}
|
||||
|
||||
try:
|
||||
@@ -376,6 +386,8 @@ def get_gpu_hardware_stats() -> Dict[str, Any]:
|
||||
"vram_mb": round(used_mem / (1024**2), 1),
|
||||
})
|
||||
|
||||
proc_breakdown["by_tenant"][kind] = (
|
||||
proc_breakdown["by_tenant"].get(kind, 0) + used_mem)
|
||||
proc_breakdown["processes"].append({
|
||||
"pid": pid,
|
||||
"name": pname,
|
||||
@@ -431,6 +443,8 @@ def get_gpu_hardware_stats() -> Dict[str, Any]:
|
||||
# reclaimed, so it is permanently unavailable headroom.
|
||||
"unmanaged_gb": round(proc_breakdown["unmanaged_bytes"] / (1024**3), 2),
|
||||
"unmanaged": proc_breakdown["unmanaged"],
|
||||
"by_tenant_gb": {k: round(b / (1024**3), 2)
|
||||
for k, b in proc_breakdown["by_tenant"].items()},
|
||||
"free_mb": round(free_vram / (1024**2), 1),
|
||||
"free_gb": round(free_vram / (1024**3), 2),
|
||||
"processes": proc_breakdown["processes"],
|
||||
@@ -849,37 +863,46 @@ async def switch_ollama_model(target_model: str, keep_alive: str = "30m",
|
||||
# idle ComfyUI and try once more.
|
||||
body = resp.text
|
||||
if looks_like_vram_oom(body) and not _retrying:
|
||||
snap = get_process_vram_bytes()
|
||||
if snap["comfyui_bytes"] >= RECLAIM_MIN_COMFY_BYTES:
|
||||
# Which application should give up memory is a question for the registry,
|
||||
# not something to answer by purging ComfyUI by name. Any reclaimable idle
|
||||
# tenant below Ollama in priority is a candidate.
|
||||
state = await arbitrator._tenant_state()
|
||||
free_gb = arbitrator._last_tenant_state["free_gb"]
|
||||
size_gb = _model_size_bytes(target_model) / (1024**3)
|
||||
needed = size_gb * 1.16 if size_gb else free_gb + 1.0
|
||||
plan = tenants_mod.plan_release("ollama", state, free_gb, needed)
|
||||
if plan["release"]:
|
||||
logger.warning(
|
||||
f"Ollama could not fit '{target_model}' with ComfyUI holding "
|
||||
f"{round(snap['comfyui_bytes'] / (1024**3), 2)} GB — reclaiming and retrying")
|
||||
purge = await instant_free_comfyui_vram()
|
||||
f"Ollama could not fit '{target_model}' — {plan['reason']}")
|
||||
freed_before = free_gb
|
||||
for victim in plan["release"]:
|
||||
await arbitrator._release_tenant(
|
||||
victim, f"Ollama could not load '{target_model}'")
|
||||
arbitrator.stats["reclaims_for_ollama"] += 1
|
||||
arbitrator.last_action = (
|
||||
f"Reclaimed {round(snap['comfyui_bytes'] / (1024**3), 2)}GB from ComfyUI so "
|
||||
f"'{target_model}' could load")
|
||||
f"Released {', '.join(plan['release'])} so '{target_model}' could load")
|
||||
_record({
|
||||
"event_type": "VRAM Reclaim for Ollama",
|
||||
"source": "ComfyUI Pipeline",
|
||||
"source": ", ".join(plan["release"]),
|
||||
"target": target_model,
|
||||
"duration_ms": purge.get("duration_ms"),
|
||||
"cache_status": "Reclaimed",
|
||||
"detail": f"Ollama OOM: {body[:160]}",
|
||||
})
|
||||
await asyncio.sleep(0.3)
|
||||
retry = await switch_ollama_model(target_model, keep_alive, _retrying=True)
|
||||
retry["reclaimed_from_comfyui_gb"] = round(
|
||||
snap["comfyui_bytes"] / (1024**3), 2)
|
||||
retry["released_tenants"] = plan["release"]
|
||||
retry["would_free_gb"] = plan.get("would_free_gb")
|
||||
retry["first_attempt_error"] = "CUDA OOM; retried after reclaiming VRAM"
|
||||
if not retry.get("success"):
|
||||
# Be specific about why the reclaim was not enough. Blaming ComfyUI
|
||||
# when a third-party process is holding the memory sends the user
|
||||
# looking in the wrong place.
|
||||
# Be specific about why the reclaim was not enough. Blaming a tenant
|
||||
# when a process nobody can release is holding the memory sends the
|
||||
# user looking in the wrong place.
|
||||
retry["blockers"] = plan.get("blockers")
|
||||
retry["unmanaged_blockers"] = describe_unmanaged()
|
||||
return retry
|
||||
return {"success": False, "error": f"HTTP {resp.status_code}: {body}",
|
||||
"duration_ms": total_duration_ms,
|
||||
"upstream_status": resp.status_code,
|
||||
"vram_oom": looks_like_vram_oom(body)}
|
||||
except Exception as e:
|
||||
return {"success": False, "error": str(e),
|
||||
@@ -911,6 +934,7 @@ class AutoArbitrator:
|
||||
def __init__(self):
|
||||
self.running = False
|
||||
self.ws_task: Optional[asyncio.Task] = None
|
||||
self.event_tasks: List[asyncio.Task] = []
|
||||
self.poll_task: Optional[asyncio.Task] = None
|
||||
self.idle_task: Optional[asyncio.Task] = None
|
||||
self.last_yield_time = 0.0
|
||||
@@ -930,6 +954,19 @@ class AutoArbitrator:
|
||||
self._yield_backoff_until: Dict[str, float] = {}
|
||||
self._yield_busy_streak: Dict[str, int] = {}
|
||||
self.last_reclaim_time = 0.0
|
||||
self.watchdog_branches = {"busy": 0, "completed": 0, "idle_check": 0,
|
||||
"bad_status": 0, "error": 0}
|
||||
self._running_id: Optional[str] = None
|
||||
self._running_since: Optional[float] = None
|
||||
self._peak_comfy_bytes = 0
|
||||
self.comfy_stale_job: Optional[str] = None
|
||||
self._idle_since: Dict[str, float] = {}
|
||||
self._last_event_wake = 0.0
|
||||
self.event_sources: Dict[str, str] = {}
|
||||
self._last_tenant_state: Optional[Dict[str, Any]] = None
|
||||
self.last_arbitration: Optional[Dict[str, Any]] = None
|
||||
self.last_handoff: Optional[Dict[str, Any]] = None
|
||||
self.last_watchdog_error: Optional[str] = None
|
||||
self.stats = {
|
||||
"yields": 0, # release confirmed
|
||||
"yield_deferred_busy": 0, # model mid-generation; unload queued behind it
|
||||
@@ -945,6 +982,12 @@ class AutoArbitrator:
|
||||
return
|
||||
self.running = True
|
||||
self.ws_task = asyncio.create_task(self._ws_listener())
|
||||
for t in tenants_mod.load_tenants():
|
||||
if t.enabled and t.events.type == "websocket" and t.events.url:
|
||||
self.event_tasks.append(asyncio.create_task(
|
||||
self._event_listener(t.name, t.events.url,
|
||||
t.events.reconnect_backoff_s,
|
||||
t.events.max_backoff_s)))
|
||||
self.poll_task = asyncio.create_task(self._poll_watchdog())
|
||||
self.idle_task = asyncio.create_task(self._idle_purge_loop())
|
||||
logger.info("AutoArbitrator background engine started (Bidirectional).")
|
||||
@@ -957,9 +1000,10 @@ class AutoArbitrator:
|
||||
|
||||
async def stop(self):
|
||||
self.running = False
|
||||
for task in (self.ws_task, self.poll_task, self.idle_task):
|
||||
for task in [self.ws_task, self.poll_task, self.idle_task, *self.event_tasks]:
|
||||
if task:
|
||||
task.cancel()
|
||||
self.event_tasks.clear()
|
||||
await close_clients()
|
||||
logger.info("AutoArbitrator background engine stopped.")
|
||||
|
||||
@@ -1069,6 +1113,51 @@ class AutoArbitrator:
|
||||
return {"purged": True, "free_gb": round(snap["free_bytes"] / (1024**3), 2)}
|
||||
return {"purged": False, "free_gb": round(free_gb, 2), "reason": "ComfyUI holds no VRAM"}
|
||||
|
||||
async def _event_listener(self, tenant_name: str, url: str,
|
||||
backoff_s: float, max_backoff_s: float) -> None:
|
||||
"""Wake on a tenant's event stream instead of waiting for the next poll.
|
||||
|
||||
Deliberately does not parse the messages. The previous listener understood
|
||||
ComfyUI's schema -- status/execution_start/executing/execution_success -- which
|
||||
tied the fast path to one application. Treating any message as "look now" and
|
||||
letting the tenant's own busy probe decide gives the same sub-second reaction
|
||||
for any application that emits anything on state change.
|
||||
"""
|
||||
backoff = backoff_s
|
||||
while self.running:
|
||||
try:
|
||||
async with websockets.connect(url, ping_interval=10, ping_timeout=10) as ws:
|
||||
self.event_sources[tenant_name] = "connected"
|
||||
self.connected_ws = True
|
||||
backoff = backoff_s
|
||||
logger.info(f"Event source connected for '{tenant_name}': {url}")
|
||||
while self.running:
|
||||
await ws.recv()
|
||||
# Coalesce bursts: a single graph emits many messages, and one
|
||||
# arbitration pass per burst is enough.
|
||||
now = time.time()
|
||||
if now - self._last_event_wake < 0.05:
|
||||
continue
|
||||
self._last_event_wake = now
|
||||
self.stats["event_wakeups"] = self.stats.get("event_wakeups", 0) + 1
|
||||
try:
|
||||
# A message on this tenant's own stream is live proof it is
|
||||
# working right now, so it is taken as busy rather than asked
|
||||
# over HTTP. A stale queue row could lie; an event arriving
|
||||
# this instant cannot.
|
||||
await self._arbitrate(active_tenant=tenant_name)
|
||||
except Exception as e:
|
||||
logger.debug(f"arbitration from event failed: {e}")
|
||||
except (websockets.exceptions.ConnectionClosed, OSError, asyncio.CancelledError):
|
||||
self.event_sources[tenant_name] = "disconnected"
|
||||
self.connected_ws = False
|
||||
except Exception as e:
|
||||
self.event_sources[tenant_name] = f"error: {str(e)[:60]}"
|
||||
self.connected_ws = False
|
||||
logger.debug(f"event source error for '{tenant_name}': {e}")
|
||||
await asyncio.sleep(backoff)
|
||||
backoff = min(backoff * 1.5, max_backoff_s)
|
||||
|
||||
async def _ws_listener(self):
|
||||
client_id = "hyperswap-arbitrator"
|
||||
ws_url = f"ws://127.0.0.1:8188/ws?clientId={client_id}"
|
||||
@@ -1121,55 +1210,209 @@ class AutoArbitrator:
|
||||
backoff = min(backoff * 1.5, 15.0)
|
||||
|
||||
RECLAIM_COOLDOWN_S = 30.0
|
||||
# A queue entry that has claimed to be running this long without the GPU ever going
|
||||
# busy is stale, not slow.
|
||||
STALE_RUNNING_S = 90.0
|
||||
# ComfyUI's own VRAM, not GPU utilisation, is what distinguishes a real job from a
|
||||
# stale row. Utilisation is shared: Ollama and any third-party process drive it too,
|
||||
# so peak utilisation stayed above any sensible threshold and a stuck entry never
|
||||
# looked stale. A real diffusion job loads gigabytes of checkpoint; a dead one holds
|
||||
# only the CUDA context.
|
||||
STALE_COMFY_BYTES = 1.5 * 1024 ** 3
|
||||
|
||||
async def _check_ollama_starved(self) -> None:
|
||||
"""The other direction: rescue an LLM that ComfyUI has squeezed onto the CPU.
|
||||
def _comfy_genuinely_busy(self, queue: Dict[str, Any]) -> bool:
|
||||
"""Decide whether ComfyUI is really working, not just claiming to be.
|
||||
|
||||
Yielding Ollama for ComfyUI was automatic; the reverse never was, despite the
|
||||
README calling the arbitration bidirectional. When Ollama cannot fit a model it
|
||||
does not fail, it silently places layers on the CPU and runs about an order of
|
||||
magnitude slower -- so this is the failure mode a user is least likely to notice
|
||||
and most likely to feel.
|
||||
ComfyUI can leave an entry in queue_running after a job dies -- observed here as
|
||||
a WAN 2.1 i2v entry that sat there with the GPU at 0% and ComfyUI holding 0.56 GB.
|
||||
Trusting that flag alone made this service believe ComfyUI was permanently busy,
|
||||
which meant it evicted the LLM on every poll, never ran the idle purge, and never
|
||||
checked whether the LLM had been squeezed onto the CPU. Half the arbitration was
|
||||
disabled by one stale row.
|
||||
|
||||
If the LLM is spilling while ComfyUI sits idle holding VRAM, ComfyUI's cached
|
||||
checkpoints are the thing to give up.
|
||||
A running entry is corroborated against GPU utilisation before it is believed.
|
||||
"""
|
||||
now = time.time()
|
||||
if self.comfy_was_active or (now - self.last_reclaim_time) < self.RECLAIM_COOLDOWN_S:
|
||||
return
|
||||
running = queue.get("queue_running") or []
|
||||
pending = queue.get("queue_pending") or []
|
||||
if pending:
|
||||
self._running_since = None
|
||||
self._running_id = None
|
||||
return True
|
||||
if not running:
|
||||
self._running_since = None
|
||||
self._running_id = None
|
||||
self.comfy_stale_job = None
|
||||
return False
|
||||
|
||||
ollama = await get_ollama_live_state()
|
||||
if not ollama.get("partially_offloaded"):
|
||||
return
|
||||
entry = running[0]
|
||||
prompt_id = entry[1] if isinstance(entry, (list, tuple)) and len(entry) > 1 else str(entry)
|
||||
now = time.time()
|
||||
if prompt_id != self._running_id:
|
||||
self._running_id = prompt_id
|
||||
self._running_since = now
|
||||
self._peak_comfy_bytes = 0
|
||||
|
||||
snap = get_process_vram_bytes()
|
||||
if snap["comfyui_bytes"] < RECLAIM_MIN_COMFY_BYTES:
|
||||
return # ComfyUI is not the one holding the memory; nothing we can do here
|
||||
self._peak_comfy_bytes = max(self._peak_comfy_bytes, snap.get("comfyui_bytes", 0))
|
||||
|
||||
self.last_reclaim_time = now
|
||||
model = ollama.get("active_model_name")
|
||||
offload = ollama.get("cpu_offload_pct")
|
||||
logger.warning(f"⚠ '{model}' is {offload}% on CPU while ComfyUI holds "
|
||||
f"{round(snap['comfyui_bytes'] / (1024**3), 2)} GB — reclaiming for the LLM")
|
||||
await self._purge_comfy_now(f"LLM spilling {offload}% to CPU")
|
||||
self.stats["reclaims_for_ollama"] += 1
|
||||
elapsed = now - (self._running_since or now)
|
||||
if elapsed > self.STALE_RUNNING_S and self._peak_comfy_bytes < self.STALE_COMFY_BYTES:
|
||||
if self.comfy_stale_job != prompt_id:
|
||||
logger.warning(
|
||||
f"ComfyUI reports prompt {prompt_id} running for {int(elapsed)}s while "
|
||||
f"holding only {self._peak_comfy_bytes / (1024**3):.2f} GB — no checkpoint "
|
||||
f"is loaded, so the queue entry is stale. Ignoring it; otherwise ComfyUI "
|
||||
f"looks permanently busy and arbitration stops working.")
|
||||
self.comfy_stale_job = prompt_id
|
||||
return False
|
||||
return True
|
||||
|
||||
# Freeing VRAM does not move layers back; only a reload re-places the model. Do
|
||||
# that only when the model is idle, never mid-generation.
|
||||
after = get_process_vram_bytes()
|
||||
if after.get("gpu_util_pct", 0) < BUSY_UTIL_PCT and model:
|
||||
logger.info(f"Reloading '{model}' to place it fully on the GPU...")
|
||||
await instant_free_ollama_vram(model, confirm=True)
|
||||
res = await switch_ollama_model(model, keep_alive="30m")
|
||||
recheck = await get_ollama_live_state()
|
||||
self.last_action = (
|
||||
f"Reclaimed {round(snap['comfyui_bytes'] / (1024**3), 2)}GB from ComfyUI and "
|
||||
f"reloaded '{model}' — now {round(recheck.get('gpu_fraction', 0) * 100)}% on GPU"
|
||||
if res.get("success") else
|
||||
f"Reclaimed VRAM from ComfyUI but reloading '{model}' failed: {res.get('error')}")
|
||||
async def _tenant_state(self, active_tenant: Optional[str] = None
|
||||
) -> List[Dict[str, Any]]:
|
||||
"""Current VRAM and busy state for every configured tenant.
|
||||
|
||||
`active_tenant` skips the HTTP busy probe for the tenant whose event stream just
|
||||
fired: the event is the evidence. That removes a round trip from the handoff,
|
||||
which is the one path where latency is the entire point.
|
||||
"""
|
||||
# Only the cheap NVML read. The full hardware snapshot also does a psutil lookup
|
||||
# per process, which is wasted work on the handoff path where latency is the
|
||||
# entire point.
|
||||
snap = get_process_vram_bytes()
|
||||
by_tenant = {k: b / (1024 ** 3) for k, b in snap["by_tenant_bytes"].items()}
|
||||
out = []
|
||||
for t in tenants_mod.load_tenants():
|
||||
if not t.enabled:
|
||||
continue
|
||||
bucket = _BUCKET_ALIASES.get(t.name, t.name)
|
||||
vram_gb = by_tenant.get(bucket, 0.0)
|
||||
if t.name == active_tenant:
|
||||
probe = {"busy": True, "reason": "event received from its own stream"}
|
||||
else:
|
||||
probe = await tenants_mod.probe_busy(t, vram_gb=vram_gb)
|
||||
out.append({
|
||||
"name": t.name,
|
||||
"priority": t.priority,
|
||||
"vram_gb": vram_gb,
|
||||
"busy": bool(probe.get("busy")),
|
||||
"below_floor": bool(probe.get("below_floor")),
|
||||
"reclaimable": t.reclaimable,
|
||||
"needs_vram_gb": t.needs_vram_gb,
|
||||
"overclock_profile": t.overclock_profile,
|
||||
"vram_floor_gb": t.vram_floor_gb,
|
||||
"idle_release_after_s": t.idle_release_after_s,
|
||||
"reason": probe.get("reason"),
|
||||
})
|
||||
self._last_tenant_state = {"ts": time.time(), "free_gb":
|
||||
round(snap["free_bytes"] / (1024**3), 2),
|
||||
"tenants": out}
|
||||
return out
|
||||
|
||||
async def _release_tenant(self, name: str, reason: str) -> Dict[str, Any]:
|
||||
"""Release one tenant's VRAM by whatever mechanism it declares."""
|
||||
t = tenants_mod.get_tenant(name)
|
||||
if not t or not t.reclaimable:
|
||||
return {"success": False, "reason": "not reclaimable"}
|
||||
models = None
|
||||
if t.release.per_model:
|
||||
state = await get_ollama_live_state()
|
||||
models = [m.get("name") for m in state.get("loaded_models", []) if m.get("name")]
|
||||
logger.info(f"Releasing VRAM from '{name}': {reason}")
|
||||
res = await tenants_mod.release_vram(t, models=models)
|
||||
self.stats["tenant_releases"] = self.stats.get("tenant_releases", 0) + 1
|
||||
return res
|
||||
|
||||
IDLE_PROFILE = "balanced"
|
||||
|
||||
def _apply_profile_for_active(self, state: List[Dict[str, Any]]) -> None:
|
||||
"""Apply the GPU profile declared by whichever tenant is currently working.
|
||||
|
||||
This used to be two calls naming 'comfy' and 'ollama' directly, so a third
|
||||
application could never get tuned clocks. The highest-priority busy tenant wins;
|
||||
with nothing working the card returns to the idle profile.
|
||||
"""
|
||||
busy = [s for s in state if s["busy"] and s.get("overclock_profile")]
|
||||
if busy:
|
||||
busy.sort(key=lambda s: -s["priority"])
|
||||
self._apply_oc_profile(busy[0]["overclock_profile"])
|
||||
else:
|
||||
self.last_action = (f"Reclaimed VRAM from ComfyUI; '{model}' is busy, so it will "
|
||||
f"stay partly on CPU until its next load")
|
||||
self._apply_oc_profile(self.IDLE_PROFILE)
|
||||
|
||||
async def _arbitrate(self, active_tenant: Optional[str] = None) -> None:
|
||||
"""Generic arbitration over any number of tenants.
|
||||
|
||||
The two-application version was a pair of hardcoded rules -- yield Ollama when
|
||||
ComfyUI is busy, purge ComfyUI when Ollama is starved -- which could not express
|
||||
a third participant at all. This works from the registry instead: a busy tenant
|
||||
that lacks the VRAM it declares it needs is starved, and the memory comes from
|
||||
idle reclaimable tenants below it in priority, lowest first.
|
||||
"""
|
||||
t_start = time.perf_counter()
|
||||
state = await self._tenant_state(active_tenant)
|
||||
free_gb = self._last_tenant_state["free_gb"]
|
||||
self._apply_profile_for_active(state)
|
||||
|
||||
# 1. Starvation: highest-priority demanding tenant first.
|
||||
for s in sorted(state, key=lambda x: -x["priority"]):
|
||||
if not s["busy"] or not s["needs_vram_gb"]:
|
||||
continue
|
||||
# Starved means it cannot reach what it needs even counting what it already
|
||||
# holds. Comparing free VRAM alone flagged a tenant that was working
|
||||
# perfectly well on 13 GB as demanding, purely because little was left over
|
||||
# -- which is the normal state of a busy GPU, and would have caused
|
||||
# pointless releases from everyone else.
|
||||
if s["vram_gb"] + free_gb >= s["needs_vram_gb"]:
|
||||
continue
|
||||
plan = tenants_mod.plan_release(s["name"], state, free_gb, s["needs_vram_gb"])
|
||||
self.last_arbitration = {"ts": time.time(), "demanding": s["name"],
|
||||
"free_gb": free_gb, **plan}
|
||||
if not plan["release"]:
|
||||
logger.debug(f"'{s['name']}' is short of VRAM but {plan['reason']}")
|
||||
return
|
||||
if time.time() - self.last_reclaim_time < self.RECLAIM_COOLDOWN_S:
|
||||
return
|
||||
self.last_reclaim_time = time.time()
|
||||
for victim in plan["release"]:
|
||||
await self._release_tenant(
|
||||
victim, f"{s['name']} needs {s['needs_vram_gb']} GB, {free_gb} GB free")
|
||||
# Wait for the memory to actually come back, and record how long the whole
|
||||
# handoff took. Swap speed is the point of this service, so it is measured
|
||||
# rather than assumed.
|
||||
target_bytes = int(s["needs_vram_gb"] * (1024 ** 3))
|
||||
deadline = time.perf_counter() + 30.0
|
||||
while time.perf_counter() < deadline:
|
||||
if get_process_vram_bytes()["free_bytes"] >= target_bytes:
|
||||
break
|
||||
await asyncio.sleep(0.02)
|
||||
handoff_ms = round((time.perf_counter() - t_start) * 1000, 1)
|
||||
self.last_handoff = {"ts": time.time(), "to": s["name"],
|
||||
"released": plan["release"], "handoff_ms": handoff_ms,
|
||||
"triggered_by": "event" if active_tenant else "poll"}
|
||||
self.stats["handoffs"] = self.stats.get("handoffs", 0) + 1
|
||||
logger.info(f"Handoff to '{s['name']}' in {handoff_ms} ms "
|
||||
f"(released {', '.join(plan['release'])})")
|
||||
self.last_action = (f"Released {', '.join(plan['release'])} so "
|
||||
f"'{s['name']}' could work — {handoff_ms} ms")
|
||||
return
|
||||
|
||||
# 2. Idle release: a tenant holding VRAM it is not using, after a grace period.
|
||||
now = time.time()
|
||||
for s in state:
|
||||
if not s["reclaimable"] or s["vram_gb"] <= 0.25:
|
||||
self._idle_since.pop(s["name"], None)
|
||||
continue
|
||||
if s["busy"]:
|
||||
self._idle_since.pop(s["name"], None)
|
||||
continue
|
||||
since = self._idle_since.setdefault(s["name"], now)
|
||||
grace = s["idle_release_after_s"]
|
||||
if grace and (now - since) >= grace:
|
||||
self._idle_since.pop(s["name"], None)
|
||||
await self._release_tenant(
|
||||
s["name"], f"idle {int(now - since)}s holding {s['vram_gb']} GB")
|
||||
self.last_action = (f"Released idle '{s['name']}' after "
|
||||
f"{int(now - since)}s")
|
||||
return
|
||||
|
||||
async def _poll_watchdog(self):
|
||||
"""Fallback for when the WebSocket is down. One cheap /queue call, 1 Hz.
|
||||
@@ -1184,15 +1427,25 @@ class AutoArbitrator:
|
||||
resp = await client.get("/queue")
|
||||
if resp.status_code == 200:
|
||||
q = resp.json()
|
||||
busy = len(q.get("queue_running", [])) > 0 or len(q.get("queue_pending", [])) > 0
|
||||
busy = self._comfy_genuinely_busy(q)
|
||||
if busy:
|
||||
self.watchdog_branches["busy"] += 1
|
||||
await self.trigger_comfy_priority("Watchdog saw an active queue")
|
||||
elif self.comfy_was_active:
|
||||
self.watchdog_branches["completed"] += 1
|
||||
await self.trigger_comfy_completed()
|
||||
else:
|
||||
await self._check_ollama_starved()
|
||||
except Exception:
|
||||
pass
|
||||
self.watchdog_branches["idle_check"] += 1
|
||||
await self._arbitrate()
|
||||
else:
|
||||
self.watchdog_branches["bad_status"] += 1
|
||||
except Exception as e:
|
||||
# This used to swallow everything silently, including anything raised by
|
||||
# the starvation check, which is why that check could appear to run and
|
||||
# do nothing.
|
||||
self.watchdog_branches["error"] += 1
|
||||
self.last_watchdog_error = str(e)[:200]
|
||||
logger.debug(f"watchdog poll error: {e}")
|
||||
await asyncio.sleep(interval)
|
||||
|
||||
def suspend_oc(self, reason: str = "tuning sweep") -> None:
|
||||
@@ -1230,6 +1483,13 @@ class AutoArbitrator:
|
||||
"idle_purge_after_s": self.COMFY_IDLE_PURGE_S,
|
||||
"oc_profile": self.oc_profile,
|
||||
"counters": dict(self.stats),
|
||||
"comfy_stale_job": self.comfy_stale_job,
|
||||
"event_sources": dict(self.event_sources),
|
||||
"last_arbitration": self.last_arbitration,
|
||||
"last_handoff": self.last_handoff,
|
||||
"tenant_state": self._last_tenant_state,
|
||||
"watchdog_branches": dict(self.watchdog_branches),
|
||||
"last_watchdog_error": self.last_watchdog_error,
|
||||
"yield_backoff": {m: round(max(t - time.time(), 0), 1)
|
||||
for m, t in self._yield_backoff_until.items()
|
||||
if t > time.time()},
|
||||
|
||||
Reference in New Issue
Block a user