Compare commits

...

13 Commits

Author SHA1 Message Date
drjones
48fbde040c Test the job queue, and surface it in the dashboard and MCP
The queue shipped with no tests despite having produced four bugs during
development, which is the wrong order. 21 tests now cover the parts whose failure
modes are not obvious from reading the code: ordering by priority then FIFO, that
depth is genuinely unbounded and survives a restart, that running work is never
cancelled, that a job abandoned mid-run is requeued rather than left RUNNING
forever, and that an LLM job's VRAM requirement comes from its own model rather
than a tenant-wide figure -- the mistake that dispatched a 14.9 GB model into 8 GB
of free memory and killed llama-server three times.

The dashboard had no view of the queue at all, so a stuck queue was
indistinguishable from an empty one. The Job Queue panel shows what is running and
what it released to get there, pending jobs in execution order with their wait time
and a cancel control, and -- when the scheduler is blocked -- how long it has been
waiting and why.

MCP gains queue_job, get_job_queue and cancel_job, so an agent can line work up
rather than firing a request and hoping the GPU is free.

Tests: 271.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 17:41:55 -07:00
drjones
0fbc3963b9 Add a durable cross-tenant job queue with VRAM-aware scheduling
The service only reacted: it noticed an application had started and scrambled to free
memory. Nothing could be lined up. Each application has its own queue but they cannot
see each other, so work submitted to one had no way to wait for the other.

Jobs are stored in SQLite, so the queue is bounded by disk rather than memory and
survives a restart. The scheduler takes the highest-priority pending job, arbitrates
VRAM for it through the same plan_release, runs it, and moves on -- one at a time,
because overlapping jobs would recreate the contention this service exists to resolve.

Four bugs found by running it rather than reasoning about it:

Dispatching without checking for room destroyed three queued LLM jobs in a row: a CUDA
OOM kills llama-server outright, it does not fail gracefully. A job that cannot run yet
now waits.

The room check used the tenant's needs_vram_gb, which cannot be right for an LLM --
the requirement is a property of the model being loaded. A flat 4 GB passed with 8 GB
free and then a 14.9 GB model was dispatched into it. The requirement is now computed
per job.

Waiting forever is also wrong. Three jobs sat pending indefinitely needing 14.93 GB on
a card where at most ~14.8 GB can ever be free, because an unreclaimable process holds
0.82 GB. A job that cannot be satisfied now fails with the ceiling and the blockers
named.

plan_release assumed releasing a tenant frees everything it holds. ComfyUI keeps its
CUDA context for as long as the process lives, so it reported that releasing ComfyUI
would free 0.37 GB against a 0.33 GB shortfall; the job was cleared and the memory
never arrived. Tenants declare vram_floor_gb and only memory above it counts.

An exception during dispatch left the row RUNNING forever while the scheduler moved on.
Failures now land on the job, and jobs left running by a previous process are requeued
at startup.

Verified end to end: five mixed jobs across both applications, queued at once, all
completed with no failures.

Tests: 250.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 17:22:34 -07:00
drjones
48bd0096ff Let a higher-priority job preempt lower-priority work in progress
Measured failure: a diffusion job took 46 s instead of 3 s, squeezed into 1.6 GB.
The LLM yielded correctly, an inference reloaded it two seconds later, and
plan_release then refused to touch it because it was "busy" -- so ComfyUI crawled
while Ollama held 13 GB for the whole run.

"Never interrupt busy work" looks like the safe rule and is not. Preempting a
lower-priority tenant is safe precisely because releasing is asynchronous: an Ollama
unload queues behind its running request and applies when that finishes, so nothing
is killed mid-flight. That is what makes fast handoff possible at all, and refusing
to do it defeats the purpose of the service.

A tenant may now be asked for memory if it is idle, whatever its rank, or if it is
busy and ranks strictly below the demander. Equal or higher priority is never
interrupted, so peers cannot fight. Idle tenants are still preferred over preempting
busy ones.

Verified against the real contention: with Ollama at 12.38 GB and 98% utilisation, a
diffusion job released it within three seconds and completed in 18 s rather than 46 s.

Tests updated to the corrected rule, and the fixture's priorities aligned with what
actually ships -- it still had the LLM outranking diffusion from before that was
swapped.

Tests: 246.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 16:57:08 -07:00
drjones
ca97f18be6 Let any tenant declare its GPU profile and event source; fix priority semantics
Two remaining pieces of the two-application coupling are gone.

Overclock profiles were switched by naming 'comfy' and 'ollama' directly, so a third
application could never get tuned clocks. A tenant declares overclock_profile and the
arbitrator applies whichever the highest-priority *working* tenant asks for, falling
back to the idle profile when nothing is running.

The websocket listener parsed ComfyUI's message schema -- status, execution_start,
executing, execution_success -- which tied the fast path to one application. An event
source is now declarative and the messages are not parsed at all: any message means
"look now", and the tenant's own busy probe decides what is true. That gives the same
sub-second reaction to any application that emits anything on state change, with no
knowledge of what it emits.

Generalising this exposed a design error in the priority rule I had introduced.
plan_release excluded candidates ranking above the demander, which broke both
directions in turn. With the LLM at priority 60 and diffusion at 50, ComfyUI could
never reclaim from Ollama -- the premise the whole service is built on, and preserved
until now only by the ComfyUI-specific trigger that was about to be removed. Swapping
the ranks then broke the reverse: a starved Ollama could no longer reclaim from an
idle ComfyUI.

Priority now orders rather than vetoes. Any idle reclaimable tenant is a candidate,
because an idle tenant is not using its VRAM; priority decides who is asked first, and
busy tenants are never interrupted whatever their rank. Diffusion outranks the LLM,
whose weights reload from page cache in seconds. All three cases are pinned by tests,
including that busy work is never interrupted even by a far higher-priority demander.

Tests: 244 (was 242).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 15:23:48 -07:00
drjones
c6455d7c6e Remove the duplicated starvation path; show every tenant; cache-bust assets
_check_ollama_starved and _arbitrate were solving the same problem, one of them
hardcoded to two applications. The Ollama-specific version is gone and the watchdog
calls only the generic loop. The OOM retry inside switch_ollama_model no longer
purges ComfyUI by name either: it asks plan_release which tenant should give up
memory, so a third application can be the one that yields, and when the reclaim is
not enough the response names the blockers instead of implying ComfyUI was at fault.

The dashboard showed exactly two engines, which no longer matched what the service
does. A GPU Tenants panel lists every configured application ordered by priority --
VRAM held, whether it is working, whether it can be reclaimed at all, and how much it
needs -- along with the last arbitration decision and why it could or could not be
satisfied.

That panel did not appear at first, and the reason is worth fixing rather than
working around: the browser kept serving a cached app.js despite the ETag, so a
reload ran the old dashboard against the new API. Assets are now stamped with their
mtime, so a changed file is always a different URL. Anyone updating this service would
have hit the same thing.

Also verified along the way that an apparent horizontal-overflow regression was a
measurement artifact from a zero-width browser pane, not a real layout fault.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 15:16:10 -07:00
drjones
2b00ab4e12 Make the live graph scale with the window
Three things pinned it small. The container was a fixed h-64, so the graph stayed
256px tall on any display; the page was capped at max-w-7xl, so a large window only
added empty margins; and the graph shared a two-column row, so it never got more than
half the available width.

The container is now h-[clamp(18rem,45vh,52rem)], the page cap is 2600px, and the
graph panel spans its row. Measured: 525x360 to 1142x360 at 1280x800, and 988x648 to
2468x648 at 3440x1440.

Chart.js needed help despite being configured responsive with maintainAspectRatio
false. It had latched onto a stale size -- the canvas sat at width:0px, height:288px
while its container had grown to 988x648. A ResizeObserver on the container fixed
growth but not shrinking, because passing explicit dimensions to resize() left the
canvas 988px wide inside a 435px container, overflowing it. Calling resize() with no
arguments lets Chart.js measure the container itself, and the canvas is absolutely
positioned within the relative container so it cannot force it wider.

Verified from 1100x700 to 3440x1440: the canvas matches its container exactly at every
size, growing and shrinking, with no horizontal page overflow.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 14:52:23 -07:00
drjones
4a38cd68b3 Arbitrate over any number of tenants by priority
Completes the generalisation. Classification and release were already data; the
decision loop was still two hardcoded rules -- yield Ollama when ComfyUI is busy,
purge ComfyUI when Ollama is starved -- which could not express a third participant.

plan_release() works from the registry instead. A busy tenant that cannot reach its
declared needs_vram_gb is starved, and the memory comes from idle reclaimable tenants
below it in priority, lowest first, stopping once enough is freed. Tenants that cannot
be released are named as blockers rather than passed over, so an impossible plan says
which process is in the way. The plan is returned before being acted on, so the
decision is testable and is logged before anything is released. Idle release is now
per-tenant too, replacing the ComfyUI-specific purge timer.

Two bugs found by running it against the live machine rather than only in tests:

Starvation was measured against free VRAM alone, so a tenant working perfectly well on
13 GB was flagged as demanding simply because little was left over -- which is the
normal state of a busy GPU, and would have caused pointless releases from everything
else. A tenant is starved only if it cannot reach what it needs counting what it
already holds.

Fields added to the tenant schema were silently absent from the config already written
to disk, so needs_vram_gb defaulted to 0 and starvation could never trigger for the two
tenants that mattered. Shipped defaults are now merged into an existing config on load,
with explicit user values still winning.

Tests: 242 (was 231), including a three-application contention case -- the property the
hardcoded pair of rules could not express.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 14:47:32 -07:00
drjones
1bcfbb2335 Describe GPU tenants as data so any application can be arbitrated
The point of this service is fast handoff of one GPU between applications. It grew up
around the two on this box, and their names ended up compiled into process matching,
VRAM attribution, busy detection and release calls alike -- about 385 references
across five modules. That made it a script for Ollama and ComfyUI rather than a GPU
arbitrator.

tenants.py describes an application as data: how to recognise its processes, how to
tell whether it is genuinely working, how to ask it for VRAM back, and how much it
matters when two want the card. Ollama, ComfyUI and the desktop compositor ship as
defaults in tenants.json, so behaviour is unchanged, but the arbitration logic no
longer knows any particular name. Endpoints are generic: GET /api/tenants,
GET /api/tenants/{name}, POST /api/tenants/{name}/release -- the last being the
general form of both the Ollama soft-yield and the ComfyUI purge.

Verified by registering a third application on this machine with no code change: the
speech relay that had been showing up only as anonymous "unmanaged VRAM" is now named,
attributed, and probed by the VRAM it holds rather than by an API it does not have.
Because it declares no release strategy, a release request returns 409 explaining that
its memory cannot be reclaimed, instead of reporting a success that did nothing.

Busy probes deliberately cannot use GPU utilisation. It is shared by every tenant, so
it cannot attribute work to one of them -- the mistake that made a stale ComfyUI queue
entry undetectable earlier in this branch. A tenant's own VRAM is the signal.

Writing the tests exposed that the suite had become non-hermetic: classification is now
configuration, so a test asserting "a third-party process is unmanaged" started failing
the moment the speech relay was registered on this machine. An autouse fixture now
isolates every test from the operator's live tenants.json.

Tests: 231 (was 206).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 11:20:44 -07:00
drjones
c3d9b36035 Stop a stale ComfyUI queue entry from disabling half the arbitration
Chasing why the reverse-direction reclaim never fired turned up something worse than
the reclaim itself.

The starvation check was never running. Instrumenting the watchdog showed busy=6,
idle_check=0: every poll took the "ComfyUI is busy" branch. ComfyUI's /queue was
reporting a WAN 2.1 i2v job in queue_running while the GPU sat at 0% and ComfyUI held
0.56 GB. The job was dead; ComfyUI had simply never cleared the row.

Believing that flag meant this service thought ComfyUI was permanently busy, so it
yielded the LLM's VRAM on every poll, never ran the idle purge, and never checked
whether the LLM had been squeezed onto the CPU. One stale row disabled half of the
arbitration, and it very likely explains the earlier burst of yields against a
cron-driven model.

A running entry is now corroborated before it is believed. The first attempt used GPU
utilisation, which does not work: utilisation is shared with Ollama and with the
third-party process on this box, so peak utilisation stayed above any sensible
threshold and a stuck entry never looked stale. ComfyUI's own VRAM is the right
signal -- a real diffusion job loads gigabytes of checkpoint, a dead one holds only
its CUDA context. After the fix the same watchdog reports busy=3, idle_check=32.

Every early return in the starvation check now records why it bailed, because with
four of them there was no way to tell which had fired. /api/health reports a stale
queue entry with its impact and how to clear it.

Also confirmed, contradicting an earlier conclusion in this branch: Ollama on this box
*does* spill to the CPU. smtek/Qwen3.8-27B:Q2_K_XL held steady at 29.2% on GPU
(size=15.59 GB, size_vram=4.56 GB) across twelve seconds of polling -- a stable
placement, not a progressive load. Both failure modes are real; which one occurs
depends on the model.

Tests: 206 (was 199).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-07 08:48:37 -07:00
drjones
aeba1b47fd Make unreclaimable VRAM actionable, and account for ComfyUI's CUDA context
The health check now reports what unmanaged VRAM actually costs rather than just
how much of it there is: "0.82 GB held by python (842 MB)" becomes "5 model(s) fit
within 15.42 GB but not the 14.60 GB actually available", naming them.

Getting that arithmetic right took a correction. The first version subtracted only
the desktop and the unmanaged process, and so reported a 14.93 GB model as fitting
against a real ceiling of 14.60 GB -- the same model the service had just refused
with 507. ComfyUI keeps a few hundred MB of CUDA context for as long as the process
lives, which a purge does not free, so it is not available either. The floor is taken
from the minimum ComfyUI VRAM in recent telemetry rather than its current value,
which could be a 7 GB checkpoint mid-generation.

The verifier's reclaim stage now re-runs a graph immediately beforehand to reset the
30 s idle window, since a large model takes longer than that to load and the purge
was freeing ComfyUI mid-load, so the reclaim path was never reached.

Tests: 199 (was 192). The new ones pin the ceiling arithmetic, including that a model
too large to fit on the card at all is not blamed on the third-party process.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-06 17:44:21 -07:00
drjones
81e5d88426 Make the verifier's results mean what they say
The first full run passed every stage, and three of those passes were worth less
than they looked.

Diffusion was reported as 0.67 it/s. That single run included loading the SDXL
checkpoint from disk, so it understated throughput roughly tenfold against a
steady-state 5.49 it/s. Cold and warm are now timed and labelled separately.

The idle-purge stage checked the flag immediately, racing the ComfyUI websocket
event that sets it, and reported "no purge pending" as a warning about ComfyUI
rather than about its own timing. It now waits for the event.

The reclaim stage passed while proving nothing: the model chosen was small enough to
fit alongside ComfyUI's checkpoint, so the reclaim path never ran. It now picks a
model that genuinely will not fit, and reports a warning rather than a pass when the
path is not exercised.

Sizing that model correctly took two corrections, both real. On-disk weight size is
not the VRAM footprint -- a 12.87 GB blob occupies 14.9 GB once context and KV cache
are allocated -- and unmanaged VRAM is not reclaimable, so it cannot count toward
what a reclaim will free. Ignoring the second picked a model that failed even after a
correct reclaim: the service returned 507 and logged "could not fit with ComfyUI
holding 7.03 GB -- reclaiming and retrying", which was right. On this box an 842 MB
third-party process is the difference between a 14.9 GB model fitting and not.

The verifier also died on the 507 instead of reporting it, since a helper called
raise_for_status() on responses a stage deliberately provokes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-06 17:31:48 -07:00
drjones
aed1c360f0 Add an end-to-end arbitration verifier; return honest HTTP status codes
The unit suite covers logic in isolation, but the promise this service exists to make
-- an LLM and a diffusion pipeline sharing one 16 GB card without either failing --
had only ever been checked by hand, piecemeal. verify_arbitration.py walks the whole
cycle against real hardware and reports what happened at each stage: load and its
bandwidth classification, the confirmed yield, a real SDXL graph, the deferred idle
purge, reclaim-and-retry, and finally whether VRAM attribution adds up and the
reported GPU state still matches the card. It restores what it changes and refuses
to start if ComfyUI is busy. Kept out of pytest deliberately: it moves real VRAM and
takes minutes.

Running it immediately found two bugs.

/api/switch-model reported every upstream failure as 500. Asking an embedding model
to generate makes Ollama return 400 -- the request is unusable, the service is fine
-- and calling that an Internal Server Error blames this service for the caller's
mistake. Failures now map to 400 for an upstream client error, 507 for a model that
will not fit (valid request, healthy service, no room), and 502 when Ollama itself
errors.

The verifier also picked the smallest installed model, which here is
nomic-embed-text -- an embedding model with no generate endpoint. It now filters
those out by family and name.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-06 17:25:42 -07:00
drjones
043d61722b Read engine configuration live instead of asserting it in the dashboard
The engine subtitles were hardcoded: "FlashAttention + Q4 KV Cache" and "DynamicVRAM
+ Pinned Async Offload". The first turned out to be accurate -- OLLAMA_FLASH_ATTENTION
and OLLAMA_KV_CACHE_TYPE really are set -- which is worse than being wrong, because it
would have gone on looking accurate after the settings changed.

engines.py reads both engines' real configuration: the ollama service environment via
systemd, and ComfyUI's own /system_stats for version, allocator, VRAM mode and argv.
Exposed at GET /api/engines, as an MCP tool, and in the dashboard subtitles with the
full settings list as a tooltip.

The settings worth surfacing are the ones that dictate how this service must behave
and that previously had to be discovered by reading journald: OLLAMA_NUM_PARALLEL=1
is why an unload queues behind a running generation and is reported as deferred
rather than failed, and OLLAMA_MAX_LOADED_MODELS=1 is why every swap evicts the
previous model. Each is reported with that explanation attached.

Writing the tests found a bug in the new code: (system.get("python_version") or
"").split()[0] raises IndexError when ComfyUI omits the field, and the surrounding
except would have swallowed it and reported ComfyUI as entirely offline.

Tests: 192 (was 182).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-06 10:56:04 -07:00
19 changed files with 3399 additions and 107 deletions

143
README.md
View File

@@ -21,6 +21,15 @@
### ⚡ Bidirectional VRAM Hot-Swapping & Arbitration
* **A stale ComfyUI queue entry no longer disables arbitration.** ComfyUI can leave a
dead job in `queue_running` indefinitely; one was found sitting there with the GPU idle
and ComfyUI holding 0.56 GB. Trusting that flag made this service believe ComfyUI was
permanently busy — so it evicted the LLM on every poll, never ran the idle purge, and
never checked for CPU spill. Instrumenting the watchdog showed `busy=6, idle_check=0`.
A running entry is now corroborated against ComfyUI's own VRAM (a real job loads
gigabytes; a dead one holds only its CUDA context) before it is believed, and a stale
entry is reported by `/api/health`. Utilisation is deliberately *not* the signal — it is
shared with Ollama and any third-party process.
* **Both directions are now automatic.** Yielding Ollama for ComfyUI always was; the
reverse was not, despite "bidirectional" in this heading. Which way an LLM fails when
it cannot fit depends on configuration: with `n_gpu_layers` left to Ollama it spills
@@ -111,6 +120,18 @@ only if the card actually needs it.
* **Honest gain reporting**: gain against the profile's *current* setting is reported separately from the spread across values tried. Conflating them turns a flat result into a headline "+102%".
* **Safety**: refuses to start while ComfyUI is executing, suspends the arbitrator's automatic profile switching for the duration (otherwise a diffusion benchmark trips `trigger_comfy_priority`, which reapplies the whole profile and overwrites the clock being measured), and restores the original profile in a `finally` block — including on exception or cancellation.
### 🔧 Live Engine Configuration (`engines.py`)
* `GET /api/engines` reports the **real, current** configuration of both engines and what
each setting implies for arbitration — because the settings that dictate this service's
behaviour live outside its own codebase.
* `OLLAMA_NUM_PARALLEL=1` is why a `keep_alive: 0` unload queues behind a running
generation and is reported as *deferred* rather than failed. `OLLAMA_MAX_LOADED_MODELS=1`
is why every swap evicts the previous model. Working these out originally meant reading
journald and the systemd unit by hand.
* The dashboard's engine subtitles now come from this endpoint. They were previously
hardcoded — and happened to be accurate, which is worse than being wrong, since they
would have kept looking accurate after the configuration changed.
### 🩺 Dependency Self-Check (`health.py`)
* `GET /api/health` verifies **everything this service depends on**: NVML, passwordless
sudo for `nvidia-smi`, fan control through the headless X server, overclock drift, the
@@ -153,10 +174,102 @@ only if the card actually needs it.
---
## 1c. Lining Work Up
Until now this service only *reacted*: it noticed an application had started and
scrambled to free memory. Nothing could be queued. Each application has its own queue,
but they cannot see each other, so work submitted to one has no way to wait for the other.
```bash
curl -X POST localhost:9090/api/jobs -H 'Content-Type: application/json' -d '{
"tenant": "ollama", "label": "nightly-summary",
"payload": {"model": "qwen3.8fast:latest", "prompt": "..."}
}'
```
Jobs live in SQLite, so the queue is bounded by disk rather than memory and survives a
restart. The scheduler takes the highest-priority pending job, arbitrates VRAM for it with
the same `plan_release`, runs it, and moves on. One at a time by design — the GPU is the
scarce resource this service exists to hand between applications, and overlapping jobs
would just recreate the contention it resolves.
`GET /api/jobs` · `GET /api/jobs/{id}` · `DELETE /api/jobs/{id}` (pending only — running
work is never killed) · `DELETE /api/jobs` to clear the queue. Agents get the same through
MCP (`queue_job`, `get_job_queue`, `cancel_job`), and the dashboard's **Job Queue** panel
shows what is running, what it released to get there, and why the scheduler is waiting if
it is.
**A job that cannot run yet waits; a job that can never run fails with the reason.**
Dispatching into insufficient VRAM does not fail gracefully — it kills `llama-server`
with a CUDA OOM. The requirement is computed per job (an LLM job needs the size of *its*
model, not a tenant-wide figure), and if the memory can never be assembled the job fails
naming what stands in the way rather than blocking the queue forever.
## 1b. Any Application, Not Just These Two
The purpose is fast handoff of one GPU between applications. It grew up around the two on
this box, and their names ended up compiled into process matching, VRAM attribution, busy
detection and release calls alike — about 385 references. That made it a script for Ollama
and ComfyUI rather than a GPU arbitrator.
A tenant is now **described as data** in `tenants.json`:
```json
{
"name": "trainer",
"kind": "other",
"priority": 80,
"match": { "cmdline": ["train.py"] },
"busy": { "type": "vram", "vram_busy_gb": 1.0 },
"release": { "type": "http_post", "url": "http://localhost:9999/release" }
}
```
| Field | What it answers |
| :--- | :--- |
| `match` | Which GPU processes belong to this application (name, cmdline substring, or suffix — ComfyUI is a bare `python main.py`) |
| `busy` | Whether it is *genuinely* working. `http_count` sums queue lists; `vram` needs no API at all. `vram_floor_gb` catches a queue that claims work while nothing is loaded |
| `release` | How to ask for VRAM back — `http_post` with a body, `per_model` for Ollama's per-model unload, or `none` |
| `priority` | Who is asked to yield **first** among idle tenants — it never protects idle memory, and never interrupts work |
| `overclock_profile` | GPU profile applied while this tenant is the active workload |
| `events` | Optional stream (e.g. a websocket) used purely as a wake-up, so reaction is sub-second rather than waiting for the next poll |
Two more fields drive the decision loop: **`needs_vram_gb`** (how much free memory the
application needs before it can work) and **`idle_release_after_s`** (how long it may sit
idle holding VRAM before being asked for it back — deliberately not immediate, so
iterating on a ComfyUI workflow does not reload the checkpoint between every run).
**Priority orders, it does not veto.** An idle tenant is not using its VRAM, so
outranking the demander is no reason to keep it; busy tenants are never interrupted
whatever their rank. Getting this wrong broke both directions in turn — with the LLM
ranked above diffusion, ComfyUI could never preempt Ollama (the service's central
behaviour), and once the ranks were swapped, a starved Ollama could no longer reclaim
from an idle ComfyUI. Diffusion now outranks the LLM, whose weights reload from page
cache in seconds.
`plan_release()` then arbitrates generically: a busy tenant that cannot reach
`needs_vram_gb` *even counting what it already holds* is starved, and the memory is taken
from idle reclaimable tenants below it in priority, lowest first, stopping as soon as
enough is freed. Tenants that cannot be released are named as blockers rather than
ignored, so `possible: false` comes with the reason. The plan is returned before it is
acted on, which makes the decision testable and loggable.
Ollama, ComfyUI and the desktop compositor ship as defaults, so behaviour is unchanged —
but nothing in the arbitration logic knows their names, and three applications can
contend for the card as easily as two. The dashboard's **GPU Tenants** panel lists all of
them ordered by the priority arbitration actually considers, with the last decision and
why it could or could not be satisfied. Endpoints are generic:
`GET /api/tenants`, `GET /api/tenants/{name}`, `POST /api/tenants/{name}/release`.
A tenant with `"release": {"type": "none"}` is still worth declaring. The 842 MB speech
relay on this box cannot be reclaimed, and naming it turns anonymous "unmanaged VRAM" into
"held by stt-relay, which exposes no release API" — and a release request returns **409**
explaining that, rather than silently doing nothing.
## 1a. Tests
```bash
/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/ -q # 182 passed in ~3.7s
/home/drjones/comfy-mcp-venv/bin/python -m pytest tests/ -q # 271 passed in ~4.0s
```
Hermetic: no GPU, no network, no sleeps. An autouse fixture stubs `overclock_manager._sh`
@@ -173,6 +286,33 @@ which contradicts the hardware fails loudly rather than silently:
| Busy yield | VRAM held at ≥50% GPU utilisation | A mid-generation model is finishing, not failing |
| Residency confidence | probe trusted only at 100% | A 12-window probe once cleared 90% on a mostly-cold file |
## 1b. End-to-End Verification
```bash
python verify_arbitration.py # full cycle, a few minutes
python verify_arbitration.py --quick # skip the diffusion stages
```
The unit suite covers logic in isolation. This exercises the promise the service exists
to make — an LLM and a diffusion pipeline sharing one 16 GB card — against real hardware,
and reports what actually happened at each stage. It restores what it changes and refuses
to start if ComfyUI is busy.
A representative run on this machine:
| Stage | Result |
| :--- | :--- |
| LLM load, classified by achieved bandwidth | 1.96 GB in 1327 ms → 1.47 GB/s → Partial Cache |
| VRAM yield confirmed against NVML | released in 43 ms, 2.39 GB freed |
| Diffusion, cold (includes checkpoint load) | 17863 ms → 1.12 it/s |
| Diffusion, warm | 3645 ms → **5.49 it/s** |
| ComfyUI retains its checkpoint | 7.03 GB held through the idle window |
| VRAM attribution adds up | 15.58 GB attributed vs 15.80 GB NVML — Ollama 7.71 + ComfyUI 7.03 coexisting |
| Reported GPU state matches hardware | profile asks 320 W, card reports 320 W |
The stages report warnings rather than passes when they did not actually prove anything —
a reclaim that was never needed is not evidence that reclaiming works.
## 2. Architectural Overview
```mermaid
@@ -264,6 +404,7 @@ The HyperSwap server runs on port `9090` by default. Interactive OpenAPI/Swagger
| `/api/analytics/models` | `GET` | Recency/frequency model ranking used to prioritise the warm budget. |
| `/api/history?durable=true` | `GET` | Swap history from the persistent store rather than the in-memory ring. |
| `/api/db` | `GET` | Store location, row counts and how many hours of history are held. |
| `/api/engines` | `GET` | Live Ollama and ComfyUI configuration, with what each setting implies for arbitration. |
| `/api/health` | `GET` | Dependency self-check: NVML, sudo, fan control, drift, store, upstreams — each with impact and remediation. |
### Governor & Autotune Endpoints

117
engines.py Normal file
View File

@@ -0,0 +1,117 @@
"""Live configuration of the two engines HyperSwap arbitrates between.
Arbitration behaviour is largely dictated by settings that live outside this codebase.
Working out why a yield behaved the way it did meant reading journald and the ollama
unit by hand: OLLAMA_NUM_PARALLEL decides whether an unload queues behind a running
generation, OLLAMA_MAX_LOADED_MODELS decides whether more than one model can be
resident, and a pinned n_gpu_layers decides whether a model that will not fit spills to
the CPU or fails outright. Those are worth reading and explaining rather than hardcoding
into a dashboard subtitle that silently goes stale.
"""
import json
import logging
import subprocess
from typing import Any, Dict, List, Optional
import httpx
import vram_arbitrator
logger = logging.getLogger("engines")
# Settings that change how the arbitrator must behave, with what they imply.
OLLAMA_SETTING_NOTES = {
"OLLAMA_NUM_PARALLEL": (
"Requests per model. At 1, a keep_alive:0 unload queues behind any running "
"generation and applies when it finishes — which is why a busy model is "
"reported as deferred rather than failed."),
"OLLAMA_MAX_LOADED_MODELS": (
"How many models may be resident at once. At 1, Ollama evicts the previous "
"model on every swap."),
"OLLAMA_KEEP_ALIVE": (
"Default residency after a request. Long values keep VRAM occupied and make "
"ComfyUI wait for an explicit yield."),
"OLLAMA_FLASH_ATTENTION": "FlashAttention kernels for attention.",
"OLLAMA_KV_CACHE_TYPE": "KV cache quantisation; smaller types cut VRAM per context.",
"OLLAMA_NUM_BATCH": "Prompt-evaluation batch size.",
}
def _ollama_unit_environment() -> Dict[str, str]:
"""Read the ollama service's environment. Empty if it is not a systemd unit."""
env: Dict[str, str] = {}
try:
proc = subprocess.run(["systemctl", "show", "ollama", "-p", "Environment",
"--value"], capture_output=True, text=True, timeout=8)
for token in proc.stdout.split():
if "=" in token and token.startswith("OLLAMA"):
k, _, v = token.partition("=")
env[k] = v
except Exception as e:
logger.debug(f"could not read ollama unit environment: {e}")
return env
async def get_engine_config() -> Dict[str, Any]:
"""Real, live configuration of both engines, with arbitration implications."""
ollama_env = _ollama_unit_environment()
ollama_settings = [
{"key": k, "value": v, "means": OLLAMA_SETTING_NOTES.get(k, "")}
for k, v in sorted(ollama_env.items())
]
# A short, honest summary line to replace the dashboard's hardcoded subtitle.
feature_bits: List[str] = []
if ollama_env.get("OLLAMA_FLASH_ATTENTION") == "1":
feature_bits.append("FlashAttention")
kv = ollama_env.get("OLLAMA_KV_CACHE_TYPE")
if kv:
feature_bits.append(f"{kv} KV cache")
host = ollama_env.get("OLLAMA_HOST", "")
port = host.rsplit(":", 1)[-1] if ":" in host else "11434"
ollama = {
"port": port,
"settings": ollama_settings,
"summary": " + ".join(feature_bits) if feature_bits else "default configuration",
"max_loaded_models": ollama_env.get("OLLAMA_MAX_LOADED_MODELS"),
"num_parallel": ollama_env.get("OLLAMA_NUM_PARALLEL"),
"keep_alive": ollama_env.get("OLLAMA_KEEP_ALIVE"),
"config_source": "systemd unit environment" if ollama_env else "unavailable",
}
comfy: Dict[str, Any] = {"online": False}
try:
async with httpx.AsyncClient(timeout=4.0) as c:
r = await c.get(f"{vram_arbitrator.COMFY_API_BASE}/system_stats")
if r.status_code == 200:
data = r.json()
system = data.get("system", {})
argv = system.get("argv") or []
devices = data.get("devices") or []
dev = devices[0] if devices else {}
# The allocator is named in the device string; it is the closest thing
# ComfyUI reports to the "async offload" the old subtitle asserted.
dev_name = dev.get("name", "")
allocator = ("cudaMallocAsync" if "cudaMallocAsync" in dev_name
else "cudaMalloc" if "cudaMalloc" in dev_name else "unknown")
vram_flags = [a for a in argv
if a in ("--lowvram", "--novram", "--highvram", "--normalvram",
"--gpu-only", "--cpu")]
comfy = {
"online": True,
"version": system.get("comfyui_version"),
"pytorch": system.get("pytorch_version"),
# split()[0] on an absent version raises IndexError, which would have
# been swallowed by the except below and reported ComfyUI as offline.
"python": ((system.get("python_version") or "").split() or [None])[0],
"argv": argv,
"vram_mode": vram_flags[0] if vram_flags else "default (auto)",
"allocator": allocator,
"device": dev_name,
"summary": f"{allocator}, {vram_flags[0] if vram_flags else 'auto VRAM'}",
}
except Exception as e:
comfy = {"online": False, "error": str(e)[:120]}
return {"ollama": ollama, "comfyui": comfy}

View File

@@ -10,6 +10,7 @@ when it is missing and how to fix it. A degraded dependency should be loud.
"""
import asyncio
import logging
import time
import os
import time
from typing import Any, Dict, List
@@ -116,6 +117,25 @@ def _check_comfy_ws() -> Dict[str, Any]:
return _check("comfyui websocket", OK, "subscribed")
def _check_comfy_queue() -> Dict[str, Any]:
"""A stale ComfyUI queue entry disables half of this service's logic."""
arb = vram_arbitrator.arbitrator
if arb.comfy_stale_job:
return _check("comfyui queue", DEGRADED,
f"prompt {arb.comfy_stale_job} claims to be running but the GPU is idle",
"ComfyUI looks permanently busy, so the LLM is evicted repeatedly, "
"the idle purge never runs and CPU-spill is never checked",
"Clear it from the ComfyUI queue, or POST /queue with "
"{\"clear\": true} to ComfyUI")
branches = arb.watchdog_branches
if branches.get("idle_check", 0) == 0 and branches.get("busy", 0) > 20:
return _check("comfyui queue", DEGRADED,
"the watchdog has only ever seen ComfyUI as busy",
"The idle purge and starvation check are not running",
"Check the ComfyUI queue for a stuck entry")
return _check("comfyui queue", OK, "queue state corroborated against GPU activity")
def _check_store() -> Dict[str, Any]:
info = telemetry_store.db_info()
if not info.get("exists"):
@@ -139,6 +159,75 @@ def _check_residency() -> Dict[str, Any]:
cap.get("hint", ""))
def _comfy_vram_floor_gb(default: float = 0.0, days: float = 1.0) -> float:
"""Lowest VRAM ComfyUI has been observed holding while alive.
A purge frees checkpoints but not the CUDA context, so ComfyUI keeps a few hundred
MB for as long as the process runs. The minimum seen in recent telemetry is a better
estimate of that floor than whatever it happens to hold right now, which could be a
7 GB checkpoint mid-generation.
"""
try:
rows = telemetry_store._rows(
"SELECT MIN(comfy_bytes) AS floor FROM telemetry "
"WHERE ts > ? AND comfy_bytes > 0",
(time.time() - days * 86400,))
if rows and rows[0].get("floor"):
return round(rows[0]["floor"] / (1024 ** 3), 2)
except Exception as e:
logger.debug(f"comfy floor lookup failed: {e}")
return default
def _check_unmanaged_vram() -> Dict[str, Any]:
"""Report unreclaimable VRAM in terms of what it actually costs.
"0.82 GB unmanaged" is a number. "0.82 GB unmanaged, which is why three of your
models can no longer fit" is something you can act on.
"""
stats = vram_arbitrator.get_gpu_hardware_stats()
if not stats.get("available"):
return _check("unmanaged VRAM", DEGRADED, "GPU unavailable")
bd = stats.get("breakdown", {})
unmanaged_gb = bd.get("unmanaged_gb", 0.0)
procs = bd.get("unmanaged", [])
if not procs:
return _check("unmanaged VRAM", OK, "no third-party GPU processes")
total_gb = stats.get("vram_total_gb", 0)
# What HyperSwap could offer at best. Three things are never available: the desktop,
# processes it cannot touch, and ComfyUI's own CUDA context, which survives a purge.
# Omitting that last one made this check claim a 14.93 GB model would fit against a
# real ceiling of 14.60 GB -- the model that had just returned 507.
comfy_floor_gb = _comfy_vram_floor_gb(default=bd.get("comfyui_gb", 0.0))
ceiling_gb = total_gb - unmanaged_gb - bd.get("desktop_gb", 0.0) - comfy_floor_gb
try:
blobs = ram_optimizer.find_ollama_model_files()
except Exception:
blobs = []
# Measured on this box: a 12.87 GB blob occupies 14.9 GB once context and KV cache
# are allocated.
VRAM_OVERHEAD = 1.16
blocked = sorted(
{b["model"]: b for b in blobs
if b["size_gb"] * VRAM_OVERHEAD > ceiling_gb
and b["size_gb"] * VRAM_OVERHEAD <= ceiling_gb + unmanaged_gb}.values(),
key=lambda b: -b["size_gb"])
names = ", ".join(b["model"] for b in blocked[:3])
who = ", ".join(f"{p['name']} ({p['vram_mb']} MB)" for p in procs[:2])
if blocked:
return _check("unmanaged VRAM", DEGRADED,
f"{unmanaged_gb} GB held by {who}",
f"{len(blocked)} model(s) fit within {ceiling_gb + unmanaged_gb:.2f} GB "
f"but not the {ceiling_gb:.2f} GB actually available: {names}",
"Stop that process to reclaim the difference, or accept that "
"these models cannot load")
return _check("unmanaged VRAM", OK,
f"{unmanaged_gb} GB held by {who}; no model is blocked by it",
"", "")
def _check_model_dirs() -> Dict[str, Any]:
comfy_dir = ram_optimizer.COMFY_MODELS_DIR
if not os.path.isdir(comfy_dir):
@@ -158,7 +247,7 @@ async def run_health_checks() -> Dict[str, Any]:
sync_checks = [_check_nvml, _check_sudo_smi, _check_fan_control,
_check_profile_drift, _check_store, _check_residency,
_check_model_dirs, _check_comfy_ws]
_check_model_dirs, _check_comfy_ws, _check_unmanaged_vram]
results: List[Dict[str, Any]] = []
for fn in sync_checks:
try:

445
jobs.py Normal file
View File

@@ -0,0 +1,445 @@
"""A durable, unbounded job queue across every GPU tenant.
Until now this service only reacted: it noticed an application had started working and
scrambled to free memory. Nothing could be *lined up*. Each application has its own queue
(ComfyUI's prompt queue, Ollama's serialised requests), but they cannot see each other, so
work submitted to one has no way to wait politely for the other.
Jobs submitted here are stored in SQLite, so the queue is limited by disk rather than
memory and survives a restart. The scheduler takes the highest-priority pending job,
makes sure its tenant actually has the VRAM to run it -- reusing the same plan_release
arbitration -- dispatches it, and moves on.
"""
import asyncio
import json
import logging
import sqlite3
import time
import uuid
from typing import Any, Dict, List, Optional
import httpx
import telemetry_store
import tenants as tenants_mod
logger = logging.getLogger("jobs")
PENDING, RUNNING, DONE, FAILED, CANCELLED = (
"pending", "running", "done", "failed", "cancelled")
SCHEMA = """
CREATE TABLE IF NOT EXISTS jobs (
id TEXT PRIMARY KEY,
tenant TEXT NOT NULL,
priority INTEGER NOT NULL DEFAULT 50,
payload TEXT NOT NULL,
state TEXT NOT NULL,
submitted_at REAL NOT NULL,
started_at REAL,
finished_at REAL,
error TEXT,
result TEXT,
label TEXT
);
CREATE INDEX IF NOT EXISTS idx_jobs_state ON jobs(state, priority DESC, submitted_at);
"""
def _conn() -> sqlite3.Connection:
c = sqlite3.connect(telemetry_store.DB_PATH, timeout=10.0)
c.row_factory = sqlite3.Row
return c
def init() -> None:
with _conn() as c:
c.executescript(SCHEMA)
def submit(tenant: str, payload: Dict[str, Any], priority: Optional[int] = None,
label: Optional[str] = None) -> Dict[str, Any]:
"""Queue a job. There is no depth limit: the queue lives on disk."""
t = tenants_mod.get_tenant(tenant)
if not t:
return {"success": False, "error": f"no tenant named '{tenant}'"}
job_id = uuid.uuid4().hex[:12]
row = {
"id": job_id, "tenant": tenant,
"priority": t.priority if priority is None else int(priority),
"payload": json.dumps(payload), "state": PENDING,
"submitted_at": time.time(), "label": label,
}
with _conn() as c:
c.execute("INSERT INTO jobs (id, tenant, priority, payload, state, submitted_at,"
" label) VALUES (:id,:tenant,:priority,:payload,:state,:submitted_at,"
":label)", row)
logger.info(f"queued job {job_id} for '{tenant}' at priority {row['priority']}")
return {"success": True, "id": job_id, "tenant": tenant,
"priority": row["priority"], "state": PENDING}
def cancel(job_id: str) -> Dict[str, Any]:
with _conn() as c:
cur = c.execute("UPDATE jobs SET state=?, finished_at=? WHERE id=? AND state=?",
(CANCELLED, time.time(), job_id, PENDING))
if cur.rowcount:
return {"success": True, "id": job_id, "state": CANCELLED}
return {"success": False, "error": "job is not pending (already running or finished)"}
def clear_pending() -> Dict[str, Any]:
with _conn() as c:
cur = c.execute("UPDATE jobs SET state=?, finished_at=? WHERE state=?",
(CANCELLED, time.time(), PENDING))
return {"success": True, "cancelled": cur.rowcount}
def get(job_id: str) -> Optional[Dict[str, Any]]:
with _conn() as c:
r = c.execute("SELECT * FROM jobs WHERE id=?", (job_id,)).fetchone()
return _row(r) if r else None
def _row(r: sqlite3.Row) -> Dict[str, Any]:
d = dict(r)
for key in ("payload", "result"):
if d.get(key):
try:
d[key] = json.loads(d[key])
except Exception:
pass
if d.get("started_at") and d.get("finished_at"):
d["duration_s"] = round(d["finished_at"] - d["started_at"], 2)
if d.get("state") == PENDING:
d["waiting_s"] = round(time.time() - d["submitted_at"], 1)
return d
def listing(state: Optional[str] = None, limit: int = 100) -> List[Dict[str, Any]]:
q = "SELECT * FROM jobs"
args: List[Any] = []
if state:
q += " WHERE state=?"
args.append(state)
# Pending jobs in the order the scheduler will take them; everything else newest first.
q += (" ORDER BY priority DESC, submitted_at ASC" if state == PENDING
else " ORDER BY submitted_at DESC")
q += " LIMIT ?"
args.append(limit)
with _conn() as c:
return [_row(r) for r in c.execute(q, args).fetchall()]
def stats() -> Dict[str, Any]:
with _conn() as c:
rows = c.execute("SELECT state, COUNT(*) n FROM jobs GROUP BY state").fetchall()
by_state = {r["state"]: r["n"] for r in rows}
pend = c.execute(
"SELECT tenant, COUNT(*) n FROM jobs WHERE state=? GROUP BY tenant",
(PENDING,)).fetchall()
oldest = c.execute(
"SELECT MIN(submitted_at) t FROM jobs WHERE state=?", (PENDING,)).fetchone()
return {
"by_state": by_state,
"pending_by_tenant": {r["tenant"]: r["n"] for r in pend},
"queue_depth": by_state.get(PENDING, 0),
"oldest_pending_s": (round(time.time() - oldest["t"], 1)
if oldest and oldest["t"] else None),
}
def requeue_orphans() -> int:
"""Return jobs abandoned mid-run to the queue.
RUNNING means "this process is working on it". If no process is, that is untrue, and
the job would otherwise never finish and never retry.
"""
with _conn() as c:
cur = c.execute("UPDATE jobs SET state=?, started_at=NULL WHERE state=?",
(PENDING, RUNNING))
return cur.rowcount
def _next_job() -> Optional[Dict[str, Any]]:
with _conn() as c:
r = c.execute(
"SELECT * FROM jobs WHERE state=? ORDER BY priority DESC, submitted_at ASC"
" LIMIT 1", (PENDING,)).fetchone()
return _row(r) if r else None
def _mark(job_id: str, state: str, **fields) -> None:
sets = ", ".join(f"{k}=?" for k in fields)
args = list(fields.values()) + [state, job_id]
with _conn() as c:
c.execute(f"UPDATE jobs SET {sets + ', ' if sets else ''}state=? WHERE id=?", args)
# ---------------------------------------------------------------- dispatch
async def _dispatch_comfy(payload: Dict[str, Any]) -> Dict[str, Any]:
"""Hand a workflow to ComfyUI and wait for it to finish."""
base = tenants_mod.get_tenant("comfyui").busy.url.rsplit("/", 1)[0]
async with httpx.AsyncClient(timeout=30.0) as c:
r = await c.post(f"{base}/prompt", json={"prompt": payload.get("prompt", payload),
"client_id": "hyperswap-jobs"})
if r.status_code != 200:
return {"ok": False, "error": f"HTTP {r.status_code}: {r.text[:200]}"}
prompt_id = r.json().get("prompt_id")
deadline = time.time() + payload.get("timeout_s", 1800)
while time.time() < deadline:
await asyncio.sleep(0.5)
h = await c.get(f"{base}/history/{prompt_id}")
entry = (h.json() or {}).get(prompt_id) if h.status_code == 200 else None
if not entry:
continue
status = entry.get("status", {})
if status.get("status_str") == "error":
return {"ok": False, "error": "ComfyUI reported an execution error"}
if status.get("completed"):
return {"ok": True, "prompt_id": prompt_id}
return {"ok": False, "error": "timed out waiting for ComfyUI"}
async def _dispatch_ollama(payload: Dict[str, Any]) -> Dict[str, Any]:
async with httpx.AsyncClient(timeout=payload.get("timeout_s", 1800)) as c:
body = {"stream": False, "keep_alive": payload.get("keep_alive", "5m"), **payload}
body.pop("timeout_s", None)
r = await c.post("http://localhost:11434/api/generate", json=body)
if r.status_code != 200:
return {"ok": False, "error": f"HTTP {r.status_code}: {r.text[:200]}"}
d = r.json()
return {"ok": True, "response": (d.get("response") or "")[:2000],
"eval_count": d.get("eval_count"),
"tokens_per_sec": (round(d.get("eval_count", 0)
/ (d.get("eval_duration", 1) / 1e9), 2)
if d.get("eval_duration") else None)}
DISPATCHERS = {
tenants_mod.KIND_DIFFUSION: _dispatch_comfy,
tenants_mod.KIND_LLM: _dispatch_ollama,
}
class Scheduler:
"""Drains the queue, making room for each job before it runs.
One job at a time by design. The GPU is the scarce resource this whole service
exists to hand between applications; running two jobs concurrently would just
recreate the contention it is meant to resolve. Throughput comes from swapping
quickly, not from overlapping.
"""
def __init__(self) -> None:
self.running = False
self.task: Optional[asyncio.Task] = None
self.current: Optional[Dict[str, Any]] = None
self.last_finished: Optional[Dict[str, Any]] = None
self.completed = 0
self.failed = 0
self.waits = 0
self.blocked: Optional[Dict[str, Any]] = None
self.idle_poll_s = 1.0
self.blocked_poll_s = 2.0
# How long a job may wait for VRAM before it is declared impossible. Long enough
# to outlast a normal diffusion run, short enough not to wedge the queue.
self.max_block_s = 120.0
async def start(self) -> None:
if self.running:
return
init()
# A job left RUNNING by a crash or a hard restart would sit there forever.
requeued = requeue_orphans()
if requeued:
logger.warning(f"requeued {requeued} job(s) left running by a previous process")
self.running = True
self.task = asyncio.create_task(self._loop())
logger.info("job scheduler started")
async def stop(self) -> None:
self.running = False
if self.task:
self.task.cancel()
def _job_vram_requirement(self, tenant_name: str, payload: Dict[str, Any]) -> float:
"""How much VRAM *this* job needs, not the tenant's generic figure.
A tenant-wide needs_vram_gb cannot be right for an LLM: the requirement is a
property of the model being loaded. Ollama's generic 4 GB passed the room check
with 8 GB free, and then a 14.9 GB model was dispatched into it and killed
llama-server with a CUDA OOM -- three queued jobs destroyed in a row.
"""
import vram_arbitrator
t = tenants_mod.get_tenant(tenant_name)
default = t.needs_vram_gb if t else 0.0
model = payload.get("model")
if t and t.kind == tenants_mod.KIND_LLM and model:
size = vram_arbitrator._model_size_bytes(model)
if size:
# Measured on this box: a 12.87 GB blob occupies 14.9 GB once context
# and KV cache are allocated.
return round((size / (1024 ** 3)) * 1.16, 2)
return default
async def _make_room(self, tenant_name: str,
payload: Optional[Dict[str, Any]] = None) -> Dict[str, Any]:
"""Ensure the job's tenant has the VRAM it needs, using the normal arbitration."""
import vram_arbitrator # imported late: it imports this module's siblings
t = tenants_mod.get_tenant(tenant_name)
needed = self._job_vram_requirement(tenant_name, payload or {})
if not t or not needed:
return {"ready": True, "reason": "no VRAM requirement declared"}
state = await vram_arbitrator.arbitrator._tenant_state()
free_gb = vram_arbitrator.arbitrator._last_tenant_state["free_gb"]
held = next((s["vram_gb"] for s in state if s["name"] == tenant_name), 0.0)
if held + free_gb >= needed:
return {"ready": True, "needed_gb": needed,
"reason": f"{free_gb:.2f} GB free, job needs {needed:.2f} GB"}
plan = tenants_mod.plan_release(tenant_name, state, free_gb, needed)
for victim in plan["release"]:
await vram_arbitrator.arbitrator._release_tenant(
victim, f"queued job for '{tenant_name}'")
if plan["release"]:
# Give the driver a moment to actually hand the memory back.
deadline = time.perf_counter() + 30
target = int(needed * (1024 ** 3))
while time.perf_counter() < deadline:
if vram_arbitrator.get_process_vram_bytes()["free_bytes"] >= target:
break
await asyncio.sleep(0.05)
# Only ready once the memory is genuinely there. A plan that *could* work is not
# the same as VRAM that *is* free, and dispatching on the former is what OOMs.
free_now = (vram_arbitrator.get_process_vram_bytes()["free_bytes"] / (1024 ** 3))
# The best this GPU could ever offer this tenant: everything currently free, plus
# what it already holds, plus everything that is reclaimable at all.
reclaimable_gb = sum(s["vram_gb"] for s in state
if s["name"] != tenant_name and s.get("reclaimable"))
max_possible = round(held + free_now + reclaimable_gb, 2)
return {"ready": plan["possible"] and (held + free_now) >= needed,
"max_possible_gb": max_possible,
"released": plan["release"], "needed_gb": needed,
"free_gb": round(free_now, 2),
"reason": f"job needs {needed:.2f} GB; {plan['reason']}",
"blockers": plan.get("blockers")}
async def _loop(self) -> None:
while self.running:
try:
job = _next_job()
if not job:
await asyncio.sleep(self.idle_poll_s)
continue
tenant = tenants_mod.get_tenant(job["tenant"])
dispatcher = DISPATCHERS.get(tenant.kind) if tenant else None
if not dispatcher:
_mark(job["id"], FAILED, finished_at=time.time(),
error=f"no dispatcher for tenant kind "
f"'{tenant.kind if tenant else '?'}'")
self.failed += 1
continue
# Check for room *before* claiming the job. Dispatching into
# insufficient VRAM does not fail gracefully -- it kills llama-server
# with a CUDA OOM, which is how three queued LLM jobs were destroyed
# while ComfyUI legitimately held the card. A job that cannot run yet
# waits; it does not fail.
room = await self._make_room(job["tenant"], job.get("payload") or {})
if not room.get("ready"):
since = (self.blocked.get("since", time.time())
if self.blocked and self.blocked.get("id") == job["id"]
else time.time())
waited = time.time() - since
self.blocked = {"id": job["id"], "tenant": job["tenant"],
"reason": room.get("reason"),
"blockers": room.get("blockers"),
"needed_gb": room.get("needed_gb"),
"since": since, "waited_s": round(waited, 1)}
# Waiting is right while the memory might still arrive. It is wrong
# when the job can never fit -- three LLM jobs sat pending forever
# needing 14.93 GB on a card where only ~14.8 GB can ever be free,
# because an unreclaimable process holds 0.82 GB. Say so and move on
# rather than blocking the queue behind an impossibility.
if waited > self.max_block_s:
ceiling = room.get("max_possible_gb")
detail = (f"needs {room.get('needed_gb')} GB but at most "
f"{ceiling} GB can ever be free on this GPU"
if ceiling is not None and room.get("needed_gb", 0) > ceiling
else f"waited {int(waited)}s for VRAM: {room.get('reason')}")
blockers = ", ".join(
f"{b['name']} ({b['vram_gb']} GB, {b['why']})"
for b in (room.get("blockers") or []))
_mark(job["id"], FAILED, finished_at=time.time(),
error=f"{detail}{'; blocked by ' + blockers if blockers else ''}")
self.failed += 1
self.blocked = None
logger.warning(f"job {job['id']} cannot run: {detail}")
continue
self.waits += 1
await asyncio.sleep(self.blocked_poll_s)
continue
self.blocked = None
t0 = time.time()
_mark(job["id"], RUNNING, started_at=t0)
self.current = {**job, "state": RUNNING, "room": room,
"started_at": t0}
logger.info(f"running job {job['id']} for '{job['tenant']}' "
f"({room.get('reason')})")
# Any failure here must land on the job. An exception used to escape to
# the loop's handler, leaving the row RUNNING forever while the scheduler
# moved on -- an orphan that never completed and never freed its slot.
try:
res = await dispatcher(job["payload"])
except asyncio.CancelledError:
_mark(job["id"], PENDING, started_at=None)
self.current = None
raise
except Exception as e:
res = {"ok": False, "error": f"dispatch raised: {e}"}
finished = time.time()
if res.get("ok"):
_mark(job["id"], DONE, finished_at=finished,
result=json.dumps(res))
self.completed += 1
else:
_mark(job["id"], FAILED, finished_at=finished,
error=str(res.get("error"))[:500])
self.failed += 1
self.last_finished = {"id": job["id"], "tenant": job["tenant"],
"ok": bool(res.get("ok")),
"duration_s": round(finished - t0, 2),
"made_room": room.get("released") or []}
self.current = None
except asyncio.CancelledError:
raise
except Exception as e:
logger.error(f"scheduler error: {e}")
self.current = None
await asyncio.sleep(1.0)
def get_status(self) -> Dict[str, Any]:
return {
"running": self.running,
"current": self.current,
"last_finished": self.last_finished,
"completed": self.completed,
"failed": self.failed,
"waits": self.waits,
"blocked": self.blocked,
**stats(),
}
scheduler = Scheduler()

View File

@@ -8,7 +8,9 @@ from typing import Dict, List, Any, Optional
from mcp.server import MCPServer
import autotune
import engines
import health
import jobs as jobs_mod
import overclock_manager
import ram_optimizer
import telemetry_store
@@ -137,6 +139,41 @@ def set_gpu_fan_speed(mode: str = "auto", percent: Optional[int] = None) -> str:
res = overclock_manager.set_fan_auto()
return json.dumps(res, indent=2)
@mcp.tool()
async def get_engine_config() -> str:
"""Live configuration of Ollama and ComfyUI (parallelism, max loaded models,
keep-alive, KV cache type, ComfyUI VRAM mode and allocator), with what each setting
implies for VRAM arbitration."""
return json.dumps(await engines.get_engine_config(), indent=2, default=str)
@mcp.tool()
def queue_job(tenant: str, payload: Dict[str, Any], priority: Optional[int] = None,
label: Optional[str] = None) -> str:
"""Queue work for a GPU application without waiting for it.
tenant: 'ollama' (payload: model, prompt, options) or 'comfyui' (payload: {"prompt":
<workflow>}). The queue is on disk, so there is no depth limit; jobs run one at a
time, highest priority first, with VRAM arbitrated before each starts."""
return json.dumps(jobs_mod.submit(tenant, payload, priority, label), indent=2,
default=str)
@mcp.tool()
def get_job_queue(state: Optional[str] = None, limit: int = 50) -> str:
"""Queued and recent jobs, plus what the scheduler is doing and why it may be
waiting. Pending jobs are listed in the order they will run."""
return json.dumps({"jobs": jobs_mod.listing(state, limit),
"scheduler": jobs_mod.scheduler.get_status()},
indent=2, default=str)
@mcp.tool()
def cancel_job(job_id: str) -> str:
"""Cancel a job that has not started. Running work is never killed."""
return json.dumps(jobs_mod.cancel(job_id), indent=2, default=str)
@mcp.tool()
async def check_system_health() -> str:
"""Check every dependency HyperSwap needs (NVML, sudo nvidia-smi, fan control via the

View File

@@ -1,7 +1,6 @@
{
"ollama": {
"label": "Ollama — LLM decode (memory-bandwidth bound; measured insensitive to power and clocks)",
"measured": "73.0-73.5 tok/s flat from 222W to 370W (qwen3.8long, 2026-08-28). Actual draw never exceeded 224W at any limit. Memory clock lock made no difference (72.6 locked vs 72.7 unlocked).",
"label": "Ollama \u2014 LLM decode (memory-bandwidth bound; measured insensitive to power and clocks)",
"power_limit_w": 320,
"core_offset_mhz": 0,
"mem_offset_mhz": 0,
@@ -9,11 +8,11 @@
"lock_core_max": 0,
"lock_mem_mhz": 0,
"fan_mode": "auto",
"fan_speed_pct": 0
"fan_speed_pct": 0,
"measured": "73.0-73.5 tok/s flat from 222W to 370W (qwen3.8long, 2026-08-28). Actual draw never exceeded 224W at any limit. Memory clock lock made no difference (72.6 locked vs 72.7 unlocked)."
},
"comfy": {
"label": "ComfyUI — diffusion (compute bound; genuinely power-scaling)",
"measured": "SDXL 1024/20-step: 5.48 it/s @222W, 6.22 @259W, 6.50 @296W, 6.52 @320W, 6.63 @333W, 6.71 @370W (2026-08-28). Worth +2.8% over the 320W stock default. Core clock lock made no difference across 2400-3105 MHz.",
"label": "ComfyUI \u2014 diffusion (compute bound; genuinely power-scaling)",
"power_limit_w": 370,
"core_offset_mhz": 0,
"mem_offset_mhz": 0,
@@ -21,18 +20,19 @@
"lock_core_max": 0,
"lock_mem_mhz": 0,
"fan_mode": "auto",
"fan_speed_pct": 0
"fan_speed_pct": 0,
"measured": "SDXL 1024/20-step: 5.48 it/s @222W, 6.22 @259W, 6.50 @296W, 6.52 @320W, 6.63 @333W, 6.71 @370W (2026-08-28). Worth +2.8% over the 320W stock default. Core clock lock made no difference across 2400-3105 MHz."
},
"balanced": {
"label": "Balanced — stock power and boost, automatic fans",
"measured": "Card's own design point. 48k telemetry samples show 67.8C average under load at 39.5% auto fan, 81C all-time max, zero thermal throttle events.",
"power_limit_w": 320,
"core_offset_mhz": 0,
"mem_offset_mhz": 0,
"label": "Balanced \u2014 stock power and boost, automatic fans",
"power_limit_w": 340,
"core_offset_mhz": 10,
"mem_offset_mhz": 150,
"lock_core_min": 0,
"lock_core_max": 0,
"lock_mem_mhz": 0,
"fan_mode": "auto",
"fan_speed_pct": 0
"fan_mode": "manual",
"fan_speed_pct": 95,
"measured": "Card's own design point. 48k telemetry samples show 67.8C average under load at 39.5% auto fan, 81C all-time max, zero thermal throttle events."
}
}

170
server.py
View File

@@ -16,10 +16,13 @@ from fastapi.middleware.cors import CORSMiddleware
from pydantic import BaseModel, Field
import autotune
import engines
import health
import jobs as jobs_mod
import overclock_manager
import ram_optimizer
import telemetry_store
import tenants as tenants_mod
import thermal_governor
import vram_arbitrator
@@ -218,10 +221,13 @@ def _install_shutdown_hook() -> None:
@asynccontextmanager
async def lifespan(app: FastAPI):
telemetry_store.start()
jobs_mod.init()
await broker.start()
await jobs_mod.scheduler.start()
await vram_arbitrator.arbitrator.start()
_install_shutdown_hook()
yield
await jobs_mod.scheduler.stop()
await vram_arbitrator.arbitrator.stop()
await broker.stop()
# Never leave the card with locked clocks and pinned fans after we exit.
@@ -321,6 +327,141 @@ async def api_health():
return await health.run_health_checks()
class TenantReleaseRequest(BaseModel):
models: Optional[List[str]] = Field(None, description="For per-model tenants (Ollama), which to unload; defaults to everything resident")
confirm: bool = Field(True, description="Wait for NVML to confirm the VRAM was actually released")
class JobRequest(BaseModel):
tenant: str = Field(..., description="Which application should run this job", example="comfyui")
payload: Dict[str, Any] = Field(..., description="What to run: a ComfyUI workflow under 'prompt', or Ollama generate parameters")
priority: Optional[int] = Field(None, description="Defaults to the tenant's priority; higher runs sooner")
label: Optional[str] = Field(None, description="Human-readable name for the queue view")
@app.post("/api/jobs", summary="Queue a Job", tags=["Jobs"])
async def api_submit_job(req: JobRequest):
"""Queue work for any tenant. The queue is on disk, so there is no depth limit and
it survives a restart. Jobs run one at a time, highest priority first, with VRAM
arbitrated before each one starts."""
res = jobs_mod.submit(req.tenant, req.payload, req.priority, req.label)
if not res.get("success"):
raise HTTPException(status_code=400, detail=res.get("error"))
return res
@app.get("/api/jobs", summary="The Job Queue", tags=["Jobs"])
async def api_jobs(state: Optional[str] = Query(None, description="pending | running | done | failed | cancelled"),
limit: int = Query(100)):
"""Queued and recent jobs. Pending jobs are listed in the order they will run."""
return {"jobs": jobs_mod.listing(state, limit), "scheduler": jobs_mod.scheduler.get_status()}
@app.get("/api/jobs/{job_id}", summary="One Job", tags=["Jobs"])
async def api_job(job_id: str):
job = jobs_mod.get(job_id)
if not job:
raise HTTPException(status_code=404, detail=f"no job '{job_id}'")
return job
@app.delete("/api/jobs/{job_id}", summary="Cancel a Pending Job", tags=["Jobs"])
async def api_cancel_job(job_id: str):
"""Cancel a job that has not started. Running jobs are left alone -- this service
frees VRAM by asking, never by killing work in flight."""
res = jobs_mod.cancel(job_id)
if not res.get("success"):
raise HTTPException(status_code=409, detail=res.get("error"))
return res
@app.delete("/api/jobs", summary="Cancel All Pending Jobs", tags=["Jobs"])
async def api_clear_jobs():
return jobs_mod.clear_pending()
@app.get("/api/tenants", summary="GPU Tenants", tags=["Tenants"])
async def api_tenants():
"""Applications competing for the GPU, as configured.
Each entry declares how its processes are recognised, how to tell whether it is
working, and how to ask it for VRAM back. Adding an application is a config change
in tenants.json, not a code change.
"""
gpu = vram_arbitrator.get_gpu_hardware_stats()
by_tenant = (gpu.get("breakdown", {}) or {}).get("by_tenant_gb", {})
out = []
for t in tenants_mod.describe():
name = t["name"]
# The two original tenants are reported under the bucket names the API has
# always used.
bucket = {"comfyui": "comfy"}.get(name, name)
t["vram_gb"] = by_tenant.get(bucket, 0.0)
out.append(t)
return {"tenants": out, "config_path": tenants_mod.CONFIG_PATH,
"unmanaged_gb": (gpu.get("breakdown", {}) or {}).get("unmanaged_gb", 0.0)}
@app.get("/api/tenants/{name}", summary="One GPU Tenant", tags=["Tenants"])
async def api_tenant(name: str):
"""A single tenant's definition, current VRAM, and whether it is genuinely busy."""
t = tenants_mod.get_tenant(name)
if not t:
raise HTTPException(status_code=404, detail=f"no tenant named '{name}'")
gpu = vram_arbitrator.get_gpu_hardware_stats()
bucket = {"comfyui": "comfy"}.get(name, name)
vram_gb = (gpu.get("breakdown", {}) or {}).get("by_tenant_gb", {}).get(bucket, 0.0)
busy = await tenants_mod.probe_busy(t, vram_gb=vram_gb)
d = t.to_dict()
d.update({"vram_gb": vram_gb, "reclaimable": t.reclaimable, "busy": busy})
return d
@app.post("/api/tenants/{name}/release", summary="Ask a Tenant for its VRAM", tags=["Tenants"])
async def api_tenant_release(name: str, req: Optional[TenantReleaseRequest] = None):
"""Release a tenant's VRAM using whatever mechanism that tenant declares.
This is the generic form of the Ollama soft-yield and the ComfyUI purge: the same
request works for any application in the registry, including ones added later.
"""
t = tenants_mod.get_tenant(name)
if not t:
raise HTTPException(status_code=404, detail=f"no tenant named '{name}'")
if not t.reclaimable:
raise HTTPException(status_code=409,
detail=f"'{name}' declares no way to release VRAM; its "
f"memory cannot be reclaimed by this service")
models = req.models if req else None
if t.release.per_model and not models:
state = await vram_arbitrator.get_ollama_live_state()
models = [m.get("name") for m in state.get("loaded_models", []) if m.get("name")]
before = vram_arbitrator.get_process_vram_bytes()
res = await tenants_mod.release_vram(t, models=models)
if (req is None or req.confirm) and res.get("released"):
bucket = {"comfyui": "comfy"}.get(name, name)
key = {"ollama": "ollama_bytes", "comfy": "comfyui_bytes"}.get(bucket)
if key:
baseline = before[key]
barrier = await vram_arbitrator._await_vram_release(baseline) \
if key == "ollama_bytes" else None
if barrier:
res.update({"outcome": barrier.get("outcome"),
"confirm_ms": barrier.get("confirm_ms")})
after = vram_arbitrator.get_process_vram_bytes()
res["free_vram_gb"] = round(after["free_bytes"] / (1024**3), 2)
return res
@app.get("/api/engines", summary="Live Engine Configuration", tags=["Telemetry"])
async def api_engines():
"""Real configuration of Ollama and ComfyUI, with what each setting implies for
arbitration. These live outside this codebase but dictate how it must behave."""
return await engines.get_engine_config()
@app.get("/api/gpu", summary="GPU Sensors and VRAM Breakdown", tags=["Telemetry"])
async def get_gpu_metrics() -> Dict[str, Any]:
"""Detailed NVML sensors (utilization, temp, power, fan, clocks, throttle reasons, per-process VRAM)."""
@@ -375,7 +516,18 @@ async def api_switch_model(req: SwitchRequest):
await vram_arbitrator.arbitrator.request_vram_for_ollama()
res = await vram_arbitrator.switch_ollama_model(req.model, keep_alive=req.keep_alive or "30m")
if not res.get("success"):
raise HTTPException(status_code=500, detail=res.get("error"))
# Reflect what actually went wrong. Ollama returns 400 for an unusable request --
# asking an embedding model to generate, say -- and reporting that as 500 blames
# this service for the caller's mistake. A model that will not fit is neither:
# the request is valid and the service is healthy, there is simply no room.
upstream = res.get("upstream_status")
if res.get("vram_oom"):
status = 507 # Insufficient Storage
elif isinstance(upstream, int) and 400 <= upstream < 500:
status = 400
else:
status = 502 if upstream else 500
raise HTTPException(status_code=status, detail=res.get("error"))
return res
@app.post("/api/free-vram", summary="Soft-Yield Ollama VRAM", tags=["Orchestration"])
@@ -595,9 +747,23 @@ app.mount("/static", StaticFiles(directory=f"{BASE_DIR}/static"), name="static")
@app.get("/", summary="Dashboard Web UI", tags=["UI"])
async def root_index():
"""Serve the dashboard with cache-busted asset URLs.
StaticFiles sends an ETag, but browsers were still serving app.js from cache after
it changed, so a reload showed the old dashboard against the new API -- a panel that
had just been added simply never appeared. Stamping each asset with its mtime means
a changed file is always a different URL.
"""
with open(f"{BASE_DIR}/static/index.html", "r") as f:
content = f.read()
return HTMLResponse(content=content)
for asset in ("app.js", "styles.css"):
try:
stamp = int(os.path.getmtime(f"{BASE_DIR}/static/{asset}"))
except OSError:
continue
content = content.replace(f"/static/{asset}", f"/static/{asset}?v={stamp}")
return HTMLResponse(content=content,
headers={"Cache-Control": "no-cache, must-revalidate"})
if __name__ == "__main__":
import uvicorn

View File

@@ -36,6 +36,7 @@ function updateDashboard(data) {
// Governor and arbitration state ride along in the shared snapshot.
if (data.governor) renderGovernor(data.governor);
if (data.arbitrator) renderArbitrator(data.arbitrator, data.gpu);
if (data.arbitrator) renderTenants(data.arbitrator, data.gpu);
// 1. GPU VRAM Stats
const gpu = data.gpu || {};
@@ -1069,3 +1070,193 @@ async function fetchLastSwapFromStore() {
document.addEventListener('DOMContentLoaded', () => {
setTimeout(fetchLastSwapFromStore, 1500);
});
// ---------------------------------------------------------------- engine config
async function fetchEngineConfig() {
// These subtitles used to be hardcoded. They happened to be accurate, which is worse
// than being wrong: they would have stayed accurate-looking after the settings changed.
try {
const d = await (await fetch('/api/engines')).json();
const o = document.getElementById('ollama-engine-sub');
if (o && d.ollama) {
const bits = [`Port :${d.ollama.port}`, d.ollama.summary];
if (d.ollama.max_loaded_models) bits.push(`${d.ollama.max_loaded_models} model resident`);
if (d.ollama.keep_alive) bits.push(`keep-alive ${d.ollama.keep_alive}`);
o.textContent = bits.join(' // ');
o.title = (d.ollama.settings || [])
.filter(s => s.means)
.map(s => `${s.key}=${s.value} — ${s.means}`)
.join('\n');
}
const c = document.getElementById('comfy-engine-sub');
if (c && d.comfyui && d.comfyui.online) {
c.textContent = `Port :8188 // v${d.comfyui.version} // ${d.comfyui.summary}`;
c.title = `torch ${d.comfyui.pytorch}\n${d.comfyui.device || ''}`;
}
} catch (e) { /* subtitles are cosmetic; never break the page over them */ }
}
document.addEventListener('DOMContentLoaded', () => {
fetchEngineConfig();
setInterval(fetchEngineConfig, 120000);
});
// ---------------------------------------------------------------- chart sizing
// Chart.js is configured responsive with maintainAspectRatio:false, so it should track
// its container on its own. In practice it latched onto a stale size -- the canvas sat
// at width:0px, height:288px (the old fixed h-64) while its container had grown to
// 988x648 on a large window. An explicit observer makes the chart follow the container
// whatever the window does.
function watchChartSize() {
const canvas = document.getElementById('oc-chart');
if (!canvas || !canvas.parentElement) return;
const container = canvas.parentElement;
const resize = () => {
if (typeof ocChart === 'undefined' || !ocChart) return;
// No explicit dimensions: with responsive + maintainAspectRatio:false, Chart.js
// measures the container itself. Passing width/height instead made the canvas grow
// but never shrink -- it ended up 988px wide inside a 435px container, overflowing
// it, which is also why the canvas is absolutely positioned now.
ocChart.resize();
};
if (typeof ResizeObserver !== 'undefined') {
new ResizeObserver(resize).observe(container);
}
window.addEventListener('resize', resize);
// Run once after layout settles, in case the chart was built before the container had
// a width (which is how it ended up at 0 in the first place).
requestAnimationFrame(resize);
setTimeout(resize, 300);
}
document.addEventListener('DOMContentLoaded', () => setTimeout(watchChartSize, 200));
// ---------------------------------------------------------------- tenants
function renderTenants(arb, gpu) {
const body = document.getElementById('tenants-body');
if (!body) return;
const state = arb && arb.tenant_state;
if (!state || !state.tenants) return;
const total = (gpu && gpu.vram_total_gb) || 16;
document.getElementById('tenants-free').textContent = `${state.free_gb} GB free`;
// Sorted by priority, the order arbitration actually considers them in.
const rows = [...state.tenants].sort((a, b) => b.priority - a.priority);
body.innerHTML = rows.map(t => {
const pct = Math.min((t.vram_gb / total) * 100, 100);
const bar = t.busy ? 'bg-emerald-500'
: t.reclaimable ? 'bg-cyan-600' : 'bg-amber-600';
const badge = t.busy
? '<span class="text-emerald-400">working</span>'
: t.reclaimable
? '<span class="text-slate-500">idle · reclaimable</span>'
: '<span class="text-amber-400">cannot be reclaimed</span>';
return `<div>
<div class="flex justify-between text-[11px] font-mono">
<span class="text-slate-200">${t.name}
<span class="text-slate-600">p${t.priority}</span></span>
<span class="text-slate-400">${t.vram_gb.toFixed(2)} GB · ${badge}</span>
</div>
<div class="w-full bg-slate-950 rounded-full h-1.5 mt-1 overflow-hidden border border-slate-800/60">
<div class="${bar} h-full transition-all duration-500" style="width:${pct}%"></div>
</div>
<div class="text-[10px] text-slate-600 mt-0.5">${t.reason || ''}${
t.needs_vram_gb ? ` · needs ${t.needs_vram_gb} GB to work` : ''}</div>
</div>`;
}).join('');
// The most recent arbitration decision, including why it could not be satisfied.
const dec = document.getElementById('tenants-decision');
const a = arb.last_arbitration;
if (!a) {
dec.innerHTML = '<span class="text-slate-600">No contention — nothing has needed to be released.</span>';
return;
}
const when = new Date(a.ts * 1000).toLocaleTimeString();
const blockers = (a.blockers || [])
.map(b => `${b.name} (${b.vram_gb} GB, ${b.why})`).join(', ');
dec.innerHTML =
`<span class="${a.possible ? 'text-cyan-400' : 'text-amber-400'}">${when} · ` +
`${a.demanding} short by ${a.shortfall_gb} GB</span> — ${a.reason}` +
(blockers ? `<div class="text-slate-600">blocked by: ${blockers}</div>` : '');
}
// ---------------------------------------------------------------- job queue
async function fetchQueue() {
const body = document.getElementById('queue-body');
if (!body) return;
try {
const d = await (await fetch('/api/jobs?limit=40')).json();
const s = d.scheduler || {};
document.getElementById('queue-summary').textContent =
`${s.queue_depth ?? 0} queued · ${s.completed ?? 0} done · ${s.failed ?? 0} failed`;
// What the scheduler is doing right now, including why it is waiting. A job that
// cannot get VRAM used to sit silent, which made a stuck queue indistinguishable
// from an empty one.
const cur = document.getElementById('queue-current');
if (s.current) {
cur.innerHTML = `<span class="text-emerald-400">running</span> ` +
`<span class="text-slate-200">${s.current.tenant}</span>` +
`<span class="text-slate-500"> · ${s.current.label || s.current.id}</span>` +
(s.current.room && s.current.room.released && s.current.room.released.length
? `<span class="text-cyan-400"> · released ${s.current.room.released.join(', ')}</span>` : '');
} else if (s.blocked) {
cur.innerHTML = `<span class="text-amber-400">waiting ${s.blocked.waited_s}s</span> ` +
`<span class="text-slate-200">${s.blocked.tenant}</span>` +
`<span class="text-slate-500"> · ${s.blocked.reason || ''}</span>`;
} else {
cur.innerHTML = '<span class="text-slate-600">scheduler idle</span>';
}
const colour = {
pending: 'text-slate-400', running: 'text-emerald-400', done: 'text-cyan-500',
failed: 'text-rose-400', cancelled: 'text-slate-600',
};
body.innerHTML = (d.jobs || []).map(job => {
const when = job.state === 'pending'
? `waiting ${job.waiting_s}s`
: (job.duration_s != null ? `${job.duration_s}s` : '');
return `<div class="flex justify-between text-[11px] font-mono gap-2">
<span class="truncate">
<span class="${colour[job.state] || 'text-slate-400'}">${job.state}</span>
<span class="text-slate-600"> p${job.priority}</span>
<span class="text-slate-200"> ${job.tenant}</span>
<span class="text-slate-500">${job.label ? ' · ' + job.label : ''}</span>
</span>
<span class="text-slate-500 whitespace-nowrap">${when}${
job.state === 'pending'
? ` <button onclick="cancelJob('${job.id}')" class="text-rose-500 hover:text-rose-300 ml-1">cancel</button>`
: ''}</span>
</div>${job.error ? `<div class="text-[10px] text-rose-500/80 pl-2 truncate">${job.error}</div>` : ''}`;
}).join('') || '<div class="text-slate-600 text-[11px]">No jobs yet.</div>';
} catch (e) {
body.innerHTML = `<div class="text-rose-400 text-[11px]">${e}</div>`;
}
}
async function cancelJob(id) {
await fetch(`/api/jobs/${id}`, { method: 'DELETE' });
fetchQueue();
}
async function clearQueue() {
await fetch('/api/jobs', { method: 'DELETE' });
fetchQueue();
}
document.addEventListener('DOMContentLoaded', () => {
fetchQueue();
setInterval(fetchQueue, 2000);
});

View File

@@ -5,6 +5,11 @@
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>HYPERSWAP // Dual-Engine Model Orchestrator & Live Telemetry</title>
<script src="https://cdn.tailwindcss.com"></script>
<script>
// The CDN build has no breakpoint above 2xl, so a very wide window kept a
// two-column layout with increasingly stretched panels.
tailwind.config = { theme: { extend: { screens: { '3xl': '2000px' } } } };
</script>
<script src="https://cdn.jsdelivr.net/npm/chart.js@4.4.1/dist/chart.umd.min.js"></script>
<link rel="stylesheet" href="/static/styles.css">
<link rel="stylesheet" href="https://cdnjs.cloudflare.com/ajax/libs/font-awesome/6.4.0/css/all.min.css">
@@ -13,7 +18,7 @@
<!-- TOP HEADER -->
<header class="border-b border-slate-800 bg-slate-900/80 backdrop-blur sticky top-0 z-50">
<div class="max-w-7xl mx-auto px-4 sm:px-6 lg:px-8 py-3 flex flex-wrap items-center justify-between gap-4">
<div class="max-w-[2600px] mx-auto px-4 sm:px-6 lg:px-8 py-3 flex flex-wrap items-center justify-between gap-4">
<div class="flex items-center space-x-3">
<div class="w-10 h-10 rounded-xl bg-gradient-to-tr from-cyan-500 via-indigo-500 to-purple-500 flex items-center justify-center shadow-lg shadow-cyan-500/20">
<i class="fa-solid fa-bolt-lightning text-white text-lg"></i>
@@ -60,7 +65,7 @@
</header>
<!-- MAIN CONTAINER -->
<main class="max-w-7xl mx-auto px-4 sm:px-6 lg:px-8 py-6 space-y-6">
<main class="max-w-[2600px] mx-auto px-4 sm:px-6 lg:px-8 py-6 space-y-6">
<!-- HERO MEMORY GAUGES -->
<div class="grid grid-cols-1 md:grid-cols-2 gap-6">
@@ -194,7 +199,7 @@
</div>
<div>
<h3 class="font-bold text-slate-100 text-sm">Ollama LLM Engine</h3>
<p class="text-xs text-slate-400">Port :11434 // FlashAttention + Q4 KV Cache</p>
<p class="text-xs text-slate-400" id="ollama-engine-sub" title="Read live from the ollama service environment">Port :11434</p>
</div>
</div>
<button onclick="freeOllamaVRAM()" class="px-2.5 py-1 text-xs font-semibold rounded-lg bg-rose-950/70 border border-rose-800 text-rose-300 hover:bg-rose-900 transition flex items-center space-x-1">
@@ -265,7 +270,7 @@
</div>
<div>
<h3 class="font-bold text-slate-100 text-sm">ComfyUI Diffusion Engine</h3>
<p class="text-xs text-slate-400">Port :8188 // DynamicVRAM + Pinned Async Offload</p>
<p class="text-xs text-slate-400" id="comfy-engine-sub" title="Read live from ComfyUI's /system_stats">Port :8188</p>
</div>
</div>
<button onclick="freeComfyVRAM()" class="px-2.5 py-1 text-xs font-semibold rounded-lg bg-rose-950/70 border border-rose-800 text-rose-300 hover:bg-rose-900 transition flex items-center space-x-1">
@@ -431,7 +436,7 @@
</div>
<!-- OVERCLOCK CONTROL PANEL -->
<div class="bg-slate-900/80 border border-fuchsia-900/50 rounded-2xl p-5 space-y-4 shadow-lg shadow-fuchsia-950/30">
<div class="lg:col-span-2 bg-slate-900/80 border border-fuchsia-900/50 rounded-2xl p-5 space-y-4 shadow-lg shadow-fuchsia-950/30">
<div class="flex flex-wrap items-center justify-between gap-3 pb-3 border-b border-slate-800">
<div class="flex items-center space-x-2">
<div class="p-2 rounded-lg bg-fuchsia-950/80 border border-fuchsia-800 text-fuchsia-400">
@@ -599,8 +604,8 @@
<span class="flex items-center space-x-1"><span class="w-2 h-2 rounded-full bg-fuchsia-400 inline-block"></span>RAM Cache GB</span>
</div>
</div>
<div class="relative h-64">
<canvas id="oc-chart"></canvas>
<div class="relative h-[clamp(18rem,45vh,52rem)]">
<canvas id="oc-chart" class="absolute inset-0 !w-full !h-full"></canvas>
</div>
</div>
@@ -672,9 +677,50 @@
</div>
<!-- ============ NEXT-LEVEL PANELS: governor / residency / analytics / autotune ============ -->
<div class="grid grid-cols-1 xl:grid-cols-2 gap-5 mt-5">
<div class="grid grid-cols-1 xl:grid-cols-2 3xl:grid-cols-3 gap-5 mt-5">
<!-- The cross-application job queue -->
<div class="bg-slate-900/80 border border-slate-800 rounded-2xl p-5 xl:col-span-2">
<div class="flex items-center justify-between pb-3 border-b border-slate-800">
<div class="flex items-center space-x-2">
<div class="p-2 rounded-lg bg-violet-950/80 border border-violet-800 text-violet-400">
<i class="fa-solid fa-list-check text-sm"></i>
</div>
<div>
<h3 class="font-bold text-slate-100 text-sm">Job Queue</h3>
<p class="text-xs text-slate-400">Work lined up across every application, highest priority first</p>
</div>
</div>
<div class="flex items-center space-x-2">
<span id="queue-summary" class="text-xs font-mono text-slate-500">—</span>
<button onclick="clearQueue()" class="px-2.5 py-1 text-xs font-semibold rounded-lg bg-slate-800 border border-slate-700 text-slate-300 hover:bg-slate-700 transition">Clear pending</button>
</div>
</div>
<div id="queue-current" class="mt-4 text-[11px] font-mono"></div>
<div id="queue-body" class="mt-3 space-y-1.5 max-h-64 overflow-y-auto pr-1"></div>
</div>
<!-- All GPU tenants, however many are configured -->
<div class="bg-slate-900/80 border border-slate-800 rounded-2xl p-5 xl:col-span-2">
<div class="flex items-center justify-between pb-3 border-b border-slate-800">
<div class="flex items-center space-x-2">
<div class="p-2 rounded-lg bg-teal-950/80 border border-teal-800 text-teal-400">
<i class="fa-solid fa-layer-group text-sm"></i>
</div>
<div>
<h3 class="font-bold text-slate-100 text-sm">GPU Tenants</h3>
<p class="text-xs text-slate-400">Every application contending for the card, from <code class="text-teal-400">tenants.json</code></p>
</div>
</div>
<span id="tenants-free" class="text-xs font-mono text-slate-500">—</span>
</div>
<div id="tenants-body" class="mt-4 space-y-2"></div>
<div id="tenants-decision" class="mt-3 text-[11px] font-mono text-slate-400"></div>
</div>
<!-- System health: makes a silently-broken dependency loud -->
<div class="bg-slate-900/80 border border-slate-800 rounded-2xl p-5 xl:col-span-2">
<div class="flex items-center justify-between pb-3 border-b border-slate-800">

108
tenants.json Normal file
View File

@@ -0,0 +1,108 @@
[
{
"name": "ollama",
"kind": "llm",
"priority": 50,
"match": {
"names": [
"ollama"
],
"cmdline": [
"llama-server",
"ollama"
]
},
"busy": {
"type": "http_count",
"url": "http://localhost:11434/api/ps",
"count_keys": [
"models"
]
},
"release": {
"type": "http_post",
"url": "http://localhost:11434/api/generate",
"body": {
"keep_alive": 0
},
"per_model": true,
"timeout_s": 120.0
},
"notes": "Unloads per model. With OLLAMA_NUM_PARALLEL=1 the request queues behind any running generation and applies when it finishes."
},
{
"name": "comfyui",
"kind": "diffusion",
"priority": 60,
"match": {
"cmdline": [
"comfyui",
"comfy"
],
"cmdline_endswith": [
"main.py"
]
},
"busy": {
"type": "http_count",
"url": "http://127.0.0.1:8188/queue",
"count_keys": [
"queue_running",
"queue_pending"
],
"vram_floor_gb": 1.5,
"stale_after_s": 90.0
},
"release": {
"type": "http_post",
"url": "http://127.0.0.1:8188/free",
"body": {
"unload_models": true,
"free_memory": true
},
"timeout_s": 30.0
},
"notes": "Leaves dead jobs in queue_running; the queue flag is corroborated against its own VRAM before being believed."
},
{
"name": "stt-relay",
"kind": "other",
"priority": 70,
"match": {
"cmdline": [
"stt_relay.py"
]
},
"busy": {
"type": "vram",
"vram_busy_gb": 1.0
},
"release": {
"type": "none"
},
"notes": "Long-running speech relay. Holds ~0.8 GB permanently and exposes no release API, so its VRAM is headroom this service can never offer. Declared so it is named rather than lumped into 'unmanaged'."
},
{
"name": "desktop",
"kind": "desktop",
"priority": 90,
"match": {
"names": [
"gnome-shell",
"xorg",
"mutter",
"kwin",
"plasmashell",
"gnome-remote-desktop",
"sddm",
"gdm",
"picom",
"weston"
]
},
"release": {
"type": "none"
},
"notes": "Compositor and display server. Small, permanent, never reclaimable."
}
]

461
tenants.py Normal file
View File

@@ -0,0 +1,461 @@
"""GPU tenants: the applications competing for the card, described as data.
The point of this service is fast handoff of a single GPU between applications. It grew
up around the two on this box, and their names ended up compiled into process matching,
VRAM attribution, busy detection and release calls alike -- roughly 385 references. That
makes it a script for Ollama and ComfyUI rather than a GPU arbitrator.
A tenant is described here instead:
* how to recognise its processes (match)
* how to tell whether it is actually working (busy probe)
* how to ask it to give VRAM back (release strategy)
* how much it matters when two want the card (priority)
Ollama and ComfyUI ship as defaults so behaviour is unchanged, but nothing about the
arbitration logic knows their names. A third application -- a training run, a speech
model, another inference server -- is a config entry, not a code change. A tenant that
cannot be released (no API to ask) is still worth declaring, because naming it turns
"unmanaged VRAM" into "held by X, which cannot be reclaimed".
"""
import json
import logging
import os
import time
from dataclasses import dataclass, field, asdict
from typing import Any, Dict, List, Optional
import httpx
import psutil
logger = logging.getLogger("tenants")
_BASE = os.path.dirname(os.path.abspath(__file__))
CONFIG_PATH = os.environ.get("HYPERSWAP_TENANTS", os.path.join(_BASE, "tenants.json"))
# Kinds are advisory: they drive presentation and sensible defaults, never control flow.
KIND_LLM, KIND_DIFFUSION, KIND_DESKTOP, KIND_OTHER = "llm", "diffusion", "desktop", "other"
@dataclass
class ProcessMatch:
"""How to recognise a tenant's processes among those NVML reports."""
names: List[str] = field(default_factory=list) # matched against process name
cmdline: List[str] = field(default_factory=list) # substrings of the full cmdline
cmdline_endswith: List[str] = field(default_factory=list)
def matches(self, pname: str, cmdline: str) -> bool:
pname, cmdline = pname.lower(), cmdline.lower()
if any(n.lower() in pname for n in self.names):
return True
if any(c.lower() in cmdline for c in self.cmdline):
return True
return any(cmdline.rstrip().endswith(c.lower()) for c in self.cmdline_endswith)
@dataclass
class BusyProbe:
"""How to tell whether a tenant is genuinely working.
`vram_floor_gb` exists because a queue flag can lie: ComfyUI leaves dead jobs in
queue_running, and only its VRAM reveals that nothing is loaded. GPU utilisation is
deliberately unavailable as a signal -- it is shared by every tenant, so it cannot
attribute work to one of them.
"""
type: str = "none" # none | http_count | vram
url: Optional[str] = None
count_keys: List[str] = field(default_factory=list) # keys whose lists are summed
vram_busy_gb: float = 0.0 # busy when its VRAM exceeds this
vram_floor_gb: float = 0.0 # below this it holds no real work
stale_after_s: float = 90.0
@dataclass
class EventSource:
"""A stream that tells us *when* to look, not what to think.
ComfyUI publishes a websocket, and the original listener parsed its message types to
decide what was happening -- which meant understanding one application's schema. Any
message is instead treated purely as a wake-up: re-run this tenant's busy probe now
rather than waiting for the next poll. That gives sub-second reaction to any
application with an event stream, with no knowledge of what it emits.
"""
type: str = "none" # none | websocket
url: Optional[str] = None
reconnect_backoff_s: float = 2.0
max_backoff_s: float = 15.0
@dataclass
class ReleaseStrategy:
"""How to ask a tenant to give VRAM back."""
type: str = "none" # none | http_post
url: Optional[str] = None
body: Dict[str, Any] = field(default_factory=dict)
# Set when the call must name the loaded model (Ollama unloads per model).
per_model: bool = False
timeout_s: float = 120.0
confirm: bool = True # wait for NVML to show the memory released
@dataclass
class GpuTenant:
name: str
kind: str = KIND_OTHER
enabled: bool = True
# Higher wins contention; a tenant yields to anything above it.
priority: int = 50
# How much free VRAM this application needs before it can work. Used to decide
# whether a busy tenant is actually being starved, rather than merely busy.
needs_vram_gb: float = 0.0
# How long a reclaimable tenant may sit idle holding VRAM before it is asked for it
# back. Iterating on a ComfyUI workflow should not pay a reload between every run,
# so this is deliberately not immediate.
idle_release_after_s: float = 30.0
# GPU profile to apply while this tenant is the active workload. Clock and power
# tuning is workload-specific -- diffusion is compute bound, LLM decode is bandwidth
# bound -- and that was previously switched by application name in the arbitrator.
overclock_profile: Optional[str] = None
# VRAM that survives a release. ComfyUI keeps its CUDA context for as long as the
# process lives, so purging it does not return everything it holds. Ignoring this
# made plan_release over-promise: it reported that releasing ComfyUI would free
# 0.37 GB against a 0.33 GB shortfall, the job was cleared to run, and the memory
# never actually arrived.
vram_floor_gb: float = 0.0
match: ProcessMatch = field(default_factory=ProcessMatch)
busy: BusyProbe = field(default_factory=BusyProbe)
release: ReleaseStrategy = field(default_factory=ReleaseStrategy)
events: EventSource = field(default_factory=EventSource)
notes: str = ""
@property
def reclaimable(self) -> bool:
return self.release.type != "none"
def to_dict(self) -> Dict[str, Any]:
return asdict(self)
def _tenant_from_dict(d: Dict[str, Any]) -> GpuTenant:
return GpuTenant(
name=d["name"],
kind=d.get("kind", KIND_OTHER),
enabled=d.get("enabled", True),
priority=int(d.get("priority", 50)),
needs_vram_gb=float(d.get("needs_vram_gb", 0.0)),
overclock_profile=d.get("overclock_profile"),
vram_floor_gb=float(d.get("vram_floor_gb", 0.0)),
idle_release_after_s=float(d.get("idle_release_after_s", 30.0)),
match=ProcessMatch(**(d.get("match") or {})),
busy=BusyProbe(**(d.get("busy") or {})),
release=ReleaseStrategy(**(d.get("release") or {})),
events=EventSource(**(d.get("events") or {})),
notes=d.get("notes", ""),
)
# Defaults reproduce today's behaviour exactly; they are data, not special cases.
DEFAULT_TENANTS: List[Dict[str, Any]] = [
{
"name": "ollama",
"kind": KIND_LLM,
# Lower than ComfyUI on purpose: an interactive diffusion job preempts the LLM,
# whose weights stay in the page cache and reload in seconds. Getting this the
# wrong way round silently disabled the service's central behaviour -- ComfyUI
# could never reclaim from Ollama.
"priority": 50,
"needs_vram_gb": 4.0,
"idle_release_after_s": 0.0,
"overclock_profile": "ollama",
"match": {"names": ["ollama"], "cmdline": ["llama-server", "ollama"]},
"busy": {"type": "http_count", "url": "http://localhost:11434/api/ps",
"count_keys": ["models"]},
"release": {"type": "http_post", "url": "http://localhost:11434/api/generate",
"body": {"keep_alive": 0}, "per_model": True, "timeout_s": 120.0},
"notes": "Unloads per model. With OLLAMA_NUM_PARALLEL=1 the request queues "
"behind any running generation and applies when it finishes.",
},
{
"name": "comfyui",
"kind": KIND_DIFFUSION,
"priority": 60,
"needs_vram_gb": 6.0,
"idle_release_after_s": 30.0,
"overclock_profile": "comfy",
"vram_floor_gb": 0.45,
"match": {"cmdline": ["comfyui", "comfy"], "cmdline_endswith": ["main.py"]},
"busy": {"type": "http_count", "url": "http://127.0.0.1:8188/queue",
"count_keys": ["queue_running", "queue_pending"],
"vram_floor_gb": 1.5, "stale_after_s": 90.0},
"release": {"type": "http_post", "url": "http://127.0.0.1:8188/free",
"body": {"unload_models": True, "free_memory": True},
"timeout_s": 30.0},
"events": {"type": "websocket", "url": "ws://127.0.0.1:8188/ws?clientId=hyperswap"},
"notes": "Leaves dead jobs in queue_running; the queue flag is corroborated "
"against its own VRAM before being believed.",
},
{
"name": "desktop",
"kind": KIND_DESKTOP,
"priority": 90,
"match": {"names": ["gnome-shell", "xorg", "mutter", "kwin", "plasmashell",
"gnome-remote-desktop", "sddm", "gdm", "picom", "weston"]},
"release": {"type": "none"},
"notes": "Compositor and display server. Small, permanent, never reclaimable.",
},
]
_cache: Dict[str, Any] = {"ts": 0.0, "tenants": None, "mtime": None}
CACHE_TTL_S = 10.0
def load_tenants(force: bool = False) -> List[GpuTenant]:
"""Load tenant definitions, writing the defaults out on first run."""
now = time.time()
try:
mtime = os.path.getmtime(CONFIG_PATH) if os.path.exists(CONFIG_PATH) else None
except OSError:
mtime = None
if (not force and _cache["tenants"] is not None
and mtime == _cache["mtime"] and (now - _cache["ts"]) < CACHE_TTL_S):
return _cache["tenants"]
raw: List[Dict[str, Any]]
if os.path.exists(CONFIG_PATH):
try:
with open(CONFIG_PATH) as f:
raw = json.load(f)
except Exception as e:
logger.error(f"could not read {CONFIG_PATH}, using defaults: {e}")
raw = DEFAULT_TENANTS
else:
raw = DEFAULT_TENANTS
try:
with open(CONFIG_PATH, "w") as f:
json.dump(DEFAULT_TENANTS, f, indent=2)
logger.info(f"wrote default tenant definitions to {CONFIG_PATH}")
except Exception as e:
logger.warning(f"could not write {CONFIG_PATH}: {e}")
# Merge in any fields a shipped default has gained since the config was written.
# Without this, adding a field silently disables the behaviour it controls for every
# existing install -- needs_vram_gb defaulted to 0, which made starvation
# undetectable for the two tenants that had been written out before it existed.
defaults_by_name = {d["name"]: d for d in DEFAULT_TENANTS}
tenants = []
for d in raw:
base = defaults_by_name.get(d.get("name"))
if base:
merged = {**base, **d}
for key in ("match", "busy", "release", "events"):
if isinstance(base.get(key), dict):
merged[key] = {**base[key], **(d.get(key) or {})}
d = merged
try:
tenants.append(_tenant_from_dict(d))
except Exception as e:
logger.error(f"skipping malformed tenant {d!r}: {e}")
_cache.update({"ts": now, "tenants": tenants, "mtime": mtime})
return tenants
def get_tenant(name: str) -> Optional[GpuTenant]:
return next((t for t in load_tenants() if t.name == name), None)
def save_tenants(tenants: List[Dict[str, Any]]) -> bool:
try:
with open(CONFIG_PATH, "w") as f:
json.dump(tenants, f, indent=2)
_cache["tenants"] = None
return True
except Exception as e:
logger.error(f"save_tenants failed: {e}")
return False
def classify_process(pname: str, cmdline: str) -> str:
"""Return the owning tenant's name, or 'unmanaged'.
'unmanaged' is meaningful rather than a dumping ground: it is VRAM this service has
no way to reclaim, and it is reported as such.
"""
for t in load_tenants():
if t.enabled and t.match.matches(pname, cmdline):
return t.name
return "unmanaged"
def classify_pid(pid: int) -> str:
try:
proc = psutil.Process(pid)
return classify_process(proc.name(), " ".join(proc.cmdline()))
except Exception:
return "unmanaged"
async def probe_busy(tenant: GpuTenant, vram_gb: float = 0.0,
state: Optional[Dict[str, Any]] = None) -> Dict[str, Any]:
"""Is this tenant actually working? Returns {busy, reason, stale}."""
probe = tenant.busy
if probe.type == "vram":
busy = vram_gb > probe.vram_busy_gb
return {"busy": busy, "reason": f"{vram_gb:.2f} GB held", "stale": False}
if probe.type != "http_count" or not probe.url:
return {"busy": False, "reason": "no busy probe configured", "stale": False}
try:
async with httpx.AsyncClient(timeout=3.0) as c:
r = await c.get(probe.url)
if r.status_code != 200:
return {"busy": False, "reason": f"probe HTTP {r.status_code}", "stale": False}
data = r.json()
count = sum(len(data.get(k) or []) for k in probe.count_keys)
except Exception as e:
return {"busy": False, "reason": f"probe failed: {str(e)[:60]}", "stale": False}
if count == 0:
return {"busy": False, "reason": "queue empty", "stale": False}
# A queue that claims work while the tenant holds no VRAM is not doing work.
if probe.vram_floor_gb and vram_gb < probe.vram_floor_gb:
return {"busy": True, "reason": f"{count} queued, holding {vram_gb:.2f} GB",
"stale": None, "below_floor": True}
return {"busy": True, "reason": f"{count} queued/running", "stale": False}
async def release_vram(tenant: GpuTenant, models: Optional[List[str]] = None
) -> Dict[str, Any]:
"""Ask a tenant to give its VRAM back, however that tenant expects to be asked."""
strategy = tenant.release
if strategy.type == "none" or not strategy.url:
return {"success": False, "tenant": tenant.name, "released": False,
"reason": "this tenant exposes no way to release VRAM"}
t0 = time.perf_counter()
payloads: List[Dict[str, Any]] = []
if strategy.per_model:
for m in (models or []):
payloads.append({**strategy.body, "model": m})
if not payloads:
return {"success": True, "tenant": tenant.name, "released": False,
"reason": "nothing loaded to release"}
else:
payloads.append(dict(strategy.body))
errors = []
try:
async with httpx.AsyncClient(timeout=strategy.timeout_s) as c:
for body in payloads:
try:
await c.post(strategy.url, json=body)
except Exception as e:
errors.append(str(e)[:80])
except Exception as e:
errors.append(str(e)[:80])
return {
"success": not errors,
"tenant": tenant.name,
"released": True,
"requests": len(payloads),
"duration_ms": round((time.perf_counter() - t0) * 1000, 2),
"errors": errors or None,
}
def describe() -> List[Dict[str, Any]]:
"""Tenant definitions for the API, with what each can and cannot do."""
out = []
for t in sorted(load_tenants(), key=lambda x: -x.priority):
d = t.to_dict()
d["reclaimable"] = t.reclaimable
d["busy_probe"] = t.busy.type
d["release_via"] = t.release.type
out.append(d)
return out
# ---------------------------------------------------------------- arbitration
def plan_release(demanding: str, tenants_state: List[Dict[str, Any]],
free_gb: float, needed_gb: float) -> Dict[str, Any]:
"""Decide who should give up VRAM so a starved tenant can work.
Generic over any number of applications: candidates are every *reclaimable* tenant
that is not itself busy and ranks below the demanding one, taken lowest priority
first, until enough would be freed. The two-application version of this was a pair
of hardcoded rules -- yield Ollama for ComfyUI, purge ComfyUI for Ollama -- which
could not express a third participant at all.
Returns the plan rather than performing it, so the decision is testable and can be
logged before anything is actually released.
"""
by_name = {s["name"]: s for s in tenants_state}
demander = by_name.get(demanding)
if not demander:
return {"possible": False, "reason": f"unknown tenant '{demanding}'", "release": []}
# The demander keeps what it already holds; only the remainder must be found.
shortfall = needed_gb - free_gb - demander.get("vram_gb", 0.0)
if shortfall <= 0:
return {"possible": True, "reason": "enough VRAM is already free",
"release": [], "shortfall_gb": 0.0}
# Who may be asked for memory:
#
# * any idle reclaimable tenant, whatever its rank -- idle memory is not in use;
# * a *busy* tenant that ranks strictly below the demander.
#
# That second clause is the point of the whole service and was nearly lost. Refusing
# to touch anything busy looks safe and is not: a diffusion job measured here ran for
# 46 s instead of 3 s, squeezed into 1.6 GB, because the LLM reloaded straight after
# yielding and was then protected as "busy" while ComfyUI starved. Preempting a
# lower-priority tenant is safe precisely because releasing is asynchronous -- an
# Ollama unload queues behind its running request and applies when that finishes, so
# nothing is killed mid-flight.
#
# Equal or higher priority is never interrupted, so peers cannot fight.
demander_priority = demander.get("priority", 0)
candidates = [
s for s in tenants_state
if s["name"] != demanding
and s.get("reclaimable")
and s.get("vram_gb", 0) > 0
and (not s.get("busy") or s.get("priority", 0) < demander_priority)
]
# Idle tenants first, then lowest priority: never disturb working software while
# something idle still has memory to give.
candidates.sort(key=lambda s: (bool(s.get("busy")), s.get("priority", 0),
-s.get("vram_gb", 0)))
plan, freed = [], 0.0
for c in candidates:
if freed >= shortfall:
break
# Only what the tenant can actually give back, not everything it holds.
releasable = max(c.get("vram_gb", 0.0) - c.get("vram_floor_gb", 0.0), 0.0)
if releasable <= 0:
continue
plan.append(c["name"])
freed += releasable
blockers = [
{"name": s["name"], "vram_gb": s.get("vram_gb", 0.0),
"why": ("busy and ranks at or above the demander" if s.get("busy") else
"declares no release mechanism" if not s.get("reclaimable") else
"enough was freed without it")}
for s in tenants_state
if s["name"] != demanding and s.get("vram_gb", 0) > 0 and s["name"] not in plan
]
return {
"possible": freed >= shortfall,
"shortfall_gb": round(shortfall, 2),
"would_free_gb": round(freed, 2),
"release": plan,
"blockers": blockers,
"reason": (f"releasing {', '.join(plan)} frees {freed:.2f} GB of the "
f"{shortfall:.2f} GB shortfall" if plan else
"no reclaimable idle tenant holds enough VRAM"),
}

View File

@@ -67,3 +67,24 @@ def temp_db(tmp_path, monkeypatch):
telemetry_store.stop()
except Exception:
pass
@pytest.fixture(autouse=True)
def isolated_tenant_registry(tmp_path, monkeypatch):
"""Never let tests read the operator's live tenants.json.
Classification is now configuration, which means a test that reads the real config
changes result when someone adds an application to their own machine -- exactly what
happened when stt-relay was registered and a "third party is unmanaged" test started
seeing it as a named tenant. Every test gets the shipped defaults unless it opts out
by pointing CONFIG_PATH somewhere itself.
"""
import json as _json
import tenants as _tenants
path = tmp_path / "tenants-default.json"
path.write_text(_json.dumps(_tenants.DEFAULT_TENANTS))
monkeypatch.setattr(_tenants, "CONFIG_PATH", str(path))
_tenants._cache.update({"ts": 0.0, "tenants": None, "mtime": None})
yield
_tenants._cache.update({"ts": 0.0, "tenants": None, "mtime": None})

133
tests/test_engines.py Normal file
View File

@@ -0,0 +1,133 @@
"""Tests for live engine-configuration reporting.
These settings live outside this codebase but dictate how arbitration must behave, and
working out why a yield behaved a certain way once meant reading journald by hand. The
dashboard previously asserted them as hardcoded text, which happened to be accurate --
worse than being wrong, because it would have stayed accurate-looking after the settings
changed.
"""
import asyncio
import pytest
import engines
class _Proc:
def __init__(self, stdout=""):
self.stdout = stdout
self.returncode = 0
class TestOllamaEnvironmentParsing:
def test_parses_the_real_unit_environment(self, monkeypatch):
# Verbatim from `systemctl show ollama -p Environment --value` on this machine.
raw = ("OLLAMA_HOST=0.0.0.0:11434 OLLAMA_FLASH_ATTENTION=1 "
"OLLAMA_KV_CACHE_TYPE=q4_0 OLLAMA_KEEP_ALIVE=30m "
"OLLAMA_MAX_LOADED_MODELS=1 OLLAMA_NUM_PARALLEL=1 OLLAMA_NUM_BATCH=2048")
monkeypatch.setattr(engines.subprocess, "run", lambda *a, **k: _Proc(raw))
env = engines._ollama_unit_environment()
assert env["OLLAMA_NUM_PARALLEL"] == "1"
assert env["OLLAMA_MAX_LOADED_MODELS"] == "1"
assert env["OLLAMA_KV_CACHE_TYPE"] == "q4_0"
def test_ignores_non_ollama_variables(self, monkeypatch):
monkeypatch.setattr(engines.subprocess, "run",
lambda *a, **k: _Proc("PATH=/usr/bin OLLAMA_HOST=x:1 HOME=/root"))
env = engines._ollama_unit_environment()
assert set(env) == {"OLLAMA_HOST"}
def test_returns_empty_rather_than_raising_when_systemctl_fails(self, monkeypatch):
def boom(*a, **k):
raise FileNotFoundError("systemctl")
monkeypatch.setattr(engines.subprocess, "run", boom)
assert engines._ollama_unit_environment() == {}
class TestEngineConfigReport:
def _run(self, monkeypatch, env, comfy_ok=True):
monkeypatch.setattr(engines, "_ollama_unit_environment", lambda: env)
class _Resp:
status_code = 200 if comfy_ok else 500
def json(self):
return {"system": {"comfyui_version": "0.33.1",
"pytorch_version": "2.11.0+cu128",
"python_version": "3.14.4 (main)",
"argv": ["main.py", "--listen", "0.0.0.0"]},
"devices": [{"name": "cuda:0 NVIDIA GeForce RTX 4080 SUPER "
": cudaMallocAsync"}]}
class _Client:
async def __aenter__(self): return self
async def __aexit__(self, *a): return False
async def get(self, url): return _Resp()
monkeypatch.setattr(engines.httpx, "AsyncClient", lambda **k: _Client())
return asyncio.run(engines.get_engine_config())
def test_surfaces_the_settings_that_drive_arbitration(self, monkeypatch):
d = self._run(monkeypatch, {"OLLAMA_NUM_PARALLEL": "1",
"OLLAMA_MAX_LOADED_MODELS": "1",
"OLLAMA_KEEP_ALIVE": "30m"})
assert d["ollama"]["num_parallel"] == "1"
assert d["ollama"]["max_loaded_models"] == "1"
assert d["ollama"]["keep_alive"] == "30m"
def test_num_parallel_explains_the_deferred_yield_behaviour(self, monkeypatch):
d = self._run(monkeypatch, {"OLLAMA_NUM_PARALLEL": "1"})
note = next(s["means"] for s in d["ollama"]["settings"]
if s["key"] == "OLLAMA_NUM_PARALLEL")
# The explanation is the point: it is why a busy model is deferred, not failed.
assert "queue" in note.lower()
def test_summary_reflects_actual_flags_not_a_fixed_string(self, monkeypatch):
on = self._run(monkeypatch, {"OLLAMA_FLASH_ATTENTION": "1",
"OLLAMA_KV_CACHE_TYPE": "q4_0"})
assert "FlashAttention" in on["ollama"]["summary"]
assert "q4_0" in on["ollama"]["summary"]
off = self._run(monkeypatch, {})
assert "FlashAttention" not in off["ollama"]["summary"]
assert off["ollama"]["config_source"] == "unavailable"
def test_port_comes_from_ollama_host(self, monkeypatch):
d = self._run(monkeypatch, {"OLLAMA_HOST": "0.0.0.0:11500"})
assert d["ollama"]["port"] == "11500"
def test_comfy_allocator_and_vram_mode_are_read_not_asserted(self, monkeypatch):
d = self._run(monkeypatch, {})
assert d["comfyui"]["allocator"] == "cudaMallocAsync"
assert d["comfyui"]["vram_mode"] == "default (auto)"
assert d["comfyui"]["version"] == "0.33.1"
def test_comfy_vram_flag_is_detected_when_present(self, monkeypatch):
monkeypatch.setattr(engines, "_ollama_unit_environment", lambda: {})
class _Resp:
status_code = 200
def json(self):
return {"system": {"argv": ["main.py", "--lowvram"]},
"devices": [{"name": "cuda:0 X : cudaMalloc"}]}
class _Client:
async def __aenter__(self): return self
async def __aexit__(self, *a): return False
async def get(self, url): return _Resp()
monkeypatch.setattr(engines.httpx, "AsyncClient", lambda **k: _Client())
d = asyncio.run(engines.get_engine_config())
assert d["comfyui"]["vram_mode"] == "--lowvram"
assert d["comfyui"]["allocator"] == "cudaMalloc"
def test_offline_comfy_is_reported_not_raised(self, monkeypatch):
monkeypatch.setattr(engines, "_ollama_unit_environment", lambda: {})
class _Client:
async def __aenter__(self): return self
async def __aexit__(self, *a): return False
async def get(self, url): raise ConnectionError("refused")
monkeypatch.setattr(engines.httpx, "AsyncClient", lambda **k: _Client())
d = asyncio.run(engines.get_engine_config())
assert d["comfyui"]["online"] is False
assert "error" in d["comfyui"]

View File

@@ -105,7 +105,8 @@ class TestAggregation:
monkeypatch.setattr(health, "_check_nvml", lambda: checks[0])
monkeypatch.setattr(health, "_check_sudo_smi", lambda: checks[1])
for fn in ("_check_fan_control", "_check_profile_drift", "_check_store",
"_check_residency", "_check_model_dirs", "_check_comfy_ws"):
"_check_residency", "_check_model_dirs", "_check_comfy_ws",
"_check_unmanaged_vram", "_check_comfy_queue"):
monkeypatch.setattr(health, fn, lambda: health._check("x", health.OK, "d"))
async def fake_http(name, url, impact, fix):
@@ -121,7 +122,7 @@ class TestAggregation:
monkeypatch.setattr(health, "_check_nvml", boom)
for fn in ("_check_sudo_smi", "_check_fan_control", "_check_profile_drift",
"_check_store", "_check_residency", "_check_model_dirs",
"_check_comfy_ws"):
"_check_comfy_ws", "_check_unmanaged_vram", "_check_comfy_queue"):
monkeypatch.setattr(health, fn, lambda: health._check("x", health.OK, "d"))
async def fake_http(name, url, impact, fix):
@@ -132,3 +133,84 @@ class TestAggregation:
# A broken check must surface as failed, not take down the endpoint.
assert res["status"] == health.FAILED
assert any("exploded" in c["detail"] for c in res["checks"])
class TestUnmanagedVramCheck:
"""Turning an unreclaimable-VRAM number into something actionable.
The arithmetic here has to be right or the check is worse than useless. A first
version omitted ComfyUI's CUDA context -- which survives a purge -- and so reported
a 14.93 GB model as fitting against a real ceiling of 14.60 GB. That was the very
model the service had just refused with 507 Insufficient Storage.
"""
def _gpu(self, unmanaged_gb=0.82, desktop_gb=0.01, comfy_gb=0.56, total=15.99,
procs=None):
return {
"available": True,
"vram_total_gb": total,
"breakdown": {
"unmanaged_gb": unmanaged_gb, "desktop_gb": desktop_gb,
"comfyui_gb": comfy_gb,
"unmanaged": procs if procs is not None else
[{"pid": 1, "name": "python", "vram_mb": unmanaged_gb * 1024,
"cmdline": "stt_relay.py"}],
},
}
def _blobs(self, sizes):
return [{"model": f"m{i}", "size_gb": s} for i, s in enumerate(sizes)]
def test_ok_when_nothing_holds_unreclaimable_vram(self, monkeypatch):
monkeypatch.setattr(health.vram_arbitrator, "get_gpu_hardware_stats",
lambda: self._gpu(unmanaged_gb=0.0, procs=[]))
assert health._check_unmanaged_vram()["status"] == health.OK
def test_comfy_cuda_context_counts_against_the_ceiling(self, monkeypatch):
# 15.99 - 0.82 unmanaged - 0.01 desktop - 0.56 comfy floor = 14.60 GB available.
# A 12.87 GB blob needs 12.87 * 1.16 = 14.93 GB, so it does not fit -- matching
# the observed 507.
monkeypatch.setattr(health.vram_arbitrator, "get_gpu_hardware_stats",
lambda: self._gpu())
monkeypatch.setattr(health, "_comfy_vram_floor_gb", lambda default=0, days=1: 0.56)
monkeypatch.setattr(health.ram_optimizer, "find_ollama_model_files",
lambda: self._blobs([12.87]))
res = health._check_unmanaged_vram()
assert res["status"] == health.DEGRADED
assert "1 model(s)" in res["impact"]
def test_model_that_fits_even_without_the_unmanaged_process_is_not_flagged(self, monkeypatch):
# A tiny model fits either way, so the unmanaged process is not what blocks it.
monkeypatch.setattr(health.vram_arbitrator, "get_gpu_hardware_stats",
lambda: self._gpu())
monkeypatch.setattr(health, "_comfy_vram_floor_gb", lambda default=0, days=1: 0.56)
monkeypatch.setattr(health.ram_optimizer, "find_ollama_model_files",
lambda: self._blobs([2.0]))
assert health._check_unmanaged_vram()["status"] == health.OK
def test_model_too_big_to_ever_fit_is_not_blamed_on_the_process(self, monkeypatch):
# A 23.7 GB model does not fit on a 16 GB card regardless; saying the 842 MB
# process is why would send the user after the wrong thing.
monkeypatch.setattr(health.vram_arbitrator, "get_gpu_hardware_stats",
lambda: self._gpu())
monkeypatch.setattr(health, "_comfy_vram_floor_gb", lambda default=0, days=1: 0.56)
monkeypatch.setattr(health.ram_optimizer, "find_ollama_model_files",
lambda: self._blobs([23.7]))
assert health._check_unmanaged_vram()["status"] == health.OK
def test_floor_uses_the_minimum_observed_not_the_current_value(self, monkeypatch):
# Current VRAM could be a 7 GB checkpoint mid-generation; the floor is what
# survives a purge.
monkeypatch.setattr(health.telemetry_store, "_rows",
lambda *a, **k: [{"floor": int(0.24 * 1024 ** 3)}])
assert health._comfy_vram_floor_gb(default=7.0) == 0.24
def test_floor_falls_back_when_history_is_empty(self, monkeypatch):
monkeypatch.setattr(health.telemetry_store, "_rows", lambda *a, **k: [])
assert health._comfy_vram_floor_gb(default=0.56) == 0.56
def test_floor_falls_back_rather_than_raising(self, monkeypatch):
def boom(*a, **k):
raise RuntimeError("db gone")
monkeypatch.setattr(health.telemetry_store, "_rows", boom)
assert health._comfy_vram_floor_gb(default=0.5) == 0.5

184
tests/test_jobs.py Normal file
View File

@@ -0,0 +1,184 @@
"""Tests for the cross-tenant job queue.
Every case below corresponds to something that actually went wrong while building this,
because the failure modes are not obvious from the code:
* dispatching without room does not fail gracefully -- a CUDA OOM kills llama-server,
and three queued jobs were destroyed in a row;
* the VRAM requirement is a property of the job's model, not of the tenant, and a flat
4 GB let a 14.9 GB model be dispatched into 8 GB of free memory;
* waiting forever is as wrong as failing immediately, when the job can never fit;
* an exception during dispatch left the row RUNNING while the scheduler moved on.
"""
import asyncio
import json
import time
import pytest
import jobs as J
import telemetry_store
@pytest.fixture
def queue(tmp_path, monkeypatch):
monkeypatch.setattr(telemetry_store, "DB_PATH", str(tmp_path / "q.db"))
J.init()
return tmp_path
class TestQueueBasics:
def test_submit_and_read_back(self, queue):
r = J.submit("ollama", {"model": "m"}, label="first")
assert r["success"]
job = J.get(r["id"])
assert job["state"] == J.PENDING and job["label"] == "first"
assert job["payload"] == {"model": "m"}
def test_unknown_tenant_is_rejected(self, queue):
assert J.submit("nope", {})["success"] is False
def test_priority_defaults_to_the_tenants_own(self, queue):
import tenants as T
r = J.submit("comfyui", {})
assert r["priority"] == T.get_tenant("comfyui").priority
def test_pending_jobs_are_listed_in_execution_order(self, queue):
J.submit("ollama", {}, priority=10, label="low")
J.submit("ollama", {}, priority=90, label="high")
J.submit("ollama", {}, priority=50, label="mid")
order = [j["label"] for j in J.listing(J.PENDING)]
assert order == ["high", "mid", "low"]
def test_equal_priority_is_first_in_first_out(self, queue):
a = J.submit("ollama", {}, priority=50, label="a")["id"]
time.sleep(0.01)
J.submit("ollama", {}, priority=50, label="b")
assert J._next_job()["id"] == a
def test_the_queue_has_no_depth_limit(self, queue):
# "Unbounded" is the point; it lives on disk, not in memory.
for i in range(500):
J.submit("ollama", {}, label=f"j{i}")
assert J.stats()["queue_depth"] == 500
def test_queue_survives_a_restart(self, queue):
J.submit("ollama", {}, label="persisted")
J._cache = None # nothing in-process is holding it
assert [j["label"] for j in J.listing(J.PENDING)] == ["persisted"]
class TestCancellation:
def test_pending_jobs_can_be_cancelled(self, queue):
jid = J.submit("ollama", {})["id"]
assert J.cancel(jid)["success"] is True
assert J.get(jid)["state"] == J.CANCELLED
def test_running_work_is_never_cancelled(self, queue):
# This service frees VRAM by asking, never by killing work in flight.
jid = J.submit("ollama", {})["id"]
J._mark(jid, J.RUNNING, started_at=time.time())
assert J.cancel(jid)["success"] is False
assert J.get(jid)["state"] == J.RUNNING
def test_clearing_the_queue_leaves_running_work_alone(self, queue):
running = J.submit("ollama", {})["id"]
J._mark(running, J.RUNNING, started_at=time.time())
J.submit("ollama", {})
J.submit("ollama", {})
assert J.clear_pending()["cancelled"] == 2
assert J.get(running)["state"] == J.RUNNING
class TestOrphanRecovery:
def test_jobs_left_running_by_a_dead_process_are_requeued(self, queue):
"""RUNNING means "this process is working on it".
An exception during dispatch left the row RUNNING while the scheduler moved on,
so the job never finished and never retried.
"""
jid = J.submit("ollama", {})["id"]
J._mark(jid, J.RUNNING, started_at=time.time())
assert J.requeue_orphans() == 1
job = J.get(jid)
assert job["state"] == J.PENDING and job["started_at"] is None
def test_finished_jobs_are_untouched_by_recovery(self, queue):
done = J.submit("ollama", {})["id"]
J._mark(done, J.DONE, finished_at=time.time())
assert J.requeue_orphans() == 0
assert J.get(done)["state"] == J.DONE
class TestPerJobVramRequirement:
"""A tenant-wide figure cannot be right for an LLM."""
def test_llm_requirement_comes_from_the_model_being_loaded(self, queue, monkeypatch):
import vram_arbitrator
monkeypatch.setattr(vram_arbitrator, "_model_size_bytes",
lambda m: int(12.87 * 1024 ** 3))
s = J.Scheduler()
# 12.87 GB on disk occupies ~14.9 GB once context and KV cache are allocated.
assert 14.5 < s._job_vram_requirement("ollama", {"model": "big"}) < 15.5
def test_a_small_model_needs_correspondingly_less(self, queue, monkeypatch):
import vram_arbitrator
monkeypatch.setattr(vram_arbitrator, "_model_size_bytes",
lambda m: int(1.96 * 1024 ** 3))
s = J.Scheduler()
assert s._job_vram_requirement("ollama", {"model": "small"}) < 3.0
def test_falls_back_to_the_tenant_figure_when_the_model_is_unknown(self, queue, monkeypatch):
import vram_arbitrator
monkeypatch.setattr(vram_arbitrator, "_model_size_bytes", lambda m: 0)
s = J.Scheduler()
import tenants as T
assert (s._job_vram_requirement("ollama", {"model": "?"})
== T.get_tenant("ollama").needs_vram_gb)
def test_non_llm_tenants_use_their_declared_figure(self, queue):
import tenants as T
s = J.Scheduler()
assert (s._job_vram_requirement("comfyui", {})
== T.get_tenant("comfyui").needs_vram_gb)
class TestSchedulerStatus:
def test_stats_report_depth_and_the_oldest_wait(self, queue):
J.submit("ollama", {})
J.submit("comfyui", {})
st = J.stats()
assert st["queue_depth"] == 2
assert st["pending_by_tenant"] == {"ollama": 1, "comfyui": 1}
assert st["oldest_pending_s"] is not None
def test_status_includes_queue_stats(self, queue):
J.submit("ollama", {})
s = J.Scheduler().get_status()
assert s["queue_depth"] == 1 and s["running"] is False
def test_a_blocked_job_reports_why_it_is_waiting(self, queue):
# Silence here is what left three jobs pending indefinitely with no explanation.
s = J.Scheduler()
s.blocked = {"id": "x", "tenant": "ollama", "reason": "needs 14.93 GB",
"waited_s": 42.0, "since": time.time()}
assert "14.93" in s.get_status()["blocked"]["reason"]
class TestDispatchers:
def test_every_tenant_kind_that_can_run_work_has_a_dispatcher(self):
import tenants as T
assert T.KIND_LLM in J.DISPATCHERS
assert T.KIND_DIFFUSION in J.DISPATCHERS
def test_ollama_dispatch_reports_failure_rather_than_raising(self, queue, monkeypatch):
class _R:
status_code = 500
text = '{"error":"llama-server process has terminated: cudaMalloc failed"}'
class _C:
async def __aenter__(self): return self
async def __aexit__(self, *a): return False
async def post(self, url, json=None): return _R()
monkeypatch.setattr(J.httpx, "AsyncClient", lambda **k: _C())
res = asyncio.run(J._dispatch_ollama({"model": "m", "prompt": "hi"}))
assert res["ok"] is False and "cudaMalloc" in res["error"]

405
tests/test_tenants.py Normal file
View File

@@ -0,0 +1,405 @@
"""Tests for the GPU tenant registry.
The point of this service is fast handoff of one GPU between applications, and it should
work for any application -- not only the two it grew up around. Their names had ended up
compiled into process matching, VRAM attribution, busy detection and release calls alike.
These tests pin the properties that make the registry generic: adding an application is
configuration, and nothing in the arbitration logic knows a particular name.
"""
import asyncio
import json
import pytest
import tenants as T
@pytest.fixture
def cfg(tmp_path, monkeypatch):
path = tmp_path / "tenants.json"
monkeypatch.setattr(T, "CONFIG_PATH", str(path))
T._cache.update({"ts": 0.0, "tenants": None, "mtime": None})
return path
class TestProcessMatching:
def test_matches_by_process_name(self):
m = T.ProcessMatch(names=["ollama"])
assert m.matches("ollama", "/usr/bin/ollama serve")
assert not m.matches("python", "main.py")
def test_matches_by_cmdline_substring(self):
m = T.ProcessMatch(cmdline=["llama-server"])
assert m.matches("python", "/usr/local/lib/ollama/llama-server --model x")
def test_matches_by_cmdline_suffix(self):
# ComfyUI is a bare `python main.py`, with nothing else distinguishing it.
m = T.ProcessMatch(cmdline_endswith=["main.py"])
assert m.matches("python", "/opt/ComfyUI/venv/bin/python main.py")
assert not m.matches("python", "/opt/other/main.py --serve")
def test_matching_is_case_insensitive(self):
assert T.ProcessMatch(names=["Xorg"]).matches("XORG", "")
class TestDefaultsPreserveExistingBehaviour:
"""The shipped defaults must classify exactly as the hardcoded version did."""
@pytest.mark.parametrize("pname,cmdline,expected", [
("llama-server", "/usr/local/lib/ollama/llama-server --model x", "ollama"),
("ollama", "/usr/bin/ollama serve", "ollama"),
("python", "/home/u/ComfyUI/venv/bin/python main.py --listen", "comfyui"),
("gnome-shell", "/usr/bin/gnome-shell --mode=ubuntu", "desktop"),
("Xorg", "/usr/lib/xorg/Xorg :8", "desktop"),
("python", "/home/u/robopest-venv/bin/python /home/u/stt_relay.py", "unmanaged"),
("trainer", "/opt/ml/bin/trainer --epochs 3", "unmanaged"),
])
def test_classification(self, cfg, pname, cmdline, expected):
assert T.classify_process(pname, cmdline) == expected
def test_unknown_process_is_unmanaged_not_silently_owned(self, cfg):
# Misattributing a third party's VRAM to a tenant would make this service
# promise headroom it cannot deliver.
assert T.classify_process("weird", "/opt/x/weird --run") == "unmanaged"
class TestAddingAnApplicationIsConfiguration:
def test_a_new_tenant_is_recognised_without_code_changes(self, cfg):
cfg.write_text(json.dumps(T.DEFAULT_TENANTS + [{
"name": "trainer",
"kind": "other",
"priority": 80,
"match": {"cmdline": ["train.py"]},
"release": {"type": "http_post", "url": "http://localhost:9999/release"},
}]))
assert T.classify_process("python", "/opt/ml/train.py --epochs 3") == "trainer"
t = T.get_tenant("trainer")
assert t.priority == 80 and t.reclaimable
def test_first_run_writes_the_defaults(self, cfg):
assert not cfg.exists()
T.load_tenants(force=True)
assert cfg.exists()
assert {t["name"] for t in json.loads(cfg.read_text())} == {
"ollama", "comfyui", "desktop"}
def test_a_malformed_entry_is_skipped_not_fatal(self, cfg):
cfg.write_text(json.dumps([{"name": "ok", "match": {"names": ["a"]}},
{"no_name": True}]))
names = [t.name for t in T.load_tenants(force=True)]
assert names == ["ok"]
def test_corrupt_config_falls_back_to_defaults(self, cfg):
cfg.write_text("{ not json")
assert {t.name for t in T.load_tenants(force=True)} >= {"ollama", "comfyui"}
class TestReclaimability:
def test_a_tenant_with_no_release_strategy_is_not_reclaimable(self, cfg):
t = T.GpuTenant(name="x", release=T.ReleaseStrategy(type="none"))
assert t.reclaimable is False
def test_release_refuses_rather_than_reporting_success(self, cfg):
t = T.GpuTenant(name="x", release=T.ReleaseStrategy(type="none"))
res = asyncio.run(T.release_vram(t))
assert res["success"] is False and res["released"] is False
assert "no way to release" in res["reason"]
def test_per_model_release_with_nothing_loaded_is_a_no_op(self, cfg):
t = T.GpuTenant(name="ollama", release=T.ReleaseStrategy(
type="http_post", url="http://x/api", per_model=True))
res = asyncio.run(T.release_vram(t, models=[]))
assert res["success"] is True and res["released"] is False
class TestBusyProbe:
def _probe(self, monkeypatch, payload, status=200):
class _R:
status_code = status
def json(self_inner): return payload
class _C:
async def __aenter__(self): return self
async def __aexit__(self, *a): return False
async def get(self, url): return _R()
monkeypatch.setattr(T.httpx, "AsyncClient", lambda **k: _C())
def test_empty_queue_is_not_busy(self, monkeypatch):
self._probe(monkeypatch, {"queue_running": [], "queue_pending": []})
t = T.GpuTenant(name="c", busy=T.BusyProbe(
type="http_count", url="http://x/queue",
count_keys=["queue_running", "queue_pending"]))
assert asyncio.run(T.probe_busy(t))["busy"] is False
def test_queued_work_while_holding_no_vram_is_flagged_below_floor(self, monkeypatch):
# ComfyUI leaves dead jobs in queue_running; only its VRAM reveals that nothing
# is loaded.
self._probe(monkeypatch, {"queue_running": [[1, "abc"]], "queue_pending": []})
t = T.GpuTenant(name="c", busy=T.BusyProbe(
type="http_count", url="http://x/queue", count_keys=["queue_running"],
vram_floor_gb=1.5))
res = asyncio.run(T.probe_busy(t, vram_gb=0.56))
assert res["busy"] is True and res.get("below_floor") is True
def test_queued_work_with_a_checkpoint_loaded_is_plainly_busy(self, monkeypatch):
self._probe(monkeypatch, {"queue_running": [[1, "abc"]], "queue_pending": []})
t = T.GpuTenant(name="c", busy=T.BusyProbe(
type="http_count", url="http://x/queue", count_keys=["queue_running"],
vram_floor_gb=1.5))
res = asyncio.run(T.probe_busy(t, vram_gb=6.8))
assert res["busy"] is True and not res.get("below_floor")
def test_vram_probe_needs_no_http_endpoint(self):
# An application with no API can still be observed by what it holds.
t = T.GpuTenant(name="x", busy=T.BusyProbe(type="vram", vram_busy_gb=1.0))
assert asyncio.run(T.probe_busy(t, vram_gb=2.0))["busy"] is True
assert asyncio.run(T.probe_busy(t, vram_gb=0.5))["busy"] is False
def test_an_unreachable_probe_reports_not_busy_rather_than_raising(self, monkeypatch):
class _C:
async def __aenter__(self): return self
async def __aexit__(self, *a): return False
async def get(self, url): raise ConnectionError("refused")
monkeypatch.setattr(T.httpx, "AsyncClient", lambda **k: _C())
t = T.GpuTenant(name="c", busy=T.BusyProbe(type="http_count", url="http://x",
count_keys=["q"]))
res = asyncio.run(T.probe_busy(t))
assert res["busy"] is False and "failed" in res["reason"]
class TestPriority:
def test_describe_orders_by_priority(self, cfg):
rows = T.describe()
prios = [r["priority"] for r in rows]
assert prios == sorted(prios, reverse=True)
assert all("reclaimable" in r for r in rows)
class TestReleasePlanning:
"""Deciding who gives up VRAM, generically over any number of applications.
The two-application version was a pair of hardcoded rules -- yield Ollama when
ComfyUI is busy, purge ComfyUI when Ollama is starved -- which could not express a
third participant at all.
"""
def _state(self, **overrides):
base = [
{"name": "desktop", "priority": 90, "vram_gb": 0.01, "busy": False,
"reclaimable": False},
{"name": "stt-relay", "priority": 70, "vram_gb": 0.8, "busy": False,
"reclaimable": False},
# Shipped priorities: diffusion outranks the LLM, whose weights reload
# from page cache in seconds.
{"name": "comfyui", "priority": 60, "vram_gb": 7.0, "busy": False,
"reclaimable": True},
{"name": "ollama", "priority": 50, "vram_gb": 0.0, "busy": True,
"reclaimable": True},
]
for s in base:
s.update(overrides.get(s["name"], {}))
return base
def test_a_tenant_already_holding_what_it_needs_is_not_starved(self):
# A busy GPU has little free by definition. Comparing free VRAM alone flagged a
# tenant working fine on 13 GB as demanding, which would have caused pointless
# releases from everything else.
state = self._state(ollama={"vram_gb": 13.0})
plan = T.plan_release("ollama", state, free_gb=1.5, needed_gb=4.0)
assert plan["release"] == []
assert "already free" in plan["reason"]
def test_starved_tenant_reclaims_from_the_idle_one_below_it(self):
plan = T.plan_release("ollama", self._state(), free_gb=1.5, needed_gb=14.9)
assert plan["release"] == ["comfyui"]
def test_a_busy_tenant_ranking_above_the_demander_is_not_a_victim(self):
# comfyui outranks ollama, so ollama may not interrupt it.
state = self._state(comfyui={"busy": True})
plan = T.plan_release("ollama", state, free_gb=1.5, needed_gb=14.9)
assert plan["release"] == []
assert any(b["name"] == "comfyui" and "busy" in b["why"]
for b in plan["blockers"])
def test_a_higher_priority_demander_preempts_busy_lower_priority_work(self):
"""The measured regression that made this rule necessary.
Refusing to touch anything busy looks safe and is not. With the LLM protected as
"busy", a diffusion job ran 46 s instead of 3 s, squeezed into 1.6 GB, because
the LLM reloaded immediately after yielding and was then untouchable. Preempting
a lower-priority tenant is safe because releasing is asynchronous: an Ollama
unload queues behind its running request rather than killing it.
"""
state = [
{"name": "comfyui", "priority": 60, "vram_gb": 1.65, "busy": True,
"reclaimable": True},
{"name": "ollama", "priority": 50, "vram_gb": 13.03, "busy": True,
"reclaimable": True},
]
plan = T.plan_release("comfyui", state, free_gb=0.28, needed_gb=6.0)
assert plan["release"] == ["ollama"]
def test_an_idle_tenant_is_preferred_over_preempting_a_busy_one(self):
state = [
{"name": "d", "priority": 60, "vram_gb": 0.0, "busy": True,
"reclaimable": True},
{"name": "busy_low", "priority": 10, "vram_gb": 8.0, "busy": True,
"reclaimable": True},
{"name": "idle_high", "priority": 90, "vram_gb": 8.0, "busy": False,
"reclaimable": True},
]
plan = T.plan_release("d", state, free_gb=0.0, needed_gb=8.0)
assert plan["release"] == ["idle_high"]
def test_unreclaimable_tenants_are_named_as_blockers_not_ignored(self):
# The user needs to know a third-party process is what stands in the way.
plan = T.plan_release("ollama", self._state(), free_gb=0.0, needed_gb=15.5)
blockers = {b["name"]: b["why"] for b in plan["blockers"]}
assert blockers["stt-relay"] == "declares no release mechanism"
assert plan["possible"] is False
def test_an_idle_tenant_yields_even_if_it_outranks_the_demander(self):
"""Priority orders who is asked first; it does not protect idle memory.
Filtering candidates by priority broke both directions in turn: with the LLM
ranked above diffusion, ComfyUI could never preempt Ollama -- the service's
central behaviour -- and once the ranks were swapped, a starved Ollama could no
longer reclaim from an idle ComfyUI. An idle tenant is not using its VRAM, so
outranking the demander is not a reason to keep it.
"""
state = self._state(comfyui={"priority": 99, "busy": False})
plan = T.plan_release("ollama", state, free_gb=1.0, needed_gb=14.9)
assert plan["release"] == ["comfyui"]
def test_priority_decides_who_is_asked_first(self):
state = [
{"name": "demander", "priority": 50, "vram_gb": 0.0, "busy": True,
"reclaimable": True},
{"name": "high", "priority": 90, "vram_gb": 4.0, "busy": False,
"reclaimable": True},
{"name": "low", "priority": 10, "vram_gb": 4.0, "busy": False,
"reclaimable": True},
]
plan = T.plan_release("demander", state, free_gb=0.0, needed_gb=5.0)
# The lowest-priority idle tenant gives up memory first.
assert plan["release"][0] == "low"
def test_peers_cannot_interrupt_each_other(self):
# Equal priority is never preempted, so two tenants at the same rank cannot
# fight over the card.
state = [
{"name": "a", "priority": 50, "vram_gb": 0.0, "busy": True,
"reclaimable": True},
{"name": "b", "priority": 50, "vram_gb": 8.0, "busy": True,
"reclaimable": True},
]
plan = T.plan_release("a", state, free_gb=0.0, needed_gb=8.0)
assert plan["release"] == []
assert "busy" in plan["blockers"][0]["why"]
def test_lowest_priority_is_released_first(self):
state = self._state() + [
{"name": "batch", "priority": 10, "vram_gb": 3.0, "busy": False,
"reclaimable": True}]
plan = T.plan_release("ollama", state, free_gb=0.0, needed_gb=5.0)
assert plan["release"][0] == "batch"
def test_releases_only_as_many_tenants_as_needed(self):
state = self._state() + [
{"name": "batch", "priority": 10, "vram_gb": 9.0, "busy": False,
"reclaimable": True}]
plan = T.plan_release("ollama", state, free_gb=0.0, needed_gb=8.0)
assert plan["release"] == ["batch"] # 9 GB covers it; comfyui is left alone
def test_three_applications_can_all_participate(self):
# The property the hardcoded pair of rules could not express.
state = [
{"name": "llm", "priority": 60, "vram_gb": 0.0, "busy": True,
"reclaimable": True},
{"name": "diffusion", "priority": 50, "vram_gb": 4.0, "busy": False,
"reclaimable": True},
{"name": "trainer", "priority": 40, "vram_gb": 5.0, "busy": False,
"reclaimable": True},
]
plan = T.plan_release("llm", state, free_gb=0.0, needed_gb=9.0)
assert set(plan["release"]) == {"trainer", "diffusion"}
assert plan["possible"] is True
def test_unknown_tenant_is_rejected_cleanly(self):
plan = T.plan_release("nope", self._state(), free_gb=0.0, needed_gb=1.0)
assert plan["possible"] is False and plan["release"] == []
class TestConfigUpgrade:
def test_fields_added_later_are_merged_into_an_existing_config(self, cfg):
# A config written before needs_vram_gb existed must not silently lose the
# behaviour that field controls.
cfg.write_text(json.dumps([{
"name": "ollama",
"match": {"names": ["ollama"]},
}]))
t = T.get_tenant("ollama")
assert t.needs_vram_gb > 0
assert t.release.type == "http_post"
def test_explicit_user_values_still_win_over_defaults(self, cfg):
cfg.write_text(json.dumps([{
"name": "ollama", "priority": 5, "needs_vram_gb": 99.0,
"match": {"names": ["ollama"]},
}]))
t = T.get_tenant("ollama")
assert t.priority == 5 and t.needs_vram_gb == 99.0
class TestVramFloor:
"""VRAM that survives a release must not be promised to anyone else.
ComfyUI keeps its CUDA context for as long as the process lives, so a purge does not
return everything it holds. Ignoring that made plan_release report it would free
0.37 GB against a 0.33 GB shortfall; the job was cleared to run and the memory never
arrived, so it waited two minutes and then failed.
"""
def test_only_memory_above_the_floor_counts_as_freeable(self):
state = [
{"name": "llm", "priority": 50, "vram_gb": 0.0, "busy": True,
"reclaimable": True, "vram_floor_gb": 0.0},
{"name": "diffusion", "priority": 60, "vram_gb": 0.44, "busy": False,
"reclaimable": True, "vram_floor_gb": 0.45},
]
plan = T.plan_release("llm", state, free_gb=14.6, needed_gb=14.93)
assert plan["possible"] is False
assert plan["release"] == []
def test_a_loaded_checkpoint_is_still_freeable_above_its_floor(self):
state = [
{"name": "llm", "priority": 50, "vram_gb": 0.0, "busy": True,
"reclaimable": True, "vram_floor_gb": 0.0},
{"name": "diffusion", "priority": 60, "vram_gb": 7.0, "busy": False,
"reclaimable": True, "vram_floor_gb": 0.45},
]
plan = T.plan_release("llm", state, free_gb=7.9, needed_gb=14.0)
assert plan["release"] == ["diffusion"]
# 7.0 held minus a 0.45 floor.
assert abs(plan["would_free_gb"] - 6.55) < 0.01
def test_a_tenant_at_its_floor_is_not_even_listed_for_release(self):
state = [
{"name": "llm", "priority": 50, "vram_gb": 0.0, "busy": True,
"reclaimable": True, "vram_floor_gb": 0.0},
{"name": "at_floor", "priority": 10, "vram_gb": 0.3, "busy": False,
"reclaimable": True, "vram_floor_gb": 0.45},
{"name": "has_room", "priority": 20, "vram_gb": 5.0, "busy": False,
"reclaimable": True, "vram_floor_gb": 0.0},
]
plan = T.plan_release("llm", state, free_gb=0.0, needed_gb=4.0)
assert plan["release"] == ["has_room"]
def test_default_floor_is_zero_so_existing_configs_are_unchanged(self):
state = [
{"name": "a", "priority": 50, "vram_gb": 0.0, "busy": True,
"reclaimable": True},
{"name": "b", "priority": 40, "vram_gb": 5.0, "busy": False,
"reclaimable": True},
]
plan = T.plan_release("a", state, free_gb=0.0, needed_gb=5.0)
assert plan["release"] == ["b"] and plan["possible"] is True

View File

@@ -13,6 +13,7 @@ were observed on real hardware before being encoded here:
out of memory ... unable to allocate CUDA0 buffer"
"""
import asyncio
import time
import pytest
@@ -132,3 +133,70 @@ class TestBusyBackoff:
assert "yield_deferred_busy" in arb.stats
assert "yield_stalled" in arb.stats
assert "yield_timeouts" not in arb.stats
class TestComfyStaleQueueDetection:
"""ComfyUI can leave a dead job in queue_running forever.
Observed on this machine: a WAN 2.1 i2v entry sat in queue_running while the GPU was
idle and ComfyUI held 0.56 GB. Trusting that flag made the watchdog believe ComfyUI
was permanently busy, so it evicted the LLM on every poll, never ran the idle purge,
and never checked whether the LLM had been pushed onto the CPU. Instrumenting the
watchdog showed busy=6, idle_check=0 -- one stale row had disabled half the logic.
"""
def _arb(self, comfy_bytes):
arb = v.AutoArbitrator()
v.get_process_vram_bytes = lambda: {
"ollama_bytes": 0, "comfyui_bytes": int(comfy_bytes), "other_bytes": 0,
"desktop_bytes": 0, "unmanaged_bytes": 0, "free_bytes": 0, "gpu_util_pct": 0}
return arb
def teardown_method(self):
import importlib
importlib.reload(v)
def test_empty_queue_is_not_busy(self):
arb = self._arb(0)
assert arb._comfy_genuinely_busy({"queue_running": [], "queue_pending": []}) is False
def test_pending_work_is_always_busy(self):
arb = self._arb(0)
assert arb._comfy_genuinely_busy(
{"queue_running": [], "queue_pending": [[1, "p"]]}) is True
def test_a_running_job_is_believed_at_first(self):
# It must not be called stale before it has had time to load anything.
arb = self._arb(0.1 * GB)
assert arb._comfy_genuinely_busy(
{"queue_running": [[1, "abc"]], "queue_pending": []}) is True
def test_long_running_job_holding_no_vram_is_stale(self):
arb = self._arb(0.56 * GB) # the observed CUDA-context floor
q = {"queue_running": [[1, "abc"]], "queue_pending": []}
arb._comfy_genuinely_busy(q)
arb._running_since = time.time() - (arb.STALE_RUNNING_S + 5)
assert arb._comfy_genuinely_busy(q) is False
assert arb.comfy_stale_job == "abc"
def test_long_running_job_holding_a_checkpoint_is_real(self):
# 6.8 GB is a loaded SDXL checkpoint; slow is not the same as stuck.
arb = self._arb(6.8 * GB)
q = {"queue_running": [[1, "abc"]], "queue_pending": []}
arb._comfy_genuinely_busy(q)
arb._running_since = time.time() - (arb.STALE_RUNNING_S + 5)
assert arb._comfy_genuinely_busy(q) is True
assert arb.comfy_stale_job is None
def test_a_new_prompt_id_resets_the_staleness_clock(self):
arb = self._arb(0.5 * GB)
arb._comfy_genuinely_busy({"queue_running": [[1, "old"]], "queue_pending": []})
arb._running_since = time.time() - 1000
assert arb._comfy_genuinely_busy(
{"queue_running": [[1, "new"]], "queue_pending": []}) is True
def test_vram_not_utilisation_is_the_signal(self):
# Utilisation is shared with Ollama and any third-party process, so it stayed
# above every sensible threshold and a stuck entry never looked stale.
assert hasattr(v.AutoArbitrator, "STALE_COMFY_BYTES")
assert not hasattr(v.AutoArbitrator, "STALE_UTIL_PCT")

338
verify_arbitration.py Executable file
View File

@@ -0,0 +1,338 @@
#!/usr/bin/env python
"""End-to-end verification of the arbitration cycle, against real hardware.
The unit suite covers logic in isolation; this exercises the promise the whole service
exists to make -- that an LLM and a diffusion pipeline can share one 16 GB card without
either failing -- and reports what actually happened at each stage.
It is deliberately not part of `pytest tests/`: it loads real models, runs a real
diffusion graph and moves real VRAM, taking a few minutes. Run it when you want proof
the system works on this machine:
python verify_arbitration.py # full cycle
python verify_arbitration.py --quick # skip the diffusion stages
Every stage restores what it changed, and the script refuses to start if ComfyUI is
already busy.
"""
import argparse
import asyncio
import sys
import time
from typing import Any, Dict, List, Optional
import httpx
BASE = "http://localhost:9090"
PASS, FAIL, SKIP, WARN = "PASS", "FAIL", "SKIP", "WARN"
_results: List[Dict[str, Any]] = []
def record(stage: str, status: str, detail: str, evidence: str = "") -> None:
_results.append({"stage": stage, "status": status, "detail": detail,
"evidence": evidence})
colour = {"PASS": "\033[32m", "FAIL": "\033[31m",
"SKIP": "\033[33m", "WARN": "\033[33m"}[status]
print(f" {colour}{status:<4}\033[0m {stage:<38} {detail}")
if evidence:
print(f" {evidence}")
async def api(client: httpx.AsyncClient, method: str, path: str,
allow_error: bool = False, **kw) -> Any:
r = await client.request(method, f"{BASE}{path}", **kw)
if allow_error:
# Some stages deliberately provoke a failure and need to read it.
body = r.json() if r.headers.get("content-type", "").startswith("application/json") else {}
return {"_status": r.status_code, **(body if isinstance(body, dict) else {})}
r.raise_for_status()
return r.json()
async def stage_preflight(c: httpx.AsyncClient) -> Optional[Dict[str, Any]]:
health = await api(c, "GET", "/api/health")
failed = health.get("failed") or []
if failed:
record("preflight: dependencies", FAIL,
f"{len(failed)} dependency check(s) failing", ", ".join(failed))
return None
record("preflight: dependencies", PASS, health["summary"])
stats = await api(c, "GET", "/api/stats")
if stats["comfyui"].get("executing") or stats["comfyui"].get("queue_remaining"):
record("preflight: ComfyUI idle", FAIL, "ComfyUI is busy; refusing to interfere")
return None
record("preflight: ComfyUI idle", PASS, "queue empty")
return stats
async def stage_llm_load(c: httpx.AsyncClient, model: str) -> bool:
res = await api(c, "POST", "/api/switch-model",
json={"model": model, "keep_alive": "5m"}, timeout=300)
if not res.get("success"):
record("LLM loads", FAIL, res.get("error", "")[:90])
return False
gbps, status = res.get("load_gbps"), res.get("cache_status")
record("LLM loads", PASS, f"{res['model_size_gb']} GB in {res['load_duration_ms']:.0f} ms",
f"{gbps} GB/s -> {status}; {res['tokens_per_sec']} tok/s")
# The classification must follow from the measured bandwidth, not a fixed duration.
if gbps is not None:
expected = ("RAM Cache Hit" if gbps >= 2.0
else "Partial Cache" if gbps >= 0.8 else "Cold Disk Load")
ok = expected.split()[0] in (status or "")
record("load classified by bandwidth", PASS if ok else FAIL,
f"{gbps} GB/s reported as '{status}'",
"" if ok else f"expected something matching '{expected}'")
return True
async def stage_yield(c: httpx.AsyncClient) -> bool:
before = await api(c, "GET", "/api/gpu")
held = before["breakdown"]["ollama_gb"]
res = await api(c, "POST", "/api/free-vram", timeout=60)
outcome = res.get("outcome")
if outcome == "released":
after = await api(c, "GET", "/api/gpu")
freed = held - after["breakdown"]["ollama_gb"]
# The barrier's promise: on return, the VRAM is genuinely gone.
ok = after["breakdown"]["ollama_gb"] < 0.3
record("VRAM yield is confirmed", PASS if ok else FAIL,
f"released in {res['confirm_ms']} ms",
f"{freed:.2f} GB freed; NVML now reports "
f"{after['breakdown']['ollama_gb']} GB held by Ollama")
return ok
if outcome == "busy":
record("VRAM yield is confirmed", WARN,
"model was mid-generation, so the unload was deferred",
res.get("error", ""))
return True
record("VRAM yield is confirmed", FAIL, f"outcome={outcome}", res.get("error", ""))
return False
async def stage_diffusion(c: httpx.AsyncClient) -> bool:
sys.path.insert(0, "/home/drjones/unified-model-manager")
import autotune, vram_arbitrator # noqa: E402 (imported late; needs the service's deps)
# The first run loads the checkpoint from disk. Timing that and calling the result
# "it/s" understates throughput by roughly 10x -- 0.67 it/s against a steady-state
# 6.7 -- so the load is measured separately and reported as what it is.
first = await autotune._diffusion_benchmark()
if not first.get("ok"):
record("diffusion runs", FAIL, first.get("error", "")[:90])
return False
record("diffusion runs (cold, includes checkpoint load)", PASS,
f"{first['exec_ms']} ms", f"{first['it_per_sec']} it/s including load")
res = await autotune._diffusion_benchmark()
if not res.get("ok"):
record("diffusion runs (warm)", FAIL, res.get("error", "")[:90])
return False
record("diffusion throughput (warm)", PASS,
f"SDXL 1024/20 steps in {res['exec_ms']} ms", f"{res['it_per_sec']} it/s")
snap = vram_arbitrator.get_process_vram_bytes()
comfy_gb = snap["comfyui_bytes"] / (1024 ** 3)
record("ComfyUI holds its checkpoint", PASS if comfy_gb > 0.5 else WARN,
f"{comfy_gb:.2f} GB retained",
"held for the idle window rather than purged between iterations")
return True
async def stage_idle_purge(c: httpx.AsyncClient) -> bool:
# The completion event arrives over the ComfyUI websocket, so the flag is set a
# moment after the graph returns. Checking instantly raced it.
arb = {}
for _ in range(12):
arb = (await api(c, "GET", "/api/stats"))["arbitrator"]
if arb.get("pending_purge"):
break
await asyncio.sleep(0.5)
if not arb.get("pending_purge"):
record("purge is deferred, not immediate", WARN,
"no purge pending after 6 s (ComfyUI may already be clean)")
return True
idle_s = arb.get("comfy_idle_s")
record("purge is deferred, not immediate", PASS,
f"holding checkpoints for {arb.get('idle_purge_after_s')} s",
f"idle {idle_s} s so far" if idle_s is not None
else "idle timer just started")
return True
async def stage_reclaim(c: httpx.AsyncClient, model: str) -> bool:
"""The direction that used to fail outright: an LLM that will not fit.
This only proves anything if the chosen model genuinely cannot fit in what ComfyUI
has left free. A small model fits alongside the checkpoint and the stage passes
without exercising the reclaim path at all, so pick the largest model that will not
fit and say plainly when no such model exists.
"""
gpu = await api(c, "GET", "/api/gpu")
comfy_gb = gpu["breakdown"]["comfyui_gb"]
free_gb = gpu["vram_free_gb"]
if comfy_gb < 0.5:
record("reclaims VRAM for the LLM", SKIP,
f"ComfyUI only holds {comfy_gb} GB; nothing to reclaim")
return True
models = (await api(c, "GET", "/api/models"))["ollama_models"]
EMBED = {"bert", "nomic-bert", "gte", "jina-bert"}
usable = [m for m in models
if (m.get("details", {}).get("family") or "").lower() not in EMBED
and "embed" not in m["name"].lower()]
# On-disk weight size is not the VRAM footprint: measured on this box, a 12.87 GB
# blob occupies 14.9 GB once context and KV cache are allocated. Sizing the test off
# disk size picks a model that cannot fit even after a successful reclaim.
VRAM_OVERHEAD = 1.18
HEADROOM_GB = 0.4
def vram_need(m):
return m.get("size", 0) / (1024 ** 3) * VRAM_OVERHEAD
# Unmanaged VRAM never comes back, so it is not part of what a reclaim can offer.
# Ignoring it picked a model that failed even after a correct reclaim -- on this box
# an 842 MB third-party process is the difference between a 14.9 GB model fitting
# and not.
reclaimable_gb = free_gb + comfy_gb - HEADROOM_GB
too_big = [m for m in usable
if vram_need(m) > free_gb and vram_need(m) < reclaimable_gb]
if too_big:
target = max(too_big, key=lambda m: m.get("size", 0))
model = target["name"]
print(f" using {model} ({target['size'] / (1024**3):.1f} GB on disk, "
f"~{vram_need(target):.1f} GB in VRAM) — will not fit in "
f"{free_gb:.1f} GB free, should fit after reclaiming {comfy_gb:.1f} GB")
else:
unmanaged = gpu["breakdown"].get("unmanaged_gb", 0)
record("reclaims VRAM for the LLM", SKIP,
f"no installed model needs between {free_gb:.1f} and "
f"{reclaimable_gb:.1f} GB of VRAM",
f"reclaimable ceiling excludes {unmanaged} GB held by processes "
f"HyperSwap cannot free")
return True
# Re-run a graph first. The idle purge fires 30 s after ComfyUI goes quiet, and a
# large model takes longer than that to load -- so without resetting the timer the
# purge frees ComfyUI mid-load and the reclaim path is never reached.
sys.path.insert(0, "/home/drjones/unified-model-manager")
import autotune # noqa: E402
await autotune._diffusion_benchmark()
gpu = await api(c, "GET", "/api/gpu")
print(f" reset the idle window; ComfyUI holds "
f"{gpu['breakdown']['comfyui_gb']} GB, {gpu['vram_free_gb']} GB free")
res = await api(c, "POST", "/api/switch-model", allow_error=True,
json={"model": model, "keep_alive": "2m"}, timeout=600)
if res.get("_status") == 507:
record("reclaims VRAM for the LLM", FAIL,
"reclaim ran but the model still did not fit",
str(res.get("detail", ""))[:150])
return False
if res.get("_status", 200) >= 400 or not res.get("success"):
record("reclaims VRAM for the LLM", FAIL,
f"HTTP {res.get('_status')}", str(res.get("detail", ""))[:120])
return False
if res.get("_status") and res.get("_status") != 200:
pass
if res.get("reclaimed_from_comfyui_gb"):
record("reclaims VRAM for the LLM", PASS,
f"reclaimed {res['reclaimed_from_comfyui_gb']} GB and retried",
f"'{model}' then loaded at {res.get('load_gbps')} GB/s")
else:
# It fit anyway, so nothing was proven; do not report that as a pass. The usual
# cause is the idle purge firing during the load and freeing ComfyUI first.
record("reclaims VRAM for the LLM", WARN,
"model fit without a reclaim, so the path was not exercised",
"the idle purge most likely freed ComfyUI during the load; "
f"loaded at {res.get('load_gbps')} GB/s")
return True
async def stage_accounting(c: httpx.AsyncClient) -> bool:
"""Reported VRAM must add up, and reported settings must match the hardware."""
gpu = await api(c, "GET", "/api/gpu")
bd = gpu["breakdown"]
parts = bd["ollama_gb"] + bd["comfyui_gb"] + bd["system_gb"]
used = gpu["vram_used_gb"]
# Driver overhead means the parts never sum exactly; a large gap means mis-attribution.
ok = abs(parts - used) < 1.5
record("VRAM attribution adds up", PASS if ok else FAIL,
f"parts {parts:.2f} GB vs NVML used {used:.2f} GB",
f"ollama {bd['ollama_gb']} + comfy {bd['comfyui_gb']} + system {bd['system_gb']} "
f"(desktop {bd.get('desktop_gb')}, unmanaged {bd.get('unmanaged_gb')})")
oc = await api(c, "GET", "/api/overclock")
drift = oc["drift"]
ok2 = not drift["drifted"]
record("reported GPU state matches hardware", PASS if ok2 else FAIL,
f"profile '{drift['profile']}' asks {drift['power_limit_intended_w']} W, "
f"card reports {drift['power_limit_actual_w']} W",
drift.get("reason") or "")
return ok and ok2
async def main() -> int:
ap = argparse.ArgumentParser(description=__doc__,
formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument("--quick", action="store_true",
help="skip the diffusion and reclaim stages")
ap.add_argument("--model", default=None,
help="Ollama model to test with (default: smallest installed)")
args = ap.parse_args()
print("\nHyperSwap arbitration verification\n" + "=" * 62)
async with httpx.AsyncClient(timeout=60.0) as c:
baseline = await stage_preflight(c)
if baseline is None:
print("\nPreflight failed; not continuing.\n")
return 2
model = args.model
if not model:
models = (await api(c, "GET", "/api/models"))["ollama_models"]
# Embedding models have no /api/generate endpoint, and the smallest model
# installed is often one of them.
EMBED_FAMILIES = {"bert", "nomic-bert", "gte", "jina-bert"}
usable = [m for m in models
if (m.get("details", {}).get("family") or "").lower()
not in EMBED_FAMILIES and "embed" not in m["name"].lower()]
if not usable:
record("choose a test model", FAIL,
"no generative Ollama models installed "
f"({len(models)} found, all embedding-only)")
return 2
model = min(usable, key=lambda m: m.get("size", 0))["name"]
print(f"\n using model: {model}\n")
await stage_llm_load(c, model)
await stage_yield(c)
if not args.quick:
if await stage_diffusion(c):
await stage_idle_purge(c)
await stage_reclaim(c, model)
else:
record("diffusion stages", SKIP, "--quick")
await stage_accounting(c)
# Leave the box as we found it.
await api(c, "POST", "/api/free-vram", timeout=60)
failed = [r for r in _results if r["status"] == FAIL]
warned = [r for r in _results if r["status"] == WARN]
print("=" * 62)
print(f" {len(_results) - len(failed) - len(warned)} passed, "
f"{len(warned)} warned, {len(failed)} failed")
for r in failed:
print(f" FAILED: {r['stage']} — {r['detail']}")
print()
return 1 if failed else 0
if __name__ == "__main__":
sys.exit(asyncio.run(main()))

View File

@@ -12,6 +12,7 @@ import websockets
import overclock_manager
import ram_optimizer
import tenants as tenants_mod
import telemetry_store
try:
@@ -72,7 +73,8 @@ RECLAIM_MIN_COMFY_BYTES = 512 * 1024 ** 2
# depends on configuration: with n_gpu_layers left to Ollama it spills layers to the CPU
# and reports size_vram < size; with n_gpu_layers pinned (99 on this box) it refuses and
# returns a hard CUDA OOM instead. Both are handled -- the spill by
# AutoArbitrator._check_ollama_starved, the hard failure by the retry below.
# AutoArbitrator._arbitrate (generically, from the tenant registry), the hard failure by
# the retry below.
OOM_SIGNATURES = ("out of memory", "cudamalloc", "unable to allocate",
"failed to allocate", "cuda error")
@@ -149,7 +151,8 @@ def get_process_vram_bytes() -> Dict[str, int]:
to actually drain.
"""
out = {"ollama_bytes": 0, "comfyui_bytes": 0, "other_bytes": 0, "free_bytes": 0,
"desktop_bytes": 0, "unmanaged_bytes": 0, "gpu_util_pct": 0}
"desktop_bytes": 0, "unmanaged_bytes": 0, "gpu_util_pct": 0,
"by_tenant_bytes": {}}
if not NVML_AVAILABLE:
return out
try:
@@ -176,6 +179,7 @@ def get_process_vram_bytes() -> Dict[str, int]:
if len(_PID_KIND_CACHE) >= _PID_KIND_CACHE_MAX:
_PID_KIND_CACHE.clear()
_PID_KIND_CACHE[key] = kind
out["by_tenant_bytes"][kind] = out["by_tenant_bytes"].get(kind, 0) + used
if kind == "ollama":
out["ollama_bytes"] += used
elif kind == "comfy":
@@ -198,6 +202,19 @@ _PID_KIND_CACHE: Dict[tuple, str] = {}
_PID_KIND_CACHE_MAX = 512
def _pid_key(pid: int) -> Optional[tuple]:
try:
return (pid, psutil.Process(pid).create_time())
except Exception:
return None
# Tenant names as used by this module's buckets. The tenant registry is the source of
# truth for *which* application a process belongs to; these two names are kept because
# the REST payloads and the dashboard have used them since the beginning.
_BUCKET_ALIASES = {"comfyui": "comfy"}
def _pid_key(pid: int) -> Optional[tuple]:
try:
return (pid, psutil.Process(pid).create_time())
@@ -214,26 +231,16 @@ DESKTOP_PROCESS_HINTS = (
def _classify_pid(pid: int) -> str:
"""Bucket a GPU process into ollama | comfy | desktop | unmanaged.
"""Which tenant owns this GPU process.
The old version had one catch-all "other" bucket, which put a 3.9 MB compositor and
an 842 MB long-running inference script in the same number. That matters: this
service can reclaim VRAM from ComfyUI, but it cannot touch a third-party workload,
and pretending otherwise makes it promise headroom it cannot deliver.
The matching rules used to be substrings compiled into this function, which made the
two applications on this box part of the arbitrator rather than input to it. They now
come from the tenant registry, so a third application is a config entry.
"unmanaged" still means something specific and useful: VRAM held by something with no
declared way to release it, and therefore headroom this service can never offer.
"""
try:
proc = psutil.Process(pid)
pname = proc.name().lower()
cmdline = " ".join(proc.cmdline()).lower()
except Exception:
return "unmanaged"
if "ollama" in pname or "llama-server" in cmdline:
return "ollama"
if "comfyui" in cmdline or "comfy" in cmdline or cmdline.rstrip().endswith("main.py"):
return "comfy"
if any(hint in pname or hint in cmdline for hint in DESKTOP_PROCESS_HINTS):
return "desktop"
return "unmanaged"
return _BUCKET_ALIASES.get(tenants_mod.classify_pid(pid), tenants_mod.classify_pid(pid))
def get_gpu_hardware_stats() -> Dict[str, Any]:
@@ -336,7 +343,10 @@ def get_gpu_hardware_stats() -> Dict[str, Any]:
"desktop_bytes": 0,
"unmanaged_bytes": 0,
"unmanaged": [],
"processes": []
"processes": [],
# Generic attribution: one entry per tenant, so an application added to the
# registry is reported without any change here.
"by_tenant": {},
}
try:
@@ -376,6 +386,8 @@ def get_gpu_hardware_stats() -> Dict[str, Any]:
"vram_mb": round(used_mem / (1024**2), 1),
})
proc_breakdown["by_tenant"][kind] = (
proc_breakdown["by_tenant"].get(kind, 0) + used_mem)
proc_breakdown["processes"].append({
"pid": pid,
"name": pname,
@@ -431,6 +443,8 @@ def get_gpu_hardware_stats() -> Dict[str, Any]:
# reclaimed, so it is permanently unavailable headroom.
"unmanaged_gb": round(proc_breakdown["unmanaged_bytes"] / (1024**3), 2),
"unmanaged": proc_breakdown["unmanaged"],
"by_tenant_gb": {k: round(b / (1024**3), 2)
for k, b in proc_breakdown["by_tenant"].items()},
"free_mb": round(free_vram / (1024**2), 1),
"free_gb": round(free_vram / (1024**3), 2),
"processes": proc_breakdown["processes"],
@@ -849,37 +863,46 @@ async def switch_ollama_model(target_model: str, keep_alive: str = "30m",
# idle ComfyUI and try once more.
body = resp.text
if looks_like_vram_oom(body) and not _retrying:
snap = get_process_vram_bytes()
if snap["comfyui_bytes"] >= RECLAIM_MIN_COMFY_BYTES:
# Which application should give up memory is a question for the registry,
# not something to answer by purging ComfyUI by name. Any reclaimable idle
# tenant below Ollama in priority is a candidate.
state = await arbitrator._tenant_state()
free_gb = arbitrator._last_tenant_state["free_gb"]
size_gb = _model_size_bytes(target_model) / (1024**3)
needed = size_gb * 1.16 if size_gb else free_gb + 1.0
plan = tenants_mod.plan_release("ollama", state, free_gb, needed)
if plan["release"]:
logger.warning(
f"Ollama could not fit '{target_model}' with ComfyUI holding "
f"{round(snap['comfyui_bytes'] / (1024**3), 2)} GB — reclaiming and retrying")
purge = await instant_free_comfyui_vram()
f"Ollama could not fit '{target_model}' — {plan['reason']}")
freed_before = free_gb
for victim in plan["release"]:
await arbitrator._release_tenant(
victim, f"Ollama could not load '{target_model}'")
arbitrator.stats["reclaims_for_ollama"] += 1
arbitrator.last_action = (
f"Reclaimed {round(snap['comfyui_bytes'] / (1024**3), 2)}GB from ComfyUI so "
f"'{target_model}' could load")
f"Released {', '.join(plan['release'])} so '{target_model}' could load")
_record({
"event_type": "VRAM Reclaim for Ollama",
"source": "ComfyUI Pipeline",
"source": ", ".join(plan["release"]),
"target": target_model,
"duration_ms": purge.get("duration_ms"),
"cache_status": "Reclaimed",
"detail": f"Ollama OOM: {body[:160]}",
})
await asyncio.sleep(0.3)
retry = await switch_ollama_model(target_model, keep_alive, _retrying=True)
retry["reclaimed_from_comfyui_gb"] = round(
snap["comfyui_bytes"] / (1024**3), 2)
retry["released_tenants"] = plan["release"]
retry["would_free_gb"] = plan.get("would_free_gb")
retry["first_attempt_error"] = "CUDA OOM; retried after reclaiming VRAM"
if not retry.get("success"):
# Be specific about why the reclaim was not enough. Blaming ComfyUI
# when a third-party process is holding the memory sends the user
# looking in the wrong place.
# Be specific about why the reclaim was not enough. Blaming a tenant
# when a process nobody can release is holding the memory sends the
# user looking in the wrong place.
retry["blockers"] = plan.get("blockers")
retry["unmanaged_blockers"] = describe_unmanaged()
return retry
return {"success": False, "error": f"HTTP {resp.status_code}: {body}",
"duration_ms": total_duration_ms,
"upstream_status": resp.status_code,
"vram_oom": looks_like_vram_oom(body)}
except Exception as e:
return {"success": False, "error": str(e),
@@ -911,6 +934,7 @@ class AutoArbitrator:
def __init__(self):
self.running = False
self.ws_task: Optional[asyncio.Task] = None
self.event_tasks: List[asyncio.Task] = []
self.poll_task: Optional[asyncio.Task] = None
self.idle_task: Optional[asyncio.Task] = None
self.last_yield_time = 0.0
@@ -930,6 +954,19 @@ class AutoArbitrator:
self._yield_backoff_until: Dict[str, float] = {}
self._yield_busy_streak: Dict[str, int] = {}
self.last_reclaim_time = 0.0
self.watchdog_branches = {"busy": 0, "completed": 0, "idle_check": 0,
"bad_status": 0, "error": 0}
self._running_id: Optional[str] = None
self._running_since: Optional[float] = None
self._peak_comfy_bytes = 0
self.comfy_stale_job: Optional[str] = None
self._idle_since: Dict[str, float] = {}
self._last_event_wake = 0.0
self.event_sources: Dict[str, str] = {}
self._last_tenant_state: Optional[Dict[str, Any]] = None
self.last_arbitration: Optional[Dict[str, Any]] = None
self.last_handoff: Optional[Dict[str, Any]] = None
self.last_watchdog_error: Optional[str] = None
self.stats = {
"yields": 0, # release confirmed
"yield_deferred_busy": 0, # model mid-generation; unload queued behind it
@@ -945,6 +982,12 @@ class AutoArbitrator:
return
self.running = True
self.ws_task = asyncio.create_task(self._ws_listener())
for t in tenants_mod.load_tenants():
if t.enabled and t.events.type == "websocket" and t.events.url:
self.event_tasks.append(asyncio.create_task(
self._event_listener(t.name, t.events.url,
t.events.reconnect_backoff_s,
t.events.max_backoff_s)))
self.poll_task = asyncio.create_task(self._poll_watchdog())
self.idle_task = asyncio.create_task(self._idle_purge_loop())
logger.info("AutoArbitrator background engine started (Bidirectional).")
@@ -957,9 +1000,10 @@ class AutoArbitrator:
async def stop(self):
self.running = False
for task in (self.ws_task, self.poll_task, self.idle_task):
for task in [self.ws_task, self.poll_task, self.idle_task, *self.event_tasks]:
if task:
task.cancel()
self.event_tasks.clear()
await close_clients()
logger.info("AutoArbitrator background engine stopped.")
@@ -1069,6 +1113,51 @@ class AutoArbitrator:
return {"purged": True, "free_gb": round(snap["free_bytes"] / (1024**3), 2)}
return {"purged": False, "free_gb": round(free_gb, 2), "reason": "ComfyUI holds no VRAM"}
async def _event_listener(self, tenant_name: str, url: str,
backoff_s: float, max_backoff_s: float) -> None:
"""Wake on a tenant's event stream instead of waiting for the next poll.
Deliberately does not parse the messages. The previous listener understood
ComfyUI's schema -- status/execution_start/executing/execution_success -- which
tied the fast path to one application. Treating any message as "look now" and
letting the tenant's own busy probe decide gives the same sub-second reaction
for any application that emits anything on state change.
"""
backoff = backoff_s
while self.running:
try:
async with websockets.connect(url, ping_interval=10, ping_timeout=10) as ws:
self.event_sources[tenant_name] = "connected"
self.connected_ws = True
backoff = backoff_s
logger.info(f"Event source connected for '{tenant_name}': {url}")
while self.running:
await ws.recv()
# Coalesce bursts: a single graph emits many messages, and one
# arbitration pass per burst is enough.
now = time.time()
if now - self._last_event_wake < 0.05:
continue
self._last_event_wake = now
self.stats["event_wakeups"] = self.stats.get("event_wakeups", 0) + 1
try:
# A message on this tenant's own stream is live proof it is
# working right now, so it is taken as busy rather than asked
# over HTTP. A stale queue row could lie; an event arriving
# this instant cannot.
await self._arbitrate(active_tenant=tenant_name)
except Exception as e:
logger.debug(f"arbitration from event failed: {e}")
except (websockets.exceptions.ConnectionClosed, OSError, asyncio.CancelledError):
self.event_sources[tenant_name] = "disconnected"
self.connected_ws = False
except Exception as e:
self.event_sources[tenant_name] = f"error: {str(e)[:60]}"
self.connected_ws = False
logger.debug(f"event source error for '{tenant_name}': {e}")
await asyncio.sleep(backoff)
backoff = min(backoff * 1.5, max_backoff_s)
async def _ws_listener(self):
client_id = "hyperswap-arbitrator"
ws_url = f"ws://127.0.0.1:8188/ws?clientId={client_id}"
@@ -1121,55 +1210,209 @@ class AutoArbitrator:
backoff = min(backoff * 1.5, 15.0)
RECLAIM_COOLDOWN_S = 30.0
# A queue entry that has claimed to be running this long without the GPU ever going
# busy is stale, not slow.
STALE_RUNNING_S = 90.0
# ComfyUI's own VRAM, not GPU utilisation, is what distinguishes a real job from a
# stale row. Utilisation is shared: Ollama and any third-party process drive it too,
# so peak utilisation stayed above any sensible threshold and a stuck entry never
# looked stale. A real diffusion job loads gigabytes of checkpoint; a dead one holds
# only the CUDA context.
STALE_COMFY_BYTES = 1.5 * 1024 ** 3
async def _check_ollama_starved(self) -> None:
"""The other direction: rescue an LLM that ComfyUI has squeezed onto the CPU.
def _comfy_genuinely_busy(self, queue: Dict[str, Any]) -> bool:
"""Decide whether ComfyUI is really working, not just claiming to be.
Yielding Ollama for ComfyUI was automatic; the reverse never was, despite the
README calling the arbitration bidirectional. When Ollama cannot fit a model it
does not fail, it silently places layers on the CPU and runs about an order of
magnitude slower -- so this is the failure mode a user is least likely to notice
and most likely to feel.
ComfyUI can leave an entry in queue_running after a job dies -- observed here as
a WAN 2.1 i2v entry that sat there with the GPU at 0% and ComfyUI holding 0.56 GB.
Trusting that flag alone made this service believe ComfyUI was permanently busy,
which meant it evicted the LLM on every poll, never ran the idle purge, and never
checked whether the LLM had been squeezed onto the CPU. Half the arbitration was
disabled by one stale row.
If the LLM is spilling while ComfyUI sits idle holding VRAM, ComfyUI's cached
checkpoints are the thing to give up.
A running entry is corroborated against GPU utilisation before it is believed.
"""
now = time.time()
if self.comfy_was_active or (now - self.last_reclaim_time) < self.RECLAIM_COOLDOWN_S:
return
running = queue.get("queue_running") or []
pending = queue.get("queue_pending") or []
if pending:
self._running_since = None
self._running_id = None
return True
if not running:
self._running_since = None
self._running_id = None
self.comfy_stale_job = None
return False
ollama = await get_ollama_live_state()
if not ollama.get("partially_offloaded"):
return
entry = running[0]
prompt_id = entry[1] if isinstance(entry, (list, tuple)) and len(entry) > 1 else str(entry)
now = time.time()
if prompt_id != self._running_id:
self._running_id = prompt_id
self._running_since = now
self._peak_comfy_bytes = 0
snap = get_process_vram_bytes()
if snap["comfyui_bytes"] < RECLAIM_MIN_COMFY_BYTES:
return # ComfyUI is not the one holding the memory; nothing we can do here
self._peak_comfy_bytes = max(self._peak_comfy_bytes, snap.get("comfyui_bytes", 0))
self.last_reclaim_time = now
model = ollama.get("active_model_name")
offload = ollama.get("cpu_offload_pct")
logger.warning(f"⚠ '{model}' is {offload}% on CPU while ComfyUI holds "
f"{round(snap['comfyui_bytes'] / (1024**3), 2)} GB — reclaiming for the LLM")
await self._purge_comfy_now(f"LLM spilling {offload}% to CPU")
self.stats["reclaims_for_ollama"] += 1
elapsed = now - (self._running_since or now)
if elapsed > self.STALE_RUNNING_S and self._peak_comfy_bytes < self.STALE_COMFY_BYTES:
if self.comfy_stale_job != prompt_id:
logger.warning(
f"ComfyUI reports prompt {prompt_id} running for {int(elapsed)}s while "
f"holding only {self._peak_comfy_bytes / (1024**3):.2f} GB — no checkpoint "
f"is loaded, so the queue entry is stale. Ignoring it; otherwise ComfyUI "
f"looks permanently busy and arbitration stops working.")
self.comfy_stale_job = prompt_id
return False
return True
# Freeing VRAM does not move layers back; only a reload re-places the model. Do
# that only when the model is idle, never mid-generation.
after = get_process_vram_bytes()
if after.get("gpu_util_pct", 0) < BUSY_UTIL_PCT and model:
logger.info(f"Reloading '{model}' to place it fully on the GPU...")
await instant_free_ollama_vram(model, confirm=True)
res = await switch_ollama_model(model, keep_alive="30m")
recheck = await get_ollama_live_state()
self.last_action = (
f"Reclaimed {round(snap['comfyui_bytes'] / (1024**3), 2)}GB from ComfyUI and "
f"reloaded '{model}' — now {round(recheck.get('gpu_fraction', 0) * 100)}% on GPU"
if res.get("success") else
f"Reclaimed VRAM from ComfyUI but reloading '{model}' failed: {res.get('error')}")
async def _tenant_state(self, active_tenant: Optional[str] = None
) -> List[Dict[str, Any]]:
"""Current VRAM and busy state for every configured tenant.
`active_tenant` skips the HTTP busy probe for the tenant whose event stream just
fired: the event is the evidence. That removes a round trip from the handoff,
which is the one path where latency is the entire point.
"""
# Only the cheap NVML read. The full hardware snapshot also does a psutil lookup
# per process, which is wasted work on the handoff path where latency is the
# entire point.
snap = get_process_vram_bytes()
by_tenant = {k: b / (1024 ** 3) for k, b in snap["by_tenant_bytes"].items()}
out = []
for t in tenants_mod.load_tenants():
if not t.enabled:
continue
bucket = _BUCKET_ALIASES.get(t.name, t.name)
vram_gb = by_tenant.get(bucket, 0.0)
if t.name == active_tenant:
probe = {"busy": True, "reason": "event received from its own stream"}
else:
probe = await tenants_mod.probe_busy(t, vram_gb=vram_gb)
out.append({
"name": t.name,
"priority": t.priority,
"vram_gb": vram_gb,
"busy": bool(probe.get("busy")),
"below_floor": bool(probe.get("below_floor")),
"reclaimable": t.reclaimable,
"needs_vram_gb": t.needs_vram_gb,
"overclock_profile": t.overclock_profile,
"vram_floor_gb": t.vram_floor_gb,
"idle_release_after_s": t.idle_release_after_s,
"reason": probe.get("reason"),
})
self._last_tenant_state = {"ts": time.time(), "free_gb":
round(snap["free_bytes"] / (1024**3), 2),
"tenants": out}
return out
async def _release_tenant(self, name: str, reason: str) -> Dict[str, Any]:
"""Release one tenant's VRAM by whatever mechanism it declares."""
t = tenants_mod.get_tenant(name)
if not t or not t.reclaimable:
return {"success": False, "reason": "not reclaimable"}
models = None
if t.release.per_model:
state = await get_ollama_live_state()
models = [m.get("name") for m in state.get("loaded_models", []) if m.get("name")]
logger.info(f"Releasing VRAM from '{name}': {reason}")
res = await tenants_mod.release_vram(t, models=models)
self.stats["tenant_releases"] = self.stats.get("tenant_releases", 0) + 1
return res
IDLE_PROFILE = "balanced"
def _apply_profile_for_active(self, state: List[Dict[str, Any]]) -> None:
"""Apply the GPU profile declared by whichever tenant is currently working.
This used to be two calls naming 'comfy' and 'ollama' directly, so a third
application could never get tuned clocks. The highest-priority busy tenant wins;
with nothing working the card returns to the idle profile.
"""
busy = [s for s in state if s["busy"] and s.get("overclock_profile")]
if busy:
busy.sort(key=lambda s: -s["priority"])
self._apply_oc_profile(busy[0]["overclock_profile"])
else:
self.last_action = (f"Reclaimed VRAM from ComfyUI; '{model}' is busy, so it will "
f"stay partly on CPU until its next load")
self._apply_oc_profile(self.IDLE_PROFILE)
async def _arbitrate(self, active_tenant: Optional[str] = None) -> None:
"""Generic arbitration over any number of tenants.
The two-application version was a pair of hardcoded rules -- yield Ollama when
ComfyUI is busy, purge ComfyUI when Ollama is starved -- which could not express
a third participant at all. This works from the registry instead: a busy tenant
that lacks the VRAM it declares it needs is starved, and the memory comes from
idle reclaimable tenants below it in priority, lowest first.
"""
t_start = time.perf_counter()
state = await self._tenant_state(active_tenant)
free_gb = self._last_tenant_state["free_gb"]
self._apply_profile_for_active(state)
# 1. Starvation: highest-priority demanding tenant first.
for s in sorted(state, key=lambda x: -x["priority"]):
if not s["busy"] or not s["needs_vram_gb"]:
continue
# Starved means it cannot reach what it needs even counting what it already
# holds. Comparing free VRAM alone flagged a tenant that was working
# perfectly well on 13 GB as demanding, purely because little was left over
# -- which is the normal state of a busy GPU, and would have caused
# pointless releases from everyone else.
if s["vram_gb"] + free_gb >= s["needs_vram_gb"]:
continue
plan = tenants_mod.plan_release(s["name"], state, free_gb, s["needs_vram_gb"])
self.last_arbitration = {"ts": time.time(), "demanding": s["name"],
"free_gb": free_gb, **plan}
if not plan["release"]:
logger.debug(f"'{s['name']}' is short of VRAM but {plan['reason']}")
return
if time.time() - self.last_reclaim_time < self.RECLAIM_COOLDOWN_S:
return
self.last_reclaim_time = time.time()
for victim in plan["release"]:
await self._release_tenant(
victim, f"{s['name']} needs {s['needs_vram_gb']} GB, {free_gb} GB free")
# Wait for the memory to actually come back, and record how long the whole
# handoff took. Swap speed is the point of this service, so it is measured
# rather than assumed.
target_bytes = int(s["needs_vram_gb"] * (1024 ** 3))
deadline = time.perf_counter() + 30.0
while time.perf_counter() < deadline:
if get_process_vram_bytes()["free_bytes"] >= target_bytes:
break
await asyncio.sleep(0.02)
handoff_ms = round((time.perf_counter() - t_start) * 1000, 1)
self.last_handoff = {"ts": time.time(), "to": s["name"],
"released": plan["release"], "handoff_ms": handoff_ms,
"triggered_by": "event" if active_tenant else "poll"}
self.stats["handoffs"] = self.stats.get("handoffs", 0) + 1
logger.info(f"Handoff to '{s['name']}' in {handoff_ms} ms "
f"(released {', '.join(plan['release'])})")
self.last_action = (f"Released {', '.join(plan['release'])} so "
f"'{s['name']}' could work — {handoff_ms} ms")
return
# 2. Idle release: a tenant holding VRAM it is not using, after a grace period.
now = time.time()
for s in state:
if not s["reclaimable"] or s["vram_gb"] <= 0.25:
self._idle_since.pop(s["name"], None)
continue
if s["busy"]:
self._idle_since.pop(s["name"], None)
continue
since = self._idle_since.setdefault(s["name"], now)
grace = s["idle_release_after_s"]
if grace and (now - since) >= grace:
self._idle_since.pop(s["name"], None)
await self._release_tenant(
s["name"], f"idle {int(now - since)}s holding {s['vram_gb']} GB")
self.last_action = (f"Released idle '{s['name']}' after "
f"{int(now - since)}s")
return
async def _poll_watchdog(self):
"""Fallback for when the WebSocket is down. One cheap /queue call, 1 Hz.
@@ -1184,15 +1427,25 @@ class AutoArbitrator:
resp = await client.get("/queue")
if resp.status_code == 200:
q = resp.json()
busy = len(q.get("queue_running", [])) > 0 or len(q.get("queue_pending", [])) > 0
busy = self._comfy_genuinely_busy(q)
if busy:
self.watchdog_branches["busy"] += 1
await self.trigger_comfy_priority("Watchdog saw an active queue")
elif self.comfy_was_active:
self.watchdog_branches["completed"] += 1
await self.trigger_comfy_completed()
else:
await self._check_ollama_starved()
except Exception:
pass
self.watchdog_branches["idle_check"] += 1
await self._arbitrate()
else:
self.watchdog_branches["bad_status"] += 1
except Exception as e:
# This used to swallow everything silently, including anything raised by
# the starvation check, which is why that check could appear to run and
# do nothing.
self.watchdog_branches["error"] += 1
self.last_watchdog_error = str(e)[:200]
logger.debug(f"watchdog poll error: {e}")
await asyncio.sleep(interval)
def suspend_oc(self, reason: str = "tuning sweep") -> None:
@@ -1230,6 +1483,13 @@ class AutoArbitrator:
"idle_purge_after_s": self.COMFY_IDLE_PURGE_S,
"oc_profile": self.oc_profile,
"counters": dict(self.stats),
"comfy_stale_job": self.comfy_stale_job,
"event_sources": dict(self.event_sources),
"last_arbitration": self.last_arbitration,
"last_handoff": self.last_handoff,
"tenant_state": self._last_tenant_state,
"watchdog_branches": dict(self.watchdog_branches),
"last_watchdog_error": self.last_watchdog_error,
"yield_backoff": {m: round(max(t - time.time(), 0), 1)
for m, t in self._yield_backoff_until.items()
if t > time.time()},