Measured failure: a diffusion job took 46 s instead of 3 s, squeezed into 1.6 GB.
The LLM yielded correctly, an inference reloaded it two seconds later, and
plan_release then refused to touch it because it was "busy" -- so ComfyUI crawled
while Ollama held 13 GB for the whole run.
"Never interrupt busy work" looks like the safe rule and is not. Preempting a
lower-priority tenant is safe precisely because releasing is asynchronous: an Ollama
unload queues behind its running request and applies when that finishes, so nothing
is killed mid-flight. That is what makes fast handoff possible at all, and refusing
to do it defeats the purpose of the service.
A tenant may now be asked for memory if it is idle, whatever its rank, or if it is
busy and ranks strictly below the demander. Equal or higher priority is never
interrupted, so peers cannot fight. Idle tenants are still preferred over preempting
busy ones.
Verified against the real contention: with Ollama at 12.38 GB and 98% utilisation, a
diffusion job released it within three seconds and completed in 18 s rather than 46 s.
Tests updated to the corrected rule, and the fixture's priorities aligned with what
actually ships -- it still had the LLM outranking diffusion from before that was
swapped.
Tests: 246.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two remaining pieces of the two-application coupling are gone.
Overclock profiles were switched by naming 'comfy' and 'ollama' directly, so a third
application could never get tuned clocks. A tenant declares overclock_profile and the
arbitrator applies whichever the highest-priority *working* tenant asks for, falling
back to the idle profile when nothing is running.
The websocket listener parsed ComfyUI's message schema -- status, execution_start,
executing, execution_success -- which tied the fast path to one application. An event
source is now declarative and the messages are not parsed at all: any message means
"look now", and the tenant's own busy probe decides what is true. That gives the same
sub-second reaction to any application that emits anything on state change, with no
knowledge of what it emits.
Generalising this exposed a design error in the priority rule I had introduced.
plan_release excluded candidates ranking above the demander, which broke both
directions in turn. With the LLM at priority 60 and diffusion at 50, ComfyUI could
never reclaim from Ollama -- the premise the whole service is built on, and preserved
until now only by the ComfyUI-specific trigger that was about to be removed. Swapping
the ranks then broke the reverse: a starved Ollama could no longer reclaim from an
idle ComfyUI.
Priority now orders rather than vetoes. Any idle reclaimable tenant is a candidate,
because an idle tenant is not using its VRAM; priority decides who is asked first, and
busy tenants are never interrupted whatever their rank. Diffusion outranks the LLM,
whose weights reload from page cache in seconds. All three cases are pinned by tests,
including that busy work is never interrupted even by a far higher-priority demander.
Tests: 244 (was 242).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
_check_ollama_starved and _arbitrate were solving the same problem, one of them
hardcoded to two applications. The Ollama-specific version is gone and the watchdog
calls only the generic loop. The OOM retry inside switch_ollama_model no longer
purges ComfyUI by name either: it asks plan_release which tenant should give up
memory, so a third application can be the one that yields, and when the reclaim is
not enough the response names the blockers instead of implying ComfyUI was at fault.
The dashboard showed exactly two engines, which no longer matched what the service
does. A GPU Tenants panel lists every configured application ordered by priority --
VRAM held, whether it is working, whether it can be reclaimed at all, and how much it
needs -- along with the last arbitration decision and why it could or could not be
satisfied.
That panel did not appear at first, and the reason is worth fixing rather than
working around: the browser kept serving a cached app.js despite the ETag, so a
reload ran the old dashboard against the new API. Assets are now stamped with their
mtime, so a changed file is always a different URL. Anyone updating this service would
have hit the same thing.
Also verified along the way that an apparent horizontal-overflow regression was a
measurement artifact from a zero-width browser pane, not a real layout fault.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three things pinned it small. The container was a fixed h-64, so the graph stayed
256px tall on any display; the page was capped at max-w-7xl, so a large window only
added empty margins; and the graph shared a two-column row, so it never got more than
half the available width.
The container is now h-[clamp(18rem,45vh,52rem)], the page cap is 2600px, and the
graph panel spans its row. Measured: 525x360 to 1142x360 at 1280x800, and 988x648 to
2468x648 at 3440x1440.
Chart.js needed help despite being configured responsive with maintainAspectRatio
false. It had latched onto a stale size -- the canvas sat at width:0px, height:288px
while its container had grown to 988x648. A ResizeObserver on the container fixed
growth but not shrinking, because passing explicit dimensions to resize() left the
canvas 988px wide inside a 435px container, overflowing it. Calling resize() with no
arguments lets Chart.js measure the container itself, and the canvas is absolutely
positioned within the relative container so it cannot force it wider.
Verified from 1100x700 to 3440x1440: the canvas matches its container exactly at every
size, growing and shrinking, with no horizontal page overflow.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Completes the generalisation. Classification and release were already data; the
decision loop was still two hardcoded rules -- yield Ollama when ComfyUI is busy,
purge ComfyUI when Ollama is starved -- which could not express a third participant.
plan_release() works from the registry instead. A busy tenant that cannot reach its
declared needs_vram_gb is starved, and the memory comes from idle reclaimable tenants
below it in priority, lowest first, stopping once enough is freed. Tenants that cannot
be released are named as blockers rather than passed over, so an impossible plan says
which process is in the way. The plan is returned before being acted on, so the
decision is testable and is logged before anything is released. Idle release is now
per-tenant too, replacing the ComfyUI-specific purge timer.
Two bugs found by running it against the live machine rather than only in tests:
Starvation was measured against free VRAM alone, so a tenant working perfectly well on
13 GB was flagged as demanding simply because little was left over -- which is the
normal state of a busy GPU, and would have caused pointless releases from everything
else. A tenant is starved only if it cannot reach what it needs counting what it
already holds.
Fields added to the tenant schema were silently absent from the config already written
to disk, so needs_vram_gb defaulted to 0 and starvation could never trigger for the two
tenants that mattered. Shipped defaults are now merged into an existing config on load,
with explicit user values still winning.
Tests: 242 (was 231), including a three-application contention case -- the property the
hardcoded pair of rules could not express.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The point of this service is fast handoff of one GPU between applications. It grew up
around the two on this box, and their names ended up compiled into process matching,
VRAM attribution, busy detection and release calls alike -- about 385 references
across five modules. That made it a script for Ollama and ComfyUI rather than a GPU
arbitrator.
tenants.py describes an application as data: how to recognise its processes, how to
tell whether it is genuinely working, how to ask it for VRAM back, and how much it
matters when two want the card. Ollama, ComfyUI and the desktop compositor ship as
defaults in tenants.json, so behaviour is unchanged, but the arbitration logic no
longer knows any particular name. Endpoints are generic: GET /api/tenants,
GET /api/tenants/{name}, POST /api/tenants/{name}/release -- the last being the
general form of both the Ollama soft-yield and the ComfyUI purge.
Verified by registering a third application on this machine with no code change: the
speech relay that had been showing up only as anonymous "unmanaged VRAM" is now named,
attributed, and probed by the VRAM it holds rather than by an API it does not have.
Because it declares no release strategy, a release request returns 409 explaining that
its memory cannot be reclaimed, instead of reporting a success that did nothing.
Busy probes deliberately cannot use GPU utilisation. It is shared by every tenant, so
it cannot attribute work to one of them -- the mistake that made a stale ComfyUI queue
entry undetectable earlier in this branch. A tenant's own VRAM is the signal.
Writing the tests exposed that the suite had become non-hermetic: classification is now
configuration, so a test asserting "a third-party process is unmanaged" started failing
the moment the speech relay was registered on this machine. An autouse fixture now
isolates every test from the operator's live tenants.json.
Tests: 231 (was 206).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Chasing why the reverse-direction reclaim never fired turned up something worse than
the reclaim itself.
The starvation check was never running. Instrumenting the watchdog showed busy=6,
idle_check=0: every poll took the "ComfyUI is busy" branch. ComfyUI's /queue was
reporting a WAN 2.1 i2v job in queue_running while the GPU sat at 0% and ComfyUI held
0.56 GB. The job was dead; ComfyUI had simply never cleared the row.
Believing that flag meant this service thought ComfyUI was permanently busy, so it
yielded the LLM's VRAM on every poll, never ran the idle purge, and never checked
whether the LLM had been squeezed onto the CPU. One stale row disabled half of the
arbitration, and it very likely explains the earlier burst of yields against a
cron-driven model.
A running entry is now corroborated before it is believed. The first attempt used GPU
utilisation, which does not work: utilisation is shared with Ollama and with the
third-party process on this box, so peak utilisation stayed above any sensible
threshold and a stuck entry never looked stale. ComfyUI's own VRAM is the right
signal -- a real diffusion job loads gigabytes of checkpoint, a dead one holds only
its CUDA context. After the fix the same watchdog reports busy=3, idle_check=32.
Every early return in the starvation check now records why it bailed, because with
four of them there was no way to tell which had fired. /api/health reports a stale
queue entry with its impact and how to clear it.
Also confirmed, contradicting an earlier conclusion in this branch: Ollama on this box
*does* spill to the CPU. smtek/Qwen3.8-27B:Q2_K_XL held steady at 29.2% on GPU
(size=15.59 GB, size_vram=4.56 GB) across twelve seconds of polling -- a stable
placement, not a progressive load. Both failure modes are real; which one occurs
depends on the model.
Tests: 206 (was 199).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The health check now reports what unmanaged VRAM actually costs rather than just
how much of it there is: "0.82 GB held by python (842 MB)" becomes "5 model(s) fit
within 15.42 GB but not the 14.60 GB actually available", naming them.
Getting that arithmetic right took a correction. The first version subtracted only
the desktop and the unmanaged process, and so reported a 14.93 GB model as fitting
against a real ceiling of 14.60 GB -- the same model the service had just refused
with 507. ComfyUI keeps a few hundred MB of CUDA context for as long as the process
lives, which a purge does not free, so it is not available either. The floor is taken
from the minimum ComfyUI VRAM in recent telemetry rather than its current value,
which could be a 7 GB checkpoint mid-generation.
The verifier's reclaim stage now re-runs a graph immediately beforehand to reset the
30 s idle window, since a large model takes longer than that to load and the purge
was freeing ComfyUI mid-load, so the reclaim path was never reached.
Tests: 199 (was 192). The new ones pin the ceiling arithmetic, including that a model
too large to fit on the card at all is not blamed on the third-party process.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The first full run passed every stage, and three of those passes were worth less
than they looked.
Diffusion was reported as 0.67 it/s. That single run included loading the SDXL
checkpoint from disk, so it understated throughput roughly tenfold against a
steady-state 5.49 it/s. Cold and warm are now timed and labelled separately.
The idle-purge stage checked the flag immediately, racing the ComfyUI websocket
event that sets it, and reported "no purge pending" as a warning about ComfyUI
rather than about its own timing. It now waits for the event.
The reclaim stage passed while proving nothing: the model chosen was small enough to
fit alongside ComfyUI's checkpoint, so the reclaim path never ran. It now picks a
model that genuinely will not fit, and reports a warning rather than a pass when the
path is not exercised.
Sizing that model correctly took two corrections, both real. On-disk weight size is
not the VRAM footprint -- a 12.87 GB blob occupies 14.9 GB once context and KV cache
are allocated -- and unmanaged VRAM is not reclaimable, so it cannot count toward
what a reclaim will free. Ignoring the second picked a model that failed even after a
correct reclaim: the service returned 507 and logged "could not fit with ComfyUI
holding 7.03 GB -- reclaiming and retrying", which was right. On this box an 842 MB
third-party process is the difference between a 14.9 GB model fitting and not.
The verifier also died on the 507 instead of reporting it, since a helper called
raise_for_status() on responses a stage deliberately provokes.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The unit suite covers logic in isolation, but the promise this service exists to make
-- an LLM and a diffusion pipeline sharing one 16 GB card without either failing --
had only ever been checked by hand, piecemeal. verify_arbitration.py walks the whole
cycle against real hardware and reports what happened at each stage: load and its
bandwidth classification, the confirmed yield, a real SDXL graph, the deferred idle
purge, reclaim-and-retry, and finally whether VRAM attribution adds up and the
reported GPU state still matches the card. It restores what it changes and refuses
to start if ComfyUI is busy. Kept out of pytest deliberately: it moves real VRAM and
takes minutes.
Running it immediately found two bugs.
/api/switch-model reported every upstream failure as 500. Asking an embedding model
to generate makes Ollama return 400 -- the request is unusable, the service is fine
-- and calling that an Internal Server Error blames this service for the caller's
mistake. Failures now map to 400 for an upstream client error, 507 for a model that
will not fit (valid request, healthy service, no room), and 502 when Ollama itself
errors.
The verifier also picked the smallest installed model, which here is
nomic-embed-text -- an embedding model with no generate endpoint. It now filters
those out by family and name.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The engine subtitles were hardcoded: "FlashAttention + Q4 KV Cache" and "DynamicVRAM
+ Pinned Async Offload". The first turned out to be accurate -- OLLAMA_FLASH_ATTENTION
and OLLAMA_KV_CACHE_TYPE really are set -- which is worse than being wrong, because it
would have gone on looking accurate after the settings changed.
engines.py reads both engines' real configuration: the ollama service environment via
systemd, and ComfyUI's own /system_stats for version, allocator, VRAM mode and argv.
Exposed at GET /api/engines, as an MCP tool, and in the dashboard subtitles with the
full settings list as a tooltip.
The settings worth surfacing are the ones that dictate how this service must behave
and that previously had to be discovered by reading journald: OLLAMA_NUM_PARALLEL=1
is why an unload queues behind a running generation and is reported as deferred
rather than failed, and OLLAMA_MAX_LOADED_MODELS=1 is why every swap evicts the
previous model. Each is reported with that explanation attached.
Writing the tests found a bug in the new code: (system.get("python_version") or
"").split()[0] raises IndexError when ComfyUI omits the field, and the surrounding
except would have swallowed it and reported ComfyUI as entirely offline.
Tests: 192 (was 182).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A screenshot of the running dashboard showed four claims that the backend does not
support. Each is the same failure the rest of this branch has been correcting:
displaying a number without checking what it measures.
"Models in RAM: 54.31 GB" was ram.cached_gb -- the entire Linux page cache, every
file the kernel has cached, not models. Measured model residency at the same moment
was 42.87 GB of a 331 GB catalogue. The two are different quantities that happened
to look similar. The legend now says "Page Cache (all files)" for what that number
is, and shows measured model residency separately beside it.
"LAST SWAP TIME 1.65 ms" and "RAM HIT STATUS: Cleaned" were read from history[0],
the most recent event of any type. Both tiles were describing a ComfyUI VRAM purge:
its duration under a swap-time label, its status under a cache-hit label. They now
select the most recent LLM Model Switch, and fall back to the persisted log when the
in-memory ring is empty, so a restart no longer blanks them.
"Host Pinned Memory: 53.6 GB Staging Buffer", "Async PCIe Offloading: Enabled (2
Streams)" and "Fast Disk RAM Mmap: Active" were hardcoded in the HTML and measured
from nothing at all. Replaced with values the service actually has: VRAM currently
held by ComfyUI, live PCIe throughput from NVML, measured checkpoint residency, and
the real idle-purge countdown.
Also: the VRAM legend's "System" bucket hid an 842 MB third-party process behind the
same label as a 3.9 MB compositor, so it is now split into Desktop and Unmanaged;
and the version badge read a hardcoded "v1.0-DEPLOY" while the API served 2.0.0, so
it now reads the version from /openapi.json.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The reconciler only compared the power limit, so a profile whose fan setting never
applied stayed wrong indefinitely. The headless X server that owns the GPU can be
starting when this unit does; in-process retries cover a short delay, but if X
arrives later nothing noticed that the fan mode had never been set. profile_drift()
now compares fan mode too, gated on fan control having worked at least once so the
check does not fire forever on a machine without it.
README documents the self-check endpoint, the ollama/comfy/desktop/unmanaged
bucketing and why unreclaimable VRAM is reported separately, the SSE trimming, and
a tests section listing the measured constants the suite pins.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Self-check. Fan control failed for an entire session -- recoverably, and completely
invisibly. It appeared once, inside one field of one log line, and nothing ever
asked whether fan control worked. health.py now checks everything this service
depends on (NVML, passwordless sudo for nvidia-smi, fan control via the headless X
server, overclock drift, the telemetry store, residency measurement capability,
model directories, the ComfyUI websocket, and both upstream HTTP services) and
reports for each one what is broken, what that breaks, and how to fix it. Exposed at
GET /api/health, as an MCP tool, and as a dashboard panel that collapses to a badge
when healthy and expands to impact-and-fix when not. Current state: 9 ok, 1 degraded
(the known cachestat permission limit on Ollama's blobs).
A self-check that returns ok while a dependency is broken is worse than none, so the
tests drive each check to its failure state -- including the exact "Error resolving
target specification 'gpu:0'" string from the original incident -- and assert that a
check which raises surfaces as failed rather than taking down the endpoint.
SSE payload. The installed-model catalog was 10.6 KB of a 13.1 KB frame, 81% of the
stream, re-sent to every subscriber every second despite changing only when a model
is pulled or removed: 135 MB/hour across three tabs. It is now sent on a
subscriber's first frame and whenever the set changes; the client keeps the last
known list. Steady-state frames dropped from 14041 to 3664 bytes, a 74% reduction,
and /api/stats still returns the complete snapshot for API consumers.
Tests: 182 (was 169).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The stale-readback bug fixed in f5917a0 was a class, not an instance. Two more:
- ram_optimizer's residency report cached for 15s and was never invalidated when
anything warmed a file, so warming a model and then looking at residency showed
the state from before the warm. warm_file_to_ram now invalidates it.
- _PID_KIND_CACHE was keyed on pid alone and never expired. Linux recycles PIDs, so
a stale entry could attribute a new process's VRAM to Ollama or ComfyUI -- inside
the very snapshot the yield barrier trusts to decide whether VRAM was released.
Now keyed by (pid, process start time) and bounded.
Unmanaged VRAM. Investigating a persistence-mode warning turned up a third GPU
consumer this service does not model: stt_relay.py, holding 842 MB for nearly three
days. It was bucketed as "system" alongside gnome-shell's 3.9 MB. That conflation
matters, because ComfyUI's memory can be reclaimed and a third party's cannot, and
the reclaim path assumed ComfyUI was always to blame for missing headroom.
Processes are now bucketed ollama | comfy | desktop | unmanaged. The breakdown
reports desktop_gb and unmanaged_gb separately and names the unmanaged processes;
when a reclaim-and-retry still fails, the error identifies them rather than
implying ComfyUI was at fault; and the dashboard shows the unreclaimable total, so
headroom the arbitrator can never give back is visible rather than inferred.
Checked and deliberately not changed: persistence mode reads Disabled, but
nvidia-persistenced is active and two clients hold the GPU open continuously, so
the driver never unloads. The nvidia-smi warning is legacy noise here and is not a
source of the profile drift.
Tests: 169 (was 164). The new ones cover the bucketing, and one existing test used
Xorg as its "unknown process" fixture -- correct before a display server had its own
bucket, wrong after.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The API claimed the card was at 320W while nvidia-smi reported 370W. Three
separate defects, all introduced by me in this branch.
The readback was stale. get_gpu_state()/get_fan_status() gained a 2s cache so the
dashboard's polling would stop forking sudo every few seconds, but apply_profile
read back through that cache and its invalidation ran afterwards. A profile that
had just moved the card 370W -> 320W therefore returned a payload whose detail
string said "set to 320.00 W from 370.00 W" next to a power_limit_w of 370.0.
Caches are now cleared before the readback, which is forced.
Fan control could fail for an entire session. On boot this unit can start before
the headless X server on :8 that owns the GPU accepts connections, and the fan
assignment fails with "Error resolving target specification 'gpu:0'". Nothing
retried and nothing surfaced it, so the fans were left unconfigured with the
failure visible only inside one log line. apply_fan_control now recognises that
specific error and retries up to 5 times.
Nothing verified the result. ACTIVE_PROFILE defaults to "balanced" at import,
which is indistinguishable from "balanced was successfully applied" -- so a failed
startup apply left the app confidently reporting a profile it had never put on the
hardware. apply_profile now returns a `verified` block comparing intent against
readback and logs a warning on mismatch; profile_drift() exposes the comparison
plus whether any profile has actually been applied since startup; and the 1Hz
sampler calls reconcile_profile() once a minute to re-apply on drift.
Verified by setting 370W externally behind the service's back: the drift was
reported immediately and corrected automatically 40s later.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Tests. First automated coverage for the project: 164 tests, 2.7s, no GPU or network.
An autouse fixture stubs overclock_manager._sh -- the single choke point for every
nvidia-smi/nvidia-settings write -- so no test can mutate the card. They deliberately
pin the empirically measured constants that would otherwise rot silently: the cold and
warm load figures behind the cache-hit thresholds, the warm_confident residency rule,
and the busy/stalled yield split. One test asserts RAM_HIT_GBPS stays at or below the
measured 2.63 GB/s warm load, so the old physically unreachable 5.0 GB/s bar cannot
come back.
Three bugs the suite surfaced, now fixed:
- autotune._subsample(values, 1) divided by zero; the early return only covered
len(values) <= max_steps.
- telemetry_store.stop() flushed its local pending list but never drained the queue,
silently losing rows submitted just before a shutdown -- exactly when the last
events matter.
- ram_optimizer.page_residency's zero-byte short-circuit omitted keys every other
return path provides, so a 0-byte file was planned for warming.
Reclaim. The README has claimed bidirectional arbitration from the start, but only one
direction was ever automatic. Establishing what actually happens took a controlled test
with the service stopped: with ComfyUI holding 6.83 GB, Ollama does not spill to the CPU
on this box -- it aborts with "cudaMalloc failed: out of memory", because n_gpu_layers is
pinned to 99 and it will not reduce the layer count. So both failure modes are handled:
_check_ollama_starved watches size_vram < size for the default configuration where Ollama
does spill, and switch_ollama_model catches the hard OOM, reclaims VRAM from an idle
ComfyUI and retries once. The request that returned HTTP 500 from Ollama directly now
succeeds through HyperSwap, loading at 3.85 GB/s after reclaiming 6.83 GB.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The persisted counters showed 19 timeouts in 20 yields. All 19 were one model,
ornith-1.5:9b-cron, in two bursts at 06:25 and 06:33. Telemetry for that window
shows the GPU pinned at 96-97% with Ollama holding 14.92 GB throughout: the model
was mid-generation. Ollama will not unload a model that is inferencing, so every
request failed, and with a 1s trigger debounce against a 10s blocking wait the
arbitrator simply asked again, four times per burst, blocking the loop for 40s.
Ollama's behaviour is correct. Ours was wrong in three ways.
Busy is now a distinct outcome. _await_vram_release returns "released", "busy" or
"stuck": VRAM that has not moved while the GPU is pinned means a generation is in
flight, which is not a failure. With OLLAMA_NUM_PARALLEL=1 our keep_alive:0 request
queues behind the running one and applies the moment it finishes, so the correct
response is to stop waiting, not to retry. Only "stuck" -- VRAM held with an idle
GPU -- is a real fault.
The wait is short again (2s, from 10s) because blocking helps nobody: ComfyUI is
not gated on our return value, and every blocked second stalls the watchdog and
profile switching. Callers who genuinely want to wait out an inference can pass
wait_for_generation=true. The unload POST itself now gets a 120s client timeout,
since a 5s one could drop the connection before Ollama ever processed a request
queued behind a long generation, losing the unload entirely.
A busy model gets per-model backoff (5s, 15s, 30s, 60s) instead of being asked
again every second, and a detached watcher confirms and logs the release when the
generation ends, so the event log tells the whole story rather than stopping at
"deferred". Measured: a mid-generation yield now returns busy in 610ms instead of
blocking 10s, and the queued unload lands on its own 3s later when the generation
completes.
Counters are honest: yields (released), yield_deferred_busy, deferred_releases,
yield_stalled. The old yield_timeouts conflated a healthy cron job with a fault
and implied a 95% failure rate.
Also adds a VRAM Arbitration panel to the dashboard. The arbitrator is the core of
this application and its state was not displayed anywhere -- there was no way to
see whether handoffs were working, which is why this went unnoticed until the
persisted counters were read by hand.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The repo is about to be pushed to a remote, so this covers what should never
travel with it and what should not be baked into the source.
.gitignore now covers credentials (.env, keys, tokens, .netrc), host-local
config (*.local.json), the SQLite telemetry store and its WAL sidecars, logs,
benchmark and sweep output, and the timestamped .bak files this project has
accumulated before. Verified that no currently tracked file is caught by the
new patterns.
Also removed /home/drjones from tracked source, which an ignore file cannot
help with. BASE_DIR now derives from the module's own location, the ComfyUI
model directory falls back to ~/ComfyUI/models, and start_manager.sh resolves
its interpreter through $HOME with a python3 fallback. All three still resolve
to exactly the same paths on this machine; they just no longer hardcode one
user's home directory into a published repository.
Note for the record: the git history was scanned across all refs and contains
no credentials. The password visible in `git remote -v` lives only in
.git/config, which is never pushed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Six commits reworking the arbitrator, memory accounting, telemetry and overclock
handling. The through-line is replacing assumptions with measurements, several of which
turned out to be wrong:
- VRAM yields are confirmed against NVML rather than fire-and-forget. The HTTP call
returns in ~63ms; the driver needs a further ~77ms to release 14.9GB. That window is
where ComfyUI could allocate into occupied VRAM.
- Page-cache residency is measured with cachestat(2), with a randomised read-rate probe
where the kernel refuses it. mincore(2) had been claiming 128GB resident on a box with
46GB of page cache.
- Cache-hit classification is calibrated against real cold and warm loads (0.38 vs 2.63
GB/s for the same 12.87GB model), not PCIe bus bandwidth.
- Overclock profiles are rebuilt from sweeps. Clock offsets turned out to be silently
ignored by this driver, clock locks changed nothing for either workload, and LLM
decode is not power-bound at all. Only ComfyUI's 370W limit earns its keep (+2.8%).
- Fans are automatic everywhere. 48k samples show 81C all-time max and zero thermal
throttle events, while the ollama profile had been holding 49.6C at 87% fan.
- Telemetry and events persist to SQLite so profile performance can be compared at all.
- A thermal governor de-escalates on sustained heat, and stock state is restored on
shutdown and via systemd ExecStopPost.
ComfyUI VRAM is no longer purged between workflow iterations, the 1Hz telemetry sample
is taken once and fanned out rather than recomputed per client, and the MCP surface is
back in parity at 23 tools and 6 resources.
Calibration. The same 12.87GB model loaded through Ollama on this box:
3.1% resident (FADV_DONTNEED) -> 34.3s -> 0.38 GB/s
100% resident (force-warmed) -> 4.9s -> 2.63 GB/s
The thresholds had been guessed from PCIe bus bandwidth: cache hit at >=5 GB/s. A fully
warm load only reaches 2.63 GB/s, because load_duration covers host-to-device transfer
and model init as well as the file read -- the page cache itself reads at 6.4 GB/s. The
5 GB/s bar was therefore unreachable, and every warm load was being reported as a
partial hit. Now 2.0 / 0.8 GB/s, either side of the measured 6.9x separation.
Warm-skip was also unsafe. A 12.87GB blob was skipped as already resident on the
strength of twelve 2MB probe windows, then loaded at 2.44 GB/s. Skipping now requires
warm_confident: an exact cachestat reading, or a probe finding every one of 32 denser
samples resident. warm_file_to_ram/warm_ollama_blob take force=True, exposed on the
warm-model endpoint, whose Pydantic model was missing the field entirely.
MCP parity: the server had drifted well behind the REST API. Adds tools for measured
residency, warm planning, VRAM requests, per-profile analytics, thermal governor
control, overclock status/apply/restore, and autotune sweeps plus status -- 23 tools
and 6 resources, up from 12 and 3. The telemetry store now starts in __main__ rather
than at import scope, since server.py imports this module for the benchmark tool.
README: replaced the remaining theoretical claims (31.5 GB/s bus rate, sub-1.5s loads,
15ms yields) with the measured numbers.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The profiles were hand-written and had never been checked against the hardware. Adding
a ComfyUI benchmark alongside the existing decode one made the compute side measurable
for the first time, and most of what the profiles configured turned out to do nothing.
Measured on this card (RTX 4080 SUPER, driver 595.84):
- LLM decode is not power-bound: 73.0-73.5 tok/s flat from 222W to 370W, with the card
never drawing more than 224W at any limit. The ollama profile's 370W did nothing.
- Diffusion is power-bound: 5.48 it/s @222W rising to 6.71 @370W, so comfy's 370W is
worth a real +2.8% over the 320W stock default.
- Clock locks did nothing for either workload: 72.6 tok/s locked at 11251MHz vs 72.7
unlocked; 6.77 it/s locked at 3105MHz vs 6.73 unlocked, and 6.78 at 2400MHz.
- Memory bandwidth is still the decode bottleneck (5001MHz halves throughput to 35.9
tok/s), confirming the profile's premise -- the card just gets there unaided.
- Fans: 48,435 samples show 81C all-time max and zero thermal throttle events, while
the ollama profile held 49.6C average by running fans at 87%. All profiles now use
automatic fans and let the thermal governor escalate on demand.
Code changes supporting that:
- _diffusion_benchmark() queues a fixed SDXL graph via ComfyUI's API. The seed must
vary per run: ComfyUI caches by node inputs, so a fixed seed returned in ~1ms without
executing. Implausibly fast results are now rejected as cache hits rather than
recorded as record scores.
- The arbitrator's automatic profile switching is suspended during a sweep. A diffusion
benchmark trips trigger_comfy_priority, which reapplies the whole profile and would
silently overwrite the clock being measured.
- _supported_clocks() queries the mem,gr pair; asking for a single field returned one
column and reading index 1 yielded an empty list rather than an error. Graphics clocks
are subsampled (the card enumerates 194 of them) and lock sweeps include an explicit
unlocked control step.
- offsets_supported() probes once and apply_profile skips inert offset levers with an
explanation instead of pretending they applied.
- Profiles carry a 'measured' field recording the evidence behind each setting.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Every connected dashboard holds a StreamingResponse open indefinitely, so uvicorn's
graceful shutdown waited on them, systemd hit its 90s stop timeout and SIGKILLed the
unit. That skipped the in-process restore hook entirely, leaving the GPU restore to
ExecStopPost alone.
- broker.stop() sets a closing flag and pushes a sentinel to every subscriber queue so
the SSE generators return instead of parking on q.get().
- uvicorn gets timeout_graceful_shutdown=10 and the unit TimeoutStopSec=20, bounding
the worst case rather than relying on the 90s default.
Restart now completes in ~11s with 'Restoring GPU to safe stock state (server
shutdown)' running in-process, and no SIGKILL.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The confirm barrier surfaced two yields that left 8.2 GB allocated after 3s. Two
separate issues behind that class of failure:
- The yield only unloaded loaded_models[0]. Ollama can hold several models resident
(OLLAMA_MAX_LOADED_MODELS), so releasing the first left the rest allocated. It now
unloads every resident model concurrently. This box runs with the limit at 1, so
the change is defensive here rather than a fix for the observed case.
- The observed 8.2 GB stalls happened while ComfyUI was starting and Ollama had a
generation in flight; Ollama will not unload mid-request. A 3s ceiling reported a
timeout for a model that was simply busy finishing. Raised to 10s -- waiting longer
is the safer failure mode, since the alternative is diffusion allocating into VRAM
that is still occupied.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sweeping a knob the driver ignores measures nothing but benchmark noise, and the
tuner would then confidently report 'best = the highest value tried'. On this box
(driver 595.84) nvidia-settings accepts GPUGraphicsClockOffset/GPUMemoryTransferRate
Offset and silently discards them: assigning 0 reports success and reads back 250.
The ollama profile's core_offset_mhz=35 and mem_offset_mhz=200 have therefore been
doing nothing.
- _knob_effective() applies a probe value and confirms the hardware actually moved
before any sweep starts, choosing the candidate furthest from the current reading
(probing with the maximum fails when the card already sits at its top clock).
- Adds discrete clock-lock knobs (lock_mem_mhz, lock_core_max) driven by the card's
own supported-clock list, since -lmc/-lgc do work where offsets do not.
- Reasoning models return their output in 'thinking' with an empty 'response', which
the degeneracy check was flagging as corruption. Token count is now the primary
signal.
- Separates gain-vs-current-setting from gain-vs-slowest-value-tried. Reporting the
latter as 'gain vs baseline' implied a +102% speedup that nobody would observe.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Nine changes, in rough order of how much they affect real behaviour:
1. VRAM yield is now a barrier. Posting keep_alive:0 only asks Ollama to unload;
measured here, the HTTP call returns in 63ms while the driver takes a further
77ms to release 14.9GB. Returning inside that window is how ComfyUI ends up
allocating into VRAM that is still occupied. instant_free_ollama_vram() polls
NVML until the allocation is actually gone and reports request/confirm split.
2. ComfyUI VRAM is no longer purged 1.5s after every prompt, which forced a full
checkpoint reload on each workflow iteration. It is held for 30s of genuinely
empty queue, with an immediate purge when Ollama actually asks for the memory.
3. Cache-hit classification uses achieved bandwidth (size / load duration) rather
than a fixed `load_duration < 2500ms`. That constant called a 12.9GB model read
at 2.9GB/s a cold load, and a 0.5GB model read from NVMe a cache hit.
4. Page-cache residency is measured, not assumed. mincore(2) reported 128GB
resident on a box with 46GB of page cache: the kernel only permits page-cache
introspection on files you own, and the Ollama blobs are owned by uid ollama,
for which mincore answers "all resident" instead of failing. Uses cachestat(2)
where permitted and a randomised read-rate probe elsewhere, labelling which was
used. Fixed-offset probing was self-fulfilling, so windows are random and cold
ones are returned with FADV_DONTNEED.
5. Warming is budgeted and ranked by recency/frequency instead of reading every
file top-to-bottom, which on 64GB of RAM just evicts whatever was warmed first.
6. Telemetry and events persist to SQLite (~0.38 MB/hour) instead of living in a
50-entry in-memory deque, so /api/analytics/profiles can finally answer whether
an overclock profile actually delivers more tok/s.
7. Thermal governor walks the overclock back on sustained heat or hardware
throttling, with hysteresis, fed from the existing sampler.
8. Autotune sweeps a clock offset, benchmarks decode at each step, watches for Xid
errors and degenerate output, and restores the profile in a finally block.
9. Stock clocks/power/fans are restored on shutdown and via systemd ExecStopPost.
Nothing previously undid a locked clock or a manually pinned fan.
Also: one shared 1Hz telemetry sampler fanned out to SSE subscribers rather than
every client re-running the whole snapshot; wall-clock timestamps in place of the
event loop's monotonic clock; cached nvidia-smi shell-outs; quieter httpx logging.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>