The repo is about to be pushed to a remote, so this covers what should never
travel with it and what should not be baked into the source.
.gitignore now covers credentials (.env, keys, tokens, .netrc), host-local
config (*.local.json), the SQLite telemetry store and its WAL sidecars, logs,
benchmark and sweep output, and the timestamped .bak files this project has
accumulated before. Verified that no currently tracked file is caught by the
new patterns.
Also removed /home/drjones from tracked source, which an ignore file cannot
help with. BASE_DIR now derives from the module's own location, the ComfyUI
model directory falls back to ~/ComfyUI/models, and start_manager.sh resolves
its interpreter through $HOME with a python3 fallback. All three still resolve
to exactly the same paths on this machine; they just no longer hardcode one
user's home directory into a published repository.
Note for the record: the git history was scanned across all refs and contains
no credentials. The password visible in `git remote -v` lives only in
.git/config, which is never pushed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Calibration. The same 12.87GB model loaded through Ollama on this box:
3.1% resident (FADV_DONTNEED) -> 34.3s -> 0.38 GB/s
100% resident (force-warmed) -> 4.9s -> 2.63 GB/s
The thresholds had been guessed from PCIe bus bandwidth: cache hit at >=5 GB/s. A fully
warm load only reaches 2.63 GB/s, because load_duration covers host-to-device transfer
and model init as well as the file read -- the page cache itself reads at 6.4 GB/s. The
5 GB/s bar was therefore unreachable, and every warm load was being reported as a
partial hit. Now 2.0 / 0.8 GB/s, either side of the measured 6.9x separation.
Warm-skip was also unsafe. A 12.87GB blob was skipped as already resident on the
strength of twelve 2MB probe windows, then loaded at 2.44 GB/s. Skipping now requires
warm_confident: an exact cachestat reading, or a probe finding every one of 32 denser
samples resident. warm_file_to_ram/warm_ollama_blob take force=True, exposed on the
warm-model endpoint, whose Pydantic model was missing the field entirely.
MCP parity: the server had drifted well behind the REST API. Adds tools for measured
residency, warm planning, VRAM requests, per-profile analytics, thermal governor
control, overclock status/apply/restore, and autotune sweeps plus status -- 23 tools
and 6 resources, up from 12 and 3. The telemetry store now starts in __main__ rather
than at import scope, since server.py imports this module for the benchmark tool.
README: replaced the remaining theoretical claims (31.5 GB/s bus rate, sub-1.5s loads,
15ms yields) with the measured numbers.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Nine changes, in rough order of how much they affect real behaviour:
1. VRAM yield is now a barrier. Posting keep_alive:0 only asks Ollama to unload;
measured here, the HTTP call returns in 63ms while the driver takes a further
77ms to release 14.9GB. Returning inside that window is how ComfyUI ends up
allocating into VRAM that is still occupied. instant_free_ollama_vram() polls
NVML until the allocation is actually gone and reports request/confirm split.
2. ComfyUI VRAM is no longer purged 1.5s after every prompt, which forced a full
checkpoint reload on each workflow iteration. It is held for 30s of genuinely
empty queue, with an immediate purge when Ollama actually asks for the memory.
3. Cache-hit classification uses achieved bandwidth (size / load duration) rather
than a fixed `load_duration < 2500ms`. That constant called a 12.9GB model read
at 2.9GB/s a cold load, and a 0.5GB model read from NVMe a cache hit.
4. Page-cache residency is measured, not assumed. mincore(2) reported 128GB
resident on a box with 46GB of page cache: the kernel only permits page-cache
introspection on files you own, and the Ollama blobs are owned by uid ollama,
for which mincore answers "all resident" instead of failing. Uses cachestat(2)
where permitted and a randomised read-rate probe elsewhere, labelling which was
used. Fixed-offset probing was self-fulfilling, so windows are random and cold
ones are returned with FADV_DONTNEED.
5. Warming is budgeted and ranked by recency/frequency instead of reading every
file top-to-bottom, which on 64GB of RAM just evicts whatever was warmed first.
6. Telemetry and events persist to SQLite (~0.38 MB/hour) instead of living in a
50-entry in-memory deque, so /api/analytics/profiles can finally answer whether
an overclock profile actually delivers more tok/s.
7. Thermal governor walks the overclock back on sustained heat or hardware
throttling, with hysteresis, fed from the existing sampler.
8. Autotune sweeps a clock offset, benchmarks decode at each step, watches for Xid
errors and degenerate output, and restores the profile in a finally block.
9. Stock clocks/power/fans are restored on shutdown and via systemd ExecStopPost.
Nothing previously undid a locked clock or a manually pinned fan.
Also: one shared 1Hz telemetry sampler fanned out to SSE subscribers rather than
every client re-running the whole snapshot; wall-clock timestamps in place of the
event loop's monotonic clock; cached nvidia-smi shell-outs; quieter httpx logging.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>