Tune profiles from measurement; add a diffusion benchmark to close the loop
The profiles were hand-written and had never been checked against the hardware. Adding a ComfyUI benchmark alongside the existing decode one made the compute side measurable for the first time, and most of what the profiles configured turned out to do nothing. Measured on this card (RTX 4080 SUPER, driver 595.84): - LLM decode is not power-bound: 73.0-73.5 tok/s flat from 222W to 370W, with the card never drawing more than 224W at any limit. The ollama profile's 370W did nothing. - Diffusion is power-bound: 5.48 it/s @222W rising to 6.71 @370W, so comfy's 370W is worth a real +2.8% over the 320W stock default. - Clock locks did nothing for either workload: 72.6 tok/s locked at 11251MHz vs 72.7 unlocked; 6.77 it/s locked at 3105MHz vs 6.73 unlocked, and 6.78 at 2400MHz. - Memory bandwidth is still the decode bottleneck (5001MHz halves throughput to 35.9 tok/s), confirming the profile's premise -- the card just gets there unaided. - Fans: 48,435 samples show 81C all-time max and zero thermal throttle events, while the ollama profile held 49.6C average by running fans at 87%. All profiles now use automatic fans and let the thermal governor escalate on demand. Code changes supporting that: - _diffusion_benchmark() queues a fixed SDXL graph via ComfyUI's API. The seed must vary per run: ComfyUI caches by node inputs, so a fixed seed returned in ~1ms without executing. Implausibly fast results are now rejected as cache hits rather than recorded as record scores. - The arbitrator's automatic profile switching is suspended during a sweep. A diffusion benchmark trips trigger_comfy_priority, which reapplies the whole profile and would silently overwrite the clock being measured. - _supported_clocks() queries the mem,gr pair; asking for a single field returned one column and reading index 1 yielded an empty list rather than an error. Graphics clocks are subsampled (the card enumerates 194 of them) and lock sweeps include an explicit unlocked control step. - offsets_supported() probes once and apply_profile skips inert offset levers with an explanation instead of pretending they applied. - Profiles carry a 'measured' field recording the evidence behind each setting. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
69
README.md
69
README.md
@@ -33,17 +33,55 @@
|
||||
* **Budgeted, Ranked Warming**: This box has 64 GB of RAM and >270 GB of model files; reading everything simply evicts whatever was warmed first. Files are ranked by recency/frequency (from the persisted event log) and warmed until a byte budget is spent, skipping anything already resident. `GET /api/warm-plan` previews the decision without executing it.
|
||||
* **Memory Telemetry**: Real-time breakdown of Total Host RAM, Applications Memory, Active Model Page Cache, Free Memory, and measured Cache Residency Ratio.
|
||||
|
||||
### 🎛️ Dynamic Overclocking & Thermal Management
|
||||
* **Workload-Aware Overclock Profiles**:
|
||||
* **`ollama` Profile (Memory-Bandwidth Bound)**: Max 370W power limit, +150 MHz Core Offset, +825 MHz Memory Offset, and 100% fan speed for maximum prompt eval / generation bandwidth.
|
||||
* **`comfy` Profile (Compute Bound)**: Max 370W power limit, +100 MHz Core Offset, +500 MHz Memory Offset, Core Clock locked to 2900–3105 MHz, and 75% fan speed for maximum diffusion compute.
|
||||
* **`balanced` Profile (Stock/General Purpose)**: Unlocked 370W power limit with stock dynamic boost curves and automatic fan control.
|
||||
* **Hardware Actuation Hierarchy**:
|
||||
* Level 1: Power Limit Control (`nvidia-smi -pl 370`).
|
||||
* Level 2: Core & Memory Clock Locking (`nvidia-smi -lgc` / `-lmc`).
|
||||
* Level 3: Clock Offsets via headless X display (`:8`) with Coolbits support (`nvidia-settings`).
|
||||
* **Hardware Fan Control**: Switch between `auto` and `manual` PWM control (30%–100%) with synchronized dual-fan actuation (`[fan:0]` and `[fan:1]`).
|
||||
* **Automated Lockstep Profile Switching**: AutoArbitrator automatically switches hardware profiles in lockstep with the active workload (`comfy` on generation start, `ollama` on completion).
|
||||
### 🎛️ Measured Overclock Profiles & Thermal Management
|
||||
|
||||
Every profile setting in this repo is now backed by a measurement from `autotune.py` on
|
||||
this specific card and driver. Several long-standing settings turned out to do nothing.
|
||||
|
||||
**What this driver actually honours** (NVIDIA 595.84, RTX 4080 SUPER):
|
||||
|
||||
| Lever | Mechanism | Works? |
|
||||
| :--- | :--- | :--- |
|
||||
| Power limit | `nvidia-smi -pl` | ✅ Yes — and it is the only lever that changes anything measurable |
|
||||
| Core / memory clock lock | `nvidia-smi -lgc` / `-lmc` | ✅ Applies correctly, but made no measurable difference to either workload |
|
||||
| Core / memory clock offsets | `nvidia-settings -a ...Offset` | ❌ **Silently ignored.** The driver reports `assigned value 0` and the attribute still reads back `250`. Detected automatically by `offsets_supported()`; `apply_profile` now skips them and says so rather than pretending. |
|
||||
| Fan control | `nvidia-settings GPUTargetFanSpeed` | ✅ Yes |
|
||||
|
||||
**Measured results** (`POST /api/autotune/sweep`):
|
||||
|
||||
*LLM decode is not power-bound.* Throughput is flat across the card's entire power range —
|
||||
the GPU never drew more than 224 W no matter what the limit allowed:
|
||||
|
||||
| Power limit | 222 W | 259 W | 296 W | 320 W | 333 W | 370 W |
|
||||
| :--- | :--- | :--- | :--- | :--- | :--- | :--- |
|
||||
| tok/s (`qwen3.8long`) | 73.17 | 73.13 | 73.43 | 73.51 | 73.10 | 73.04 |
|
||||
|
||||
*Diffusion is power-bound.* Here the watts genuinely buy throughput:
|
||||
|
||||
| Power limit | 222 W | 259 W | 296 W | 320 W | 333 W | 370 W |
|
||||
| :--- | :--- | :--- | :--- | :--- | :--- | :--- |
|
||||
| it/s (SDXL 1024, 20 steps) | 5.48 | 6.22 | 6.50 | 6.52 | 6.63 | **6.71** |
|
||||
|
||||
*Clock locks changed nothing for either workload.* Memory clock: 72.6 tok/s locked at
|
||||
11251 MHz vs 72.7 unlocked. Core clock: 6.73 it/s unlocked vs 6.77 locked at 3105 MHz —
|
||||
and 6.78 at 2400 MHz, so diffusion here is not core-clock-bound at all.
|
||||
|
||||
*Memory bandwidth is the decode bottleneck*, confirming the profile's original premise —
|
||||
dropping the memory clock to 5001 MHz halves throughput (35.9 tok/s vs 72.6). The card
|
||||
simply reaches its top memory clock on its own; pinning it there adds nothing.
|
||||
|
||||
**Resulting profiles**:
|
||||
* **`ollama`** — 320 W (stock), no locks, automatic fans. Decode draws ~224 W and is
|
||||
bandwidth-bound, so the previous 370 W limit and 100% fan pinning bought nothing.
|
||||
* **`comfy`** — 370 W, no locks, automatic fans. The extra power is worth a measured
|
||||
**+2.8%** over the 320 W stock default.
|
||||
* **`balanced`** — stock power and boost, automatic fans.
|
||||
|
||||
**On fans**: all three profiles previously pinned the fans to manual 100%. Across 48,435
|
||||
telemetry samples this card has never exceeded **81 °C** and has logged **zero** thermal
|
||||
throttle events; the `ollama` profile was holding 49.6 °C average by running the fans at
|
||||
87%. Fans are now automatic in every profile, with the thermal governor escalating them
|
||||
only if the card actually needs it.
|
||||
|
||||
### 🌡️ Thermal Governor (closed-loop de-escalation)
|
||||
* Every overclock lever here is sticky: a profile locks clocks and pins the fans to a manual PWM, and nothing used to undo that. The governor watches the telemetry the sampler already collects (so it costs no extra NVML calls) and walks the overclock back through a four-step derate ladder when the card runs hot or reports a hardware throttle.
|
||||
@@ -51,9 +89,12 @@
|
||||
* **Guaranteed restore**: stock clocks, default power limit and automatic fans are restored by the server's shutdown hook *and* by a systemd `ExecStopPost=`, so a `SIGKILL` cannot leave the card with locked clocks and fans pinned at 100%.
|
||||
|
||||
### 🔬 Overclock Autotune (`autotune.py`)
|
||||
* Walks a clock offset upward, running a fixed decode benchmark at each step, and reports the **fastest stable** value with its measured gain over baseline.
|
||||
* Sweeps a knob (`power_limit_w`, `lock_mem_mhz`, `lock_core_max`, clock offsets) and reports the **fastest stable** value.
|
||||
* **Both workloads are measurable.** `workload=ollama` benchmarks decode throughput in tok/s; `workload=comfy` queues a fixed SDXL 1024/20-step graph through ComfyUI's API and measures it/s. Without the second one there was no way to tell whether the compute-oriented `comfy` profile was doing anything at all — and it was not.
|
||||
* **Refuses to sweep a knob the driver ignores.** A preflight applies a probe value and confirms the hardware moved; the probe is chosen as the candidate furthest from the current reading, since probing with the maximum proves nothing when the card already sits there. This is what caught the silently-discarded clock offsets.
|
||||
* **Instability detection**: kernel `Xid`/`NVRM` messages via `journalctl -k`, benchmark failure, degenerate output, and a temperature ceiling. The sweep stops climbing the moment a step looks unstable.
|
||||
* **Safety**: refuses to start while ComfyUI is executing, and restores the original profile in a `finally` block — including on exception or cancellation.
|
||||
* **Honest gain reporting**: gain against the profile's *current* setting is reported separately from the spread across values tried. Conflating them turns a flat result into a headline "+102%".
|
||||
* **Safety**: refuses to start while ComfyUI is executing, suspends the arbitrator's automatic profile switching for the duration (otherwise a diffusion benchmark trips `trigger_comfy_priority`, which reapplies the whole profile and overwrites the clock being measured), and restores the original profile in a `finally` block — including on exception or cancellation.
|
||||
|
||||
### 🗄️ Persistent Telemetry Store (`telemetry_store.py`)
|
||||
* Swap history used to be an in-memory `deque(maxlen=50)` that evaporated on every restart. Telemetry and events now persist to SQLite (WAL, single writer thread, batched 1 Hz inserts, automatic retention pruning) at roughly **0.4 MB per hour**.
|
||||
@@ -161,7 +202,7 @@ The HyperSwap server runs on port `9090` by default. Interactive OpenAPI/Swagger
|
||||
| `/api/governor` | `GET` / `POST` | Current derate level and why; enable/disable, or clear an active derate. |
|
||||
| `/api/overclock/restore` | `POST` | Drop all clock locks and offsets, restore default power limit and automatic fans. |
|
||||
| `/api/autotune` | `GET` | Sweep progress, last result, and every recorded autotune step. |
|
||||
| `/api/autotune/sweep` | `POST` | Walk a clock offset upward, measuring tok/s and watching for instability at each step. |
|
||||
| `/api/autotune/sweep` | `POST` | Sweep a knob against a real workload (`workload`: `ollama` decode tok/s, `comfy` SDXL it/s), verifying the knob moves the hardware first. |
|
||||
| `/api/autotune/cancel` | `POST` | Stop the current sweep after the step in flight; the profile is restored either way. |
|
||||
|
||||
---
|
||||
|
||||
Reference in New Issue
Block a user