Tune profiles from measurement; add a diffusion benchmark to close the loop
The profiles were hand-written and had never been checked against the hardware. Adding a ComfyUI benchmark alongside the existing decode one made the compute side measurable for the first time, and most of what the profiles configured turned out to do nothing. Measured on this card (RTX 4080 SUPER, driver 595.84): - LLM decode is not power-bound: 73.0-73.5 tok/s flat from 222W to 370W, with the card never drawing more than 224W at any limit. The ollama profile's 370W did nothing. - Diffusion is power-bound: 5.48 it/s @222W rising to 6.71 @370W, so comfy's 370W is worth a real +2.8% over the 320W stock default. - Clock locks did nothing for either workload: 72.6 tok/s locked at 11251MHz vs 72.7 unlocked; 6.77 it/s locked at 3105MHz vs 6.73 unlocked, and 6.78 at 2400MHz. - Memory bandwidth is still the decode bottleneck (5001MHz halves throughput to 35.9 tok/s), confirming the profile's premise -- the card just gets there unaided. - Fans: 48,435 samples show 81C all-time max and zero thermal throttle events, while the ollama profile held 49.6C average by running fans at 87%. All profiles now use automatic fans and let the thermal governor escalate on demand. Code changes supporting that: - _diffusion_benchmark() queues a fixed SDXL graph via ComfyUI's API. The seed must vary per run: ComfyUI caches by node inputs, so a fixed seed returned in ~1ms without executing. Implausibly fast results are now rejected as cache hits rather than recorded as record scores. - The arbitrator's automatic profile switching is suspended during a sweep. A diffusion benchmark trips trigger_comfy_priority, which reapplies the whole profile and would silently overwrite the clock being measured. - _supported_clocks() queries the mem,gr pair; asking for a single field returned one column and reading index 1 yielded an empty list rather than an error. Graphics clocks are subsampled (the card enumerates 194 of them) and lock sweeps include an explicit unlocked control step. - offsets_supported() probes once and apply_profile skips inert offset levers with an explanation instead of pretending they applied. - Profiles carry a 'measured' field recording the evidence behind each setting. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
69
README.md
69
README.md
@@ -33,17 +33,55 @@
|
|||||||
* **Budgeted, Ranked Warming**: This box has 64 GB of RAM and >270 GB of model files; reading everything simply evicts whatever was warmed first. Files are ranked by recency/frequency (from the persisted event log) and warmed until a byte budget is spent, skipping anything already resident. `GET /api/warm-plan` previews the decision without executing it.
|
* **Budgeted, Ranked Warming**: This box has 64 GB of RAM and >270 GB of model files; reading everything simply evicts whatever was warmed first. Files are ranked by recency/frequency (from the persisted event log) and warmed until a byte budget is spent, skipping anything already resident. `GET /api/warm-plan` previews the decision without executing it.
|
||||||
* **Memory Telemetry**: Real-time breakdown of Total Host RAM, Applications Memory, Active Model Page Cache, Free Memory, and measured Cache Residency Ratio.
|
* **Memory Telemetry**: Real-time breakdown of Total Host RAM, Applications Memory, Active Model Page Cache, Free Memory, and measured Cache Residency Ratio.
|
||||||
|
|
||||||
### 🎛️ Dynamic Overclocking & Thermal Management
|
### 🎛️ Measured Overclock Profiles & Thermal Management
|
||||||
* **Workload-Aware Overclock Profiles**:
|
|
||||||
* **`ollama` Profile (Memory-Bandwidth Bound)**: Max 370W power limit, +150 MHz Core Offset, +825 MHz Memory Offset, and 100% fan speed for maximum prompt eval / generation bandwidth.
|
Every profile setting in this repo is now backed by a measurement from `autotune.py` on
|
||||||
* **`comfy` Profile (Compute Bound)**: Max 370W power limit, +100 MHz Core Offset, +500 MHz Memory Offset, Core Clock locked to 2900–3105 MHz, and 75% fan speed for maximum diffusion compute.
|
this specific card and driver. Several long-standing settings turned out to do nothing.
|
||||||
* **`balanced` Profile (Stock/General Purpose)**: Unlocked 370W power limit with stock dynamic boost curves and automatic fan control.
|
|
||||||
* **Hardware Actuation Hierarchy**:
|
**What this driver actually honours** (NVIDIA 595.84, RTX 4080 SUPER):
|
||||||
* Level 1: Power Limit Control (`nvidia-smi -pl 370`).
|
|
||||||
* Level 2: Core & Memory Clock Locking (`nvidia-smi -lgc` / `-lmc`).
|
| Lever | Mechanism | Works? |
|
||||||
* Level 3: Clock Offsets via headless X display (`:8`) with Coolbits support (`nvidia-settings`).
|
| :--- | :--- | :--- |
|
||||||
* **Hardware Fan Control**: Switch between `auto` and `manual` PWM control (30%–100%) with synchronized dual-fan actuation (`[fan:0]` and `[fan:1]`).
|
| Power limit | `nvidia-smi -pl` | ✅ Yes — and it is the only lever that changes anything measurable |
|
||||||
* **Automated Lockstep Profile Switching**: AutoArbitrator automatically switches hardware profiles in lockstep with the active workload (`comfy` on generation start, `ollama` on completion).
|
| Core / memory clock lock | `nvidia-smi -lgc` / `-lmc` | ✅ Applies correctly, but made no measurable difference to either workload |
|
||||||
|
| Core / memory clock offsets | `nvidia-settings -a ...Offset` | ❌ **Silently ignored.** The driver reports `assigned value 0` and the attribute still reads back `250`. Detected automatically by `offsets_supported()`; `apply_profile` now skips them and says so rather than pretending. |
|
||||||
|
| Fan control | `nvidia-settings GPUTargetFanSpeed` | ✅ Yes |
|
||||||
|
|
||||||
|
**Measured results** (`POST /api/autotune/sweep`):
|
||||||
|
|
||||||
|
*LLM decode is not power-bound.* Throughput is flat across the card's entire power range —
|
||||||
|
the GPU never drew more than 224 W no matter what the limit allowed:
|
||||||
|
|
||||||
|
| Power limit | 222 W | 259 W | 296 W | 320 W | 333 W | 370 W |
|
||||||
|
| :--- | :--- | :--- | :--- | :--- | :--- | :--- |
|
||||||
|
| tok/s (`qwen3.8long`) | 73.17 | 73.13 | 73.43 | 73.51 | 73.10 | 73.04 |
|
||||||
|
|
||||||
|
*Diffusion is power-bound.* Here the watts genuinely buy throughput:
|
||||||
|
|
||||||
|
| Power limit | 222 W | 259 W | 296 W | 320 W | 333 W | 370 W |
|
||||||
|
| :--- | :--- | :--- | :--- | :--- | :--- | :--- |
|
||||||
|
| it/s (SDXL 1024, 20 steps) | 5.48 | 6.22 | 6.50 | 6.52 | 6.63 | **6.71** |
|
||||||
|
|
||||||
|
*Clock locks changed nothing for either workload.* Memory clock: 72.6 tok/s locked at
|
||||||
|
11251 MHz vs 72.7 unlocked. Core clock: 6.73 it/s unlocked vs 6.77 locked at 3105 MHz —
|
||||||
|
and 6.78 at 2400 MHz, so diffusion here is not core-clock-bound at all.
|
||||||
|
|
||||||
|
*Memory bandwidth is the decode bottleneck*, confirming the profile's original premise —
|
||||||
|
dropping the memory clock to 5001 MHz halves throughput (35.9 tok/s vs 72.6). The card
|
||||||
|
simply reaches its top memory clock on its own; pinning it there adds nothing.
|
||||||
|
|
||||||
|
**Resulting profiles**:
|
||||||
|
* **`ollama`** — 320 W (stock), no locks, automatic fans. Decode draws ~224 W and is
|
||||||
|
bandwidth-bound, so the previous 370 W limit and 100% fan pinning bought nothing.
|
||||||
|
* **`comfy`** — 370 W, no locks, automatic fans. The extra power is worth a measured
|
||||||
|
**+2.8%** over the 320 W stock default.
|
||||||
|
* **`balanced`** — stock power and boost, automatic fans.
|
||||||
|
|
||||||
|
**On fans**: all three profiles previously pinned the fans to manual 100%. Across 48,435
|
||||||
|
telemetry samples this card has never exceeded **81 °C** and has logged **zero** thermal
|
||||||
|
throttle events; the `ollama` profile was holding 49.6 °C average by running the fans at
|
||||||
|
87%. Fans are now automatic in every profile, with the thermal governor escalating them
|
||||||
|
only if the card actually needs it.
|
||||||
|
|
||||||
### 🌡️ Thermal Governor (closed-loop de-escalation)
|
### 🌡️ Thermal Governor (closed-loop de-escalation)
|
||||||
* Every overclock lever here is sticky: a profile locks clocks and pins the fans to a manual PWM, and nothing used to undo that. The governor watches the telemetry the sampler already collects (so it costs no extra NVML calls) and walks the overclock back through a four-step derate ladder when the card runs hot or reports a hardware throttle.
|
* Every overclock lever here is sticky: a profile locks clocks and pins the fans to a manual PWM, and nothing used to undo that. The governor watches the telemetry the sampler already collects (so it costs no extra NVML calls) and walks the overclock back through a four-step derate ladder when the card runs hot or reports a hardware throttle.
|
||||||
@@ -51,9 +89,12 @@
|
|||||||
* **Guaranteed restore**: stock clocks, default power limit and automatic fans are restored by the server's shutdown hook *and* by a systemd `ExecStopPost=`, so a `SIGKILL` cannot leave the card with locked clocks and fans pinned at 100%.
|
* **Guaranteed restore**: stock clocks, default power limit and automatic fans are restored by the server's shutdown hook *and* by a systemd `ExecStopPost=`, so a `SIGKILL` cannot leave the card with locked clocks and fans pinned at 100%.
|
||||||
|
|
||||||
### 🔬 Overclock Autotune (`autotune.py`)
|
### 🔬 Overclock Autotune (`autotune.py`)
|
||||||
* Walks a clock offset upward, running a fixed decode benchmark at each step, and reports the **fastest stable** value with its measured gain over baseline.
|
* Sweeps a knob (`power_limit_w`, `lock_mem_mhz`, `lock_core_max`, clock offsets) and reports the **fastest stable** value.
|
||||||
|
* **Both workloads are measurable.** `workload=ollama` benchmarks decode throughput in tok/s; `workload=comfy` queues a fixed SDXL 1024/20-step graph through ComfyUI's API and measures it/s. Without the second one there was no way to tell whether the compute-oriented `comfy` profile was doing anything at all — and it was not.
|
||||||
|
* **Refuses to sweep a knob the driver ignores.** A preflight applies a probe value and confirms the hardware moved; the probe is chosen as the candidate furthest from the current reading, since probing with the maximum proves nothing when the card already sits there. This is what caught the silently-discarded clock offsets.
|
||||||
* **Instability detection**: kernel `Xid`/`NVRM` messages via `journalctl -k`, benchmark failure, degenerate output, and a temperature ceiling. The sweep stops climbing the moment a step looks unstable.
|
* **Instability detection**: kernel `Xid`/`NVRM` messages via `journalctl -k`, benchmark failure, degenerate output, and a temperature ceiling. The sweep stops climbing the moment a step looks unstable.
|
||||||
* **Safety**: refuses to start while ComfyUI is executing, and restores the original profile in a `finally` block — including on exception or cancellation.
|
* **Honest gain reporting**: gain against the profile's *current* setting is reported separately from the spread across values tried. Conflating them turns a flat result into a headline "+102%".
|
||||||
|
* **Safety**: refuses to start while ComfyUI is executing, suspends the arbitrator's automatic profile switching for the duration (otherwise a diffusion benchmark trips `trigger_comfy_priority`, which reapplies the whole profile and overwrites the clock being measured), and restores the original profile in a `finally` block — including on exception or cancellation.
|
||||||
|
|
||||||
### 🗄️ Persistent Telemetry Store (`telemetry_store.py`)
|
### 🗄️ Persistent Telemetry Store (`telemetry_store.py`)
|
||||||
* Swap history used to be an in-memory `deque(maxlen=50)` that evaporated on every restart. Telemetry and events now persist to SQLite (WAL, single writer thread, batched 1 Hz inserts, automatic retention pruning) at roughly **0.4 MB per hour**.
|
* Swap history used to be an in-memory `deque(maxlen=50)` that evaporated on every restart. Telemetry and events now persist to SQLite (WAL, single writer thread, batched 1 Hz inserts, automatic retention pruning) at roughly **0.4 MB per hour**.
|
||||||
@@ -161,7 +202,7 @@ The HyperSwap server runs on port `9090` by default. Interactive OpenAPI/Swagger
|
|||||||
| `/api/governor` | `GET` / `POST` | Current derate level and why; enable/disable, or clear an active derate. |
|
| `/api/governor` | `GET` / `POST` | Current derate level and why; enable/disable, or clear an active derate. |
|
||||||
| `/api/overclock/restore` | `POST` | Drop all clock locks and offsets, restore default power limit and automatic fans. |
|
| `/api/overclock/restore` | `POST` | Drop all clock locks and offsets, restore default power limit and automatic fans. |
|
||||||
| `/api/autotune` | `GET` | Sweep progress, last result, and every recorded autotune step. |
|
| `/api/autotune` | `GET` | Sweep progress, last result, and every recorded autotune step. |
|
||||||
| `/api/autotune/sweep` | `POST` | Walk a clock offset upward, measuring tok/s and watching for instability at each step. |
|
| `/api/autotune/sweep` | `POST` | Sweep a knob against a real workload (`workload`: `ollama` decode tok/s, `comfy` SDXL it/s), verifying the knob moves the hardware first. |
|
||||||
| `/api/autotune/cancel` | `POST` | Stop the current sweep after the step in flight; the profile is restored either way. |
|
| `/api/autotune/cancel` | `POST` | Stop the current sweep after the step in flight; the profile is restored either way. |
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|||||||
209
autotune.py
209
autotune.py
@@ -15,6 +15,7 @@ Safety properties:
|
|||||||
"""
|
"""
|
||||||
import asyncio
|
import asyncio
|
||||||
import logging
|
import logging
|
||||||
|
import random
|
||||||
import subprocess
|
import subprocess
|
||||||
import time
|
import time
|
||||||
from typing import Any, Dict, List, Optional
|
from typing import Any, Dict, List, Optional
|
||||||
@@ -46,28 +47,74 @@ KNOBS = {
|
|||||||
"values": None}, # filled from the card's supported clock list
|
"values": None}, # filled from the card's supported clock list
|
||||||
"lock_core_max": {"kind": "discrete", "verify": "clock_sm",
|
"lock_core_max": {"kind": "discrete", "verify": "clock_sm",
|
||||||
"values": None},
|
"values": None},
|
||||||
|
# The one lever this driver definitely honours. Worth knowing whether the extra
|
||||||
|
# watts actually buy throughput, or just heat and fan noise.
|
||||||
|
"power_limit_w": {"kind": "discrete", "verify": "power_limit",
|
||||||
|
"values": None, "no_unlocked": True},
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
||||||
def _supported_clocks(which: str = "mem") -> List[int]:
|
def _supported_clocks(which: str = "mem") -> List[int]:
|
||||||
"""Discrete clock values the card will actually accept for -lmc / -lgc."""
|
"""Discrete clock values the card will actually accept for -lmc / -lgc.
|
||||||
|
|
||||||
|
Always queried as the mem,gr pair: asking for a single field returns one column, and
|
||||||
|
reading index 1 from it silently yields an empty list rather than an error.
|
||||||
|
"""
|
||||||
try:
|
try:
|
||||||
proc = subprocess.run(
|
proc = subprocess.run(
|
||||||
["nvidia-smi", f"--query-supported-clocks={'mem' if which == 'mem' else 'gr'}",
|
["nvidia-smi", "--query-supported-clocks=mem,gr", "--format=csv,noheader,nounits"],
|
||||||
"--format=csv,noheader,nounits"],
|
capture_output=True, text=True, timeout=15)
|
||||||
capture_output=True, text=True, timeout=10)
|
rows = []
|
||||||
col = 0 if which == "mem" else 1
|
|
||||||
vals = set()
|
|
||||||
for line in proc.stdout.splitlines():
|
for line in proc.stdout.splitlines():
|
||||||
parts = [p.strip() for p in line.split(",")]
|
parts = [p.strip() for p in line.split(",")]
|
||||||
if len(parts) > col and parts[col].isdigit():
|
if len(parts) >= 2 and parts[0].isdigit() and parts[1].isdigit():
|
||||||
vals.add(int(parts[col]))
|
rows.append((int(parts[0]), int(parts[1])))
|
||||||
return sorted(vals)
|
if not rows:
|
||||||
|
return []
|
||||||
|
if which == "mem":
|
||||||
|
return sorted({m for m, _ in rows})
|
||||||
|
# Graphics clocks are enumerated per memory clock; take the list for the highest
|
||||||
|
# memory clock, which is the one any real workload runs at.
|
||||||
|
top_mem = max(m for m, _ in rows)
|
||||||
|
return sorted({g for m, g in rows if m == top_mem})
|
||||||
except Exception as e:
|
except Exception as e:
|
||||||
logger.debug(f"supported clock query failed: {e}")
|
logger.debug(f"supported clock query failed: {e}")
|
||||||
return []
|
return []
|
||||||
|
|
||||||
|
|
||||||
|
def _supported_power_limits(steps: int = 5) -> List[int]:
|
||||||
|
"""Power limits between the card's minimum and maximum, in even increments."""
|
||||||
|
try:
|
||||||
|
proc = subprocess.run(
|
||||||
|
["nvidia-smi", "--query-gpu=power.min_limit,power.max_limit,power.default_limit",
|
||||||
|
"--format=csv,noheader,nounits"],
|
||||||
|
capture_output=True, text=True, timeout=10)
|
||||||
|
parts = [p.strip() for p in proc.stdout.strip().split(",")]
|
||||||
|
lo, hi, default = (int(float(parts[0])), int(float(parts[1])), int(float(parts[2])))
|
||||||
|
except Exception as e:
|
||||||
|
logger.debug(f"power limit query failed: {e}")
|
||||||
|
return []
|
||||||
|
# Start at 60% of max -- below that the card is not doing useful work for these
|
||||||
|
# workloads -- and always include the stock default as a reference point.
|
||||||
|
lo = max(lo, int(hi * 0.6))
|
||||||
|
span = hi - lo
|
||||||
|
vals = {lo + round(i * span / (steps - 1)) for i in range(steps)}
|
||||||
|
vals.add(default)
|
||||||
|
return sorted(v for v in vals if lo <= v <= hi)
|
||||||
|
|
||||||
|
|
||||||
|
def _subsample(values: List[int], max_steps: int) -> List[int]:
|
||||||
|
"""Evenly spaced subset, always keeping the endpoints.
|
||||||
|
|
||||||
|
The card enumerates ~194 graphics clocks in 15 MHz increments; benchmarking every one
|
||||||
|
would take hours and tell us nothing that a handful of well-spread points does not.
|
||||||
|
"""
|
||||||
|
if len(values) <= max_steps:
|
||||||
|
return values
|
||||||
|
idx = [round(i * (len(values) - 1) / (max_steps - 1)) for i in range(max_steps)]
|
||||||
|
return sorted({values[i] for i in idx})
|
||||||
|
|
||||||
|
|
||||||
def _read_hw(field: str) -> Optional[float]:
|
def _read_hw(field: str) -> Optional[float]:
|
||||||
"""Read back the hardware state a knob is supposed to move."""
|
"""Read back the hardware state a knob is supposed to move."""
|
||||||
gpu = vram_arbitrator.get_gpu_hardware_stats()
|
gpu = vram_arbitrator.get_gpu_hardware_stats()
|
||||||
@@ -75,6 +122,8 @@ def _read_hw(field: str) -> Optional[float]:
|
|||||||
return gpu.get("clock_mem_mhz")
|
return gpu.get("clock_mem_mhz")
|
||||||
if field == "clock_sm":
|
if field == "clock_sm":
|
||||||
return gpu.get("clock_graphics_mhz")
|
return gpu.get("clock_graphics_mhz")
|
||||||
|
if field == "power_limit":
|
||||||
|
return gpu.get("power_limit_w")
|
||||||
if field in ("mem_offset", "core_offset"):
|
if field in ("mem_offset", "core_offset"):
|
||||||
r = overclock_manager._nvidia_settings(
|
r = overclock_manager._nvidia_settings(
|
||||||
"-q", f"[gpu:0]/{'GPUMemoryTransferRateOffset' if field == 'mem_offset' else 'GPUGraphicsClockOffset'}[3]")
|
"-q", f"[gpu:0]/{'GPUMemoryTransferRateOffset' if field == 'mem_offset' else 'GPUGraphicsClockOffset'}[3]")
|
||||||
@@ -174,6 +223,97 @@ async def _decode_benchmark(model: str) -> Dict[str, Any]:
|
|||||||
}
|
}
|
||||||
|
|
||||||
|
|
||||||
|
# A fixed SDXL txt2img graph. Deterministic seed/steps/resolution so every step of a
|
||||||
|
# sweep does identical work and the only variable is the clock. PreviewImage rather than
|
||||||
|
# SaveImage keeps benchmark runs out of the user's output gallery.
|
||||||
|
COMFY_BENCH_CKPT = "sd_xl_base_1.0.safetensors"
|
||||||
|
COMFY_BENCH_STEPS = 20
|
||||||
|
COMFY_BENCH_SIZE = 1024
|
||||||
|
|
||||||
|
|
||||||
|
def _comfy_workflow(ckpt: str = COMFY_BENCH_CKPT, seed: Optional[int] = None) -> Dict[str, Any]:
|
||||||
|
# The seed must vary per run. ComfyUI caches by node inputs, so a fixed seed makes the
|
||||||
|
# second and later benchmarks return in ~1ms without executing anything at all. The
|
||||||
|
# cost of the graph is identical regardless of seed, so this costs no comparability.
|
||||||
|
seed = random.randint(1, 2**31) if seed is None else seed
|
||||||
|
return {
|
||||||
|
"1": {"class_type": "CheckpointLoaderSimple", "inputs": {"ckpt_name": ckpt}},
|
||||||
|
"2": {"class_type": "CLIPTextEncode",
|
||||||
|
"inputs": {"clip": ["1", 1],
|
||||||
|
"text": "a detailed photograph of a mountain range at sunrise"}},
|
||||||
|
"3": {"class_type": "CLIPTextEncode",
|
||||||
|
"inputs": {"clip": ["1", 1], "text": "blurry, low quality"}},
|
||||||
|
"4": {"class_type": "EmptyLatentImage",
|
||||||
|
"inputs": {"width": COMFY_BENCH_SIZE, "height": COMFY_BENCH_SIZE, "batch_size": 1}},
|
||||||
|
"5": {"class_type": "KSampler",
|
||||||
|
"inputs": {"model": ["1", 0], "positive": ["2", 0], "negative": ["3", 0],
|
||||||
|
"latent_image": ["4", 0], "seed": seed, "steps": COMFY_BENCH_STEPS,
|
||||||
|
"cfg": 7.0, "sampler_name": "euler", "scheduler": "normal",
|
||||||
|
"denoise": 1.0}},
|
||||||
|
"6": {"class_type": "VAEDecode", "inputs": {"samples": ["5", 0], "vae": ["1", 2]}},
|
||||||
|
"7": {"class_type": "PreviewImage", "inputs": {"images": ["6", 0]}},
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
async def _diffusion_benchmark(ckpt: str = COMFY_BENCH_CKPT,
|
||||||
|
timeout_s: float = 300.0) -> Dict[str, Any]:
|
||||||
|
"""Queue one fixed SDXL graph and time it. This is the compute-bound counterpart to
|
||||||
|
the decode benchmark, and the only way to tell whether the 'comfy' profile helps."""
|
||||||
|
client = vram_arbitrator._client(vram_arbitrator.COMFY_API_BASE, 30.0)
|
||||||
|
t0 = time.perf_counter()
|
||||||
|
try:
|
||||||
|
resp = await client.post("/prompt", json={"prompt": _comfy_workflow(ckpt),
|
||||||
|
"client_id": "hyperswap-autotune"})
|
||||||
|
if resp.status_code != 200:
|
||||||
|
return {"ok": False, "error": f"queue failed HTTP {resp.status_code}: {resp.text[:200]}"}
|
||||||
|
prompt_id = resp.json().get("prompt_id")
|
||||||
|
except Exception as e:
|
||||||
|
return {"ok": False, "error": f"queue failed: {e}"}
|
||||||
|
|
||||||
|
while (time.perf_counter() - t0) < timeout_s:
|
||||||
|
await asyncio.sleep(0.25)
|
||||||
|
try:
|
||||||
|
h = await client.get(f"/history/{prompt_id}")
|
||||||
|
if h.status_code != 200:
|
||||||
|
continue
|
||||||
|
entry = (h.json() or {}).get(prompt_id)
|
||||||
|
if not entry:
|
||||||
|
continue
|
||||||
|
status = entry.get("status", {})
|
||||||
|
if status.get("status_str") == "error" or not status.get("completed", True):
|
||||||
|
if status.get("status_str") == "error":
|
||||||
|
return {"ok": False, "error": "ComfyUI reported an execution error",
|
||||||
|
"wall_ms": round((time.perf_counter() - t0) * 1000, 2)}
|
||||||
|
if status.get("completed"):
|
||||||
|
wall = time.perf_counter() - t0
|
||||||
|
# ComfyUI stamps execution_start/success in the status messages; the delta
|
||||||
|
# between them excludes our polling overhead and the queue wait.
|
||||||
|
stamps = {}
|
||||||
|
for msg in status.get("messages", []):
|
||||||
|
if isinstance(msg, list) and len(msg) >= 2 and isinstance(msg[1], dict):
|
||||||
|
if "timestamp" in msg[1]:
|
||||||
|
stamps[msg[0]] = msg[1]["timestamp"]
|
||||||
|
exec_ms = None
|
||||||
|
if "execution_start" in stamps and "execution_success" in stamps:
|
||||||
|
exec_ms = round(stamps["execution_success"] - stamps["execution_start"], 2)
|
||||||
|
effective_ms = exec_ms or wall * 1000
|
||||||
|
# A graph that "finished" implausibly fast was served from ComfyUI's cache
|
||||||
|
# rather than executed; treat it as an invalid sample, not a record score.
|
||||||
|
cached = effective_ms < 250
|
||||||
|
return {
|
||||||
|
"ok": not cached,
|
||||||
|
"error": "result served from ComfyUI cache, not executed" if cached else None,
|
||||||
|
"wall_ms": round(wall * 1000, 2),
|
||||||
|
"exec_ms": exec_ms,
|
||||||
|
"steps": COMFY_BENCH_STEPS,
|
||||||
|
"it_per_sec": round(COMFY_BENCH_STEPS / (effective_ms / 1000), 3),
|
||||||
|
"degenerate": cached,
|
||||||
|
}
|
||||||
|
except Exception:
|
||||||
|
continue
|
||||||
|
return {"ok": False, "error": f"diffusion benchmark timed out after {timeout_s}s"}
|
||||||
|
|
||||||
|
|
||||||
class SweepState:
|
class SweepState:
|
||||||
def __init__(self) -> None:
|
def __init__(self) -> None:
|
||||||
self.running = False
|
self.running = False
|
||||||
@@ -187,11 +327,14 @@ state = SweepState()
|
|||||||
|
|
||||||
async def sweep(knob: str = "mem_offset_mhz",
|
async def sweep(knob: str = "mem_offset_mhz",
|
||||||
profile: str = "ollama",
|
profile: str = "ollama",
|
||||||
|
workload: str = "auto",
|
||||||
model: Optional[str] = None,
|
model: Optional[str] = None,
|
||||||
start: Optional[int] = None,
|
start: Optional[int] = None,
|
||||||
stop: Optional[int] = None,
|
stop: Optional[int] = None,
|
||||||
step: Optional[int] = None,
|
step: Optional[int] = None,
|
||||||
repeats: int = 1,
|
repeats: int = 1,
|
||||||
|
max_steps: int = 6,
|
||||||
|
include_unlocked: bool = True,
|
||||||
apply_best: bool = False) -> Dict[str, Any]:
|
apply_best: bool = False) -> Dict[str, Any]:
|
||||||
"""Sweep one clock offset and return the fastest stable value."""
|
"""Sweep one clock offset and return the fastest stable value."""
|
||||||
if knob not in KNOBS:
|
if knob not in KNOBS:
|
||||||
@@ -203,7 +346,23 @@ async def sweep(knob: str = "mem_offset_mhz",
|
|||||||
if comfy.get("executing") or comfy.get("queue_remaining"):
|
if comfy.get("executing") or comfy.get("queue_remaining"):
|
||||||
return {"success": False, "error": "ComfyUI is busy; refusing to change clocks mid-render"}
|
return {"success": False, "error": "ComfyUI is busy; refusing to change clocks mid-render"}
|
||||||
|
|
||||||
if not model:
|
# 'auto': tune the workload the profile is actually for.
|
||||||
|
if workload == "auto":
|
||||||
|
workload = "comfy" if profile == "comfy" else "ollama"
|
||||||
|
if workload not in ("ollama", "comfy"):
|
||||||
|
return {"success": False, "error": "workload must be 'ollama', 'comfy' or 'auto'"}
|
||||||
|
if workload == "comfy" and not comfy.get("online"):
|
||||||
|
return {"success": False, "error": "ComfyUI is not reachable; cannot run a diffusion sweep"}
|
||||||
|
|
||||||
|
if workload == "comfy":
|
||||||
|
model = model or COMFY_BENCH_CKPT
|
||||||
|
benchmark = lambda: _diffusion_benchmark(model)
|
||||||
|
metric = "it_per_sec"
|
||||||
|
else:
|
||||||
|
benchmark = lambda: _decode_benchmark(model)
|
||||||
|
metric = "tokens_per_sec"
|
||||||
|
|
||||||
|
if workload == "ollama" and not model:
|
||||||
ollama = await vram_arbitrator.get_ollama_live_state()
|
ollama = await vram_arbitrator.get_ollama_live_state()
|
||||||
model = ollama.get("active_model_name")
|
model = ollama.get("active_model_name")
|
||||||
if not model:
|
if not model:
|
||||||
@@ -211,17 +370,26 @@ async def sweep(knob: str = "mem_offset_mhz",
|
|||||||
if not installed:
|
if not installed:
|
||||||
return {"success": False, "error": "no Ollama model available to benchmark"}
|
return {"success": False, "error": "no Ollama model available to benchmark"}
|
||||||
model = installed[0].get("name")
|
model = installed[0].get("name")
|
||||||
|
benchmark = lambda: _decode_benchmark(model)
|
||||||
|
|
||||||
defaults = KNOBS[knob]
|
defaults = KNOBS[knob]
|
||||||
if defaults.get("kind") == "discrete":
|
if defaults.get("kind") == "discrete":
|
||||||
supported = defaults.get("values") or _supported_clocks(
|
if knob == "power_limit_w":
|
||||||
"mem" if knob == "lock_mem_mhz" else "gr")
|
supported = _supported_power_limits()
|
||||||
|
else:
|
||||||
|
supported = defaults.get("values") or _supported_clocks(
|
||||||
|
"mem" if knob == "lock_mem_mhz" else "gr")
|
||||||
if not supported:
|
if not supported:
|
||||||
return {"success": False, "error": f"card reported no supported clocks for {knob}"}
|
return {"success": False, "error": f"card reported no supported clocks for {knob}"}
|
||||||
values = [v for v in supported
|
values = [v for v in supported
|
||||||
if (start is None or v >= start) and (stop is None or v <= stop)]
|
if (start is None or v >= start) and (stop is None or v <= stop)]
|
||||||
if not values:
|
if not values:
|
||||||
return {"success": False, "error": f"no supported values in range; card offers {supported}"}
|
return {"success": False, "error": f"no supported values in range; card offers {supported}"}
|
||||||
|
values = _subsample(values, max_steps or 6)
|
||||||
|
# 0 means "no lock at all". That is the honest control for a profile whose whole
|
||||||
|
# premise is that locking the clock beats letting the card boost on its own.
|
||||||
|
if include_unlocked and not defaults.get("no_unlocked"):
|
||||||
|
values = [0] + values
|
||||||
start, stop, step = values[0], values[-1], None
|
start, stop, step = values[0], values[-1], None
|
||||||
else:
|
else:
|
||||||
start = defaults["default_start"] if start is None else start
|
start = defaults["default_start"] if start is None else start
|
||||||
@@ -235,7 +403,8 @@ async def sweep(knob: str = "mem_offset_mhz",
|
|||||||
baseline_value = int(baseline_cfg.get(knob, 0) or 0)
|
baseline_value = int(baseline_cfg.get(knob, 0) or 0)
|
||||||
|
|
||||||
# Refuse to sweep a knob the driver is going to ignore.
|
# Refuse to sweep a knob the driver is going to ignore.
|
||||||
effectiveness = _knob_effective(knob, profile, values, baseline_value)
|
effectiveness = _knob_effective(knob, profile, [v for v in values if v] or values,
|
||||||
|
baseline_value)
|
||||||
if not effectiveness["effective"]:
|
if not effectiveness["effective"]:
|
||||||
return {
|
return {
|
||||||
"success": False,
|
"success": False,
|
||||||
@@ -249,8 +418,9 @@ async def sweep(knob: str = "mem_offset_mhz",
|
|||||||
t_start = time.time()
|
t_start = time.time()
|
||||||
|
|
||||||
try:
|
try:
|
||||||
# Load the model once up front so the first step does not pay the load cost.
|
# Warm-up: load weights once up front so the first step does not pay the load cost.
|
||||||
await _decode_benchmark(model)
|
vram_arbitrator.arbitrator.suspend_oc("autotune sweep")
|
||||||
|
await benchmark()
|
||||||
|
|
||||||
for value in values:
|
for value in values:
|
||||||
if state.cancel:
|
if state.cancel:
|
||||||
@@ -261,7 +431,7 @@ async def sweep(knob: str = "mem_offset_mhz",
|
|||||||
|
|
||||||
samples = []
|
samples = []
|
||||||
for _ in range(max(repeats, 1)):
|
for _ in range(max(repeats, 1)):
|
||||||
samples.append(await _decode_benchmark(model))
|
samples.append(await benchmark())
|
||||||
if state.cancel:
|
if state.cancel:
|
||||||
break
|
break
|
||||||
|
|
||||||
@@ -278,11 +448,13 @@ async def sweep(knob: str = "mem_offset_mhz",
|
|||||||
if temp >= TEMP_CEILING_C:
|
if temp >= TEMP_CEILING_C:
|
||||||
instability.append(f"temperature ceiling hit ({temp}°C)")
|
instability.append(f"temperature ceiling hit ({temp}°C)")
|
||||||
|
|
||||||
tok_s = round(max((s["tokens_per_sec"] for s in ok_samples), default=0.0), 2)
|
tok_s = round(max((s.get(metric, 0.0) for s in ok_samples), default=0.0), 2)
|
||||||
row = {
|
row = {
|
||||||
"knob": knob,
|
"knob": knob,
|
||||||
"value": value,
|
"value": value,
|
||||||
"profile": profile,
|
"profile": profile,
|
||||||
|
"workload": workload,
|
||||||
|
"metric": metric,
|
||||||
"model": model,
|
"model": model,
|
||||||
"tokens_per_sec": tok_s,
|
"tokens_per_sec": tok_s,
|
||||||
"temp_c": temp,
|
"temp_c": temp,
|
||||||
@@ -340,6 +512,8 @@ async def sweep(knob: str = "mem_offset_mhz",
|
|||||||
"success": True,
|
"success": True,
|
||||||
"knob": knob,
|
"knob": knob,
|
||||||
"profile": profile,
|
"profile": profile,
|
||||||
|
"workload": workload,
|
||||||
|
"metric": metric,
|
||||||
"model": model,
|
"model": model,
|
||||||
"range": {"start": start, "stop": stop, "step": step, "values": values},
|
"range": {"start": start, "stop": stop, "step": step, "values": values},
|
||||||
"effectiveness": effectiveness,
|
"effectiveness": effectiveness,
|
||||||
@@ -370,6 +544,7 @@ async def sweep(knob: str = "mem_offset_mhz",
|
|||||||
# Always hand the card back exactly as we found it.
|
# Always hand the card back exactly as we found it.
|
||||||
state.running = False
|
state.running = False
|
||||||
state.current = None
|
state.current = None
|
||||||
|
vram_arbitrator.arbitrator.resume_oc(profile)
|
||||||
try:
|
try:
|
||||||
overclock_manager.apply_profile(profile)
|
overclock_manager.apply_profile(profile)
|
||||||
logger.info(f"autotune restored profile '{profile}'")
|
logger.info(f"autotune restored profile '{profile}'")
|
||||||
|
|||||||
@@ -67,6 +67,46 @@ DEFAULT_PROFILES: Dict[str, Dict[str, Any]] = {
|
|||||||
},
|
},
|
||||||
}
|
}
|
||||||
|
|
||||||
|
_OFFSETS_SUPPORTED: Optional[bool] = None
|
||||||
|
|
||||||
|
|
||||||
|
def offsets_supported(recheck: bool = False) -> bool:
|
||||||
|
"""Whether nvidia-settings clock offsets actually take effect on this driver.
|
||||||
|
|
||||||
|
Driver 595.84 accepts GPUGraphicsClockOffset/GPUMemoryTransferRateOffset and silently
|
||||||
|
discards them: assigning 0 returns success and the attribute still reads back its old
|
||||||
|
value. Profiles carrying core_offset_mhz/mem_offset_mhz were therefore configuring
|
||||||
|
nothing. Probed once and cached.
|
||||||
|
"""
|
||||||
|
global _OFFSETS_SUPPORTED
|
||||||
|
if _OFFSETS_SUPPORTED is not None and not recheck:
|
||||||
|
return _OFFSETS_SUPPORTED
|
||||||
|
|
||||||
|
def _read() -> Optional[int]:
|
||||||
|
q = _nvidia_settings("-q", "[gpu:0]/GPUGraphicsClockOffset[3]")
|
||||||
|
for line in (q.get("out") or "").splitlines():
|
||||||
|
if "Attribute" in line and "):" in line:
|
||||||
|
try:
|
||||||
|
return int(line.split("):")[-1].split(".")[0].strip())
|
||||||
|
except Exception:
|
||||||
|
pass
|
||||||
|
return None
|
||||||
|
|
||||||
|
before = _read()
|
||||||
|
if before is None:
|
||||||
|
_OFFSETS_SUPPORTED = False
|
||||||
|
return False
|
||||||
|
probe = before + 25
|
||||||
|
_nvidia_settings("-a", f"[gpu:0]/GPUGraphicsClockOffset[3]={probe}")
|
||||||
|
after = _read()
|
||||||
|
_nvidia_settings("-a", f"[gpu:0]/GPUGraphicsClockOffset[3]={before}")
|
||||||
|
_OFFSETS_SUPPORTED = (after is not None and after != before)
|
||||||
|
if not _OFFSETS_SUPPORTED:
|
||||||
|
logger.warning("Clock offsets are not honoured by this driver "
|
||||||
|
f"(set {probe}, read back {after}); profile offset fields are inert.")
|
||||||
|
return _OFFSETS_SUPPORTED
|
||||||
|
|
||||||
|
|
||||||
ACTIVE_PROFILE = "balanced"
|
ACTIVE_PROFILE = "balanced"
|
||||||
_LAST_RESULT: Dict[str, Any] = {}
|
_LAST_RESULT: Dict[str, Any] = {}
|
||||||
FAN_MANUAL = False
|
FAN_MANUAL = False
|
||||||
@@ -231,7 +271,11 @@ def apply_profile(name: str, overrides: Optional[Dict[str, Any]] = None) -> Dict
|
|||||||
"power_limit": _apply_power_limit(int(cfg.get("power_limit_w", 370))),
|
"power_limit": _apply_power_limit(int(cfg.get("power_limit_w", 370))),
|
||||||
"clock_lock": _apply_clock_lock(int(cfg.get("lock_core_min", 0)), int(cfg.get("lock_core_max", 0))),
|
"clock_lock": _apply_clock_lock(int(cfg.get("lock_core_min", 0)), int(cfg.get("lock_core_max", 0))),
|
||||||
"mem_lock": _apply_mem_lock(int(cfg.get("lock_mem_mhz", 0))),
|
"mem_lock": _apply_mem_lock(int(cfg.get("lock_mem_mhz", 0))),
|
||||||
"offsets": _apply_offsets(int(cfg.get("core_offset_mhz", 0)), int(cfg.get("mem_offset_mhz", 0))),
|
"offsets": (_apply_offsets(int(cfg.get("core_offset_mhz", 0)),
|
||||||
|
int(cfg.get("mem_offset_mhz", 0)))
|
||||||
|
if offsets_supported() else
|
||||||
|
{"applied": False, "supported": False,
|
||||||
|
"detail": "skipped: this driver accepts clock offsets and ignores them"}),
|
||||||
"fan": apply_fan_control(fan_mode, fan_speed),
|
"fan": apply_fan_control(fan_mode, fan_speed),
|
||||||
}
|
}
|
||||||
result["gpu"] = get_gpu_state()
|
result["gpu"] = get_gpu_state()
|
||||||
@@ -382,6 +426,9 @@ def get_status() -> Dict[str, Any]:
|
|||||||
"""Full overclock status for the dashboard."""
|
"""Full overclock status for the dashboard."""
|
||||||
return {
|
return {
|
||||||
"active_profile": ACTIVE_PROFILE,
|
"active_profile": ACTIVE_PROFILE,
|
||||||
|
"offsets_supported": offsets_supported(),
|
||||||
|
"effective_levers": (["power_limit", "clock_lock", "mem_lock", "fan"]
|
||||||
|
+ (["offsets"] if offsets_supported() else [])),
|
||||||
"profiles": load_profiles(),
|
"profiles": load_profiles(),
|
||||||
"gpu": get_gpu_state(),
|
"gpu": get_gpu_state(),
|
||||||
"fan": get_fan_status(),
|
"fan": get_fan_status(),
|
||||||
|
|||||||
@@ -1,29 +1,32 @@
|
|||||||
{
|
{
|
||||||
"ollama": {
|
"ollama": {
|
||||||
"label": "Ollama \u2014 LLM decode (memory-bandwidth bound)",
|
"label": "Ollama — LLM decode (memory-bandwidth bound; measured insensitive to power and clocks)",
|
||||||
"power_limit_w": 370,
|
"measured": "73.0-73.5 tok/s flat from 222W to 370W (qwen3.8long, 2026-08-28). Actual draw never exceeded 224W at any limit. Memory clock lock made no difference (72.6 locked vs 72.7 unlocked).",
|
||||||
"core_offset_mhz": 35,
|
"power_limit_w": 320,
|
||||||
"mem_offset_mhz": 200,
|
"core_offset_mhz": 0,
|
||||||
|
"mem_offset_mhz": 0,
|
||||||
"lock_core_min": 0,
|
"lock_core_min": 0,
|
||||||
"lock_core_max": 0,
|
"lock_core_max": 0,
|
||||||
"lock_mem_mhz": 0,
|
"lock_mem_mhz": 0,
|
||||||
"fan_mode": "manual",
|
"fan_mode": "auto",
|
||||||
"fan_speed_pct": 100
|
"fan_speed_pct": 0
|
||||||
},
|
},
|
||||||
"comfy": {
|
"comfy": {
|
||||||
"label": "ComfyUI \u2014 diffusion (core-compute bound)",
|
"label": "ComfyUI — diffusion (compute bound; genuinely power-scaling)",
|
||||||
|
"measured": "SDXL 1024/20-step: 5.48 it/s @222W, 6.22 @259W, 6.50 @296W, 6.52 @320W, 6.63 @333W, 6.71 @370W (2026-08-28). Worth +2.8% over the 320W stock default. Core clock lock made no difference across 2400-3105 MHz.",
|
||||||
"power_limit_w": 370,
|
"power_limit_w": 370,
|
||||||
"core_offset_mhz": 100,
|
"core_offset_mhz": 0,
|
||||||
"mem_offset_mhz": 150,
|
"mem_offset_mhz": 0,
|
||||||
"lock_core_min": 2900,
|
"lock_core_min": 0,
|
||||||
"lock_core_max": 3105,
|
"lock_core_max": 0,
|
||||||
"lock_mem_mhz": 0,
|
"lock_mem_mhz": 0,
|
||||||
"fan_mode": "manual",
|
"fan_mode": "auto",
|
||||||
"fan_speed_pct": 100
|
"fan_speed_pct": 0
|
||||||
},
|
},
|
||||||
"balanced": {
|
"balanced": {
|
||||||
"label": "Balanced \u2014 stock boost, power unlocked",
|
"label": "Balanced — stock power and boost, automatic fans",
|
||||||
"power_limit_w": 370,
|
"measured": "Card's own design point. 48k telemetry samples show 67.8C average under load at 39.5% auto fan, 81C all-time max, zero thermal throttle events.",
|
||||||
|
"power_limit_w": 320,
|
||||||
"core_offset_mhz": 0,
|
"core_offset_mhz": 0,
|
||||||
"mem_offset_mhz": 0,
|
"mem_offset_mhz": 0,
|
||||||
"lock_core_min": 0,
|
"lock_core_min": 0,
|
||||||
|
|||||||
45
server.py
45
server.py
@@ -3,6 +3,7 @@ import asyncio
|
|||||||
import contextlib
|
import contextlib
|
||||||
import json
|
import json
|
||||||
import logging
|
import logging
|
||||||
|
import signal
|
||||||
import time
|
import time
|
||||||
from contextlib import asynccontextmanager
|
from contextlib import asynccontextmanager
|
||||||
from typing import Dict, Any, Optional, List, Set
|
from typing import Dict, Any, Optional, List, Set
|
||||||
@@ -64,6 +65,13 @@ class TelemetryBroker:
|
|||||||
self.running = True
|
self.running = True
|
||||||
self.task = asyncio.create_task(self._loop())
|
self.task = asyncio.create_task(self._loop())
|
||||||
|
|
||||||
|
def begin_shutdown(self) -> None:
|
||||||
|
"""Release every SSE subscriber. Safe to call from a signal handler."""
|
||||||
|
self.closing = True
|
||||||
|
for q in list(self.subscribers):
|
||||||
|
with contextlib.suppress(asyncio.QueueFull):
|
||||||
|
q.put_nowait(None)
|
||||||
|
|
||||||
async def stop(self) -> None:
|
async def stop(self) -> None:
|
||||||
self.running = False
|
self.running = False
|
||||||
self.closing = True
|
self.closing = True
|
||||||
@@ -150,11 +158,36 @@ class TelemetryBroker:
|
|||||||
broker = TelemetryBroker()
|
broker = TelemetryBroker()
|
||||||
|
|
||||||
|
|
||||||
|
def _install_shutdown_hook() -> None:
|
||||||
|
"""Close SSE streams the moment a shutdown signal arrives.
|
||||||
|
|
||||||
|
uvicorn runs the lifespan shutdown only after it has finished waiting on open
|
||||||
|
connections, so releasing subscribers from there is too late: the streams keep the
|
||||||
|
server busy until the graceful timeout expires and every one of them is force
|
||||||
|
cancelled, which logs a CancelledError traceback apiece. Chaining onto the existing
|
||||||
|
signal handler lets us drain them first and leaves uvicorn's own shutdown intact.
|
||||||
|
"""
|
||||||
|
loop = asyncio.get_running_loop()
|
||||||
|
for sig in (signal.SIGTERM, signal.SIGINT):
|
||||||
|
previous = signal.getsignal(sig)
|
||||||
|
|
||||||
|
def handler(signum, frame, _prev=previous):
|
||||||
|
broker.begin_shutdown()
|
||||||
|
if callable(_prev):
|
||||||
|
_prev(signum, frame)
|
||||||
|
|
||||||
|
try:
|
||||||
|
signal.signal(sig, handler)
|
||||||
|
except (ValueError, OSError):
|
||||||
|
pass # not on the main thread; the lifespan path still cleans up
|
||||||
|
|
||||||
|
|
||||||
@asynccontextmanager
|
@asynccontextmanager
|
||||||
async def lifespan(app: FastAPI):
|
async def lifespan(app: FastAPI):
|
||||||
telemetry_store.start()
|
telemetry_store.start()
|
||||||
await broker.start()
|
await broker.start()
|
||||||
await vram_arbitrator.arbitrator.start()
|
await vram_arbitrator.arbitrator.start()
|
||||||
|
_install_shutdown_hook()
|
||||||
yield
|
yield
|
||||||
await vram_arbitrator.arbitrator.stop()
|
await vram_arbitrator.arbitrator.stop()
|
||||||
await broker.stop()
|
await broker.stop()
|
||||||
@@ -218,13 +251,16 @@ class GovernorRequest(BaseModel):
|
|||||||
reset: bool = Field(False, description="Clear any active derate and reapply the full profile")
|
reset: bool = Field(False, description="Clear any active derate and reapply the full profile")
|
||||||
|
|
||||||
class SweepRequest(BaseModel):
|
class SweepRequest(BaseModel):
|
||||||
knob: str = Field("mem_offset_mhz", description="mem_offset_mhz | core_offset_mhz")
|
knob: str = Field("mem_offset_mhz", description="mem_offset_mhz | core_offset_mhz | lock_mem_mhz | lock_core_max")
|
||||||
profile: str = Field("ollama", description="Profile to tune")
|
profile: str = Field("ollama", description="Profile to tune")
|
||||||
|
workload: str = Field("auto", description="ollama (decode tok/s) | comfy (diffusion it/s) | auto")
|
||||||
model: Optional[str] = Field(None, description="Model to benchmark with; defaults to the loaded one")
|
model: Optional[str] = Field(None, description="Model to benchmark with; defaults to the loaded one")
|
||||||
start: Optional[int] = Field(None, description="First offset value")
|
start: Optional[int] = Field(None, description="First offset value")
|
||||||
stop: Optional[int] = Field(None, description="Last offset value")
|
stop: Optional[int] = Field(None, description="Last offset value")
|
||||||
step: Optional[int] = Field(None, description="Offset increment")
|
step: Optional[int] = Field(None, description="Offset increment")
|
||||||
repeats: int = Field(1, description="Benchmark runs per step")
|
repeats: int = Field(1, description="Benchmark runs per step")
|
||||||
|
max_steps: int = Field(6, description="Cap on swept values for discrete clock knobs")
|
||||||
|
include_unlocked: bool = Field(True, description="Include an unlocked (0) control step")
|
||||||
apply_best: bool = Field(False, description="Write the winning value into the profile")
|
apply_best: bool = Field(False, description="Write the winning value into the profile")
|
||||||
|
|
||||||
class RequestVramRequest(BaseModel):
|
class RequestVramRequest(BaseModel):
|
||||||
@@ -263,7 +299,7 @@ async def sse_telemetry_stream(request: Request):
|
|||||||
if await request.is_disconnected():
|
if await request.is_disconnected():
|
||||||
break
|
break
|
||||||
try:
|
try:
|
||||||
snap = await asyncio.wait_for(q.get(), timeout=15.0)
|
snap = await asyncio.wait_for(q.get(), timeout=5.0)
|
||||||
if snap is None: # shutdown sentinel
|
if snap is None: # shutdown sentinel
|
||||||
break
|
break
|
||||||
yield f"data: {json.dumps(snap)}\n\n"
|
yield f"data: {json.dumps(snap)}\n\n"
|
||||||
@@ -483,9 +519,10 @@ async def api_autotune_status():
|
|||||||
async def api_autotune_sweep(req: SweepRequest):
|
async def api_autotune_sweep(req: SweepRequest):
|
||||||
"""Walk a clock offset upward, measuring tok/s and watching for instability at each step."""
|
"""Walk a clock offset upward, measuring tok/s and watching for instability at each step."""
|
||||||
res = await autotune.sweep(
|
res = await autotune.sweep(
|
||||||
knob=req.knob, profile=req.profile, model=req.model,
|
knob=req.knob, profile=req.profile, workload=req.workload, model=req.model,
|
||||||
start=req.start, stop=req.stop, step=req.step,
|
start=req.start, stop=req.stop, step=req.step,
|
||||||
repeats=req.repeats, apply_best=req.apply_best,
|
repeats=req.repeats, max_steps=req.max_steps,
|
||||||
|
include_unlocked=req.include_unlocked, apply_best=req.apply_best,
|
||||||
)
|
)
|
||||||
if not res.get("success"):
|
if not res.get("success"):
|
||||||
raise HTTPException(status_code=400, detail=res.get("error"))
|
raise HTTPException(status_code=400, detail=res.get("error"))
|
||||||
|
|||||||
@@ -666,6 +666,10 @@ class AutoArbitrator:
|
|||||||
self.comfy_idle_since: Optional[float] = None
|
self.comfy_idle_since: Optional[float] = None
|
||||||
self.oc_profile = None
|
self.oc_profile = None
|
||||||
self.pending_purge = False
|
self.pending_purge = False
|
||||||
|
# While a tuning sweep is running, the arbitrator must not fight it: a ComfyUI
|
||||||
|
# benchmark would otherwise trip trigger_comfy_priority, which reapplies the whole
|
||||||
|
# 'comfy' profile and silently overwrites the clock the sweep is measuring.
|
||||||
|
self.oc_suspended = False
|
||||||
self.stats = {"yields": 0, "purges": 0, "yield_timeouts": 0, "deferred_purges": 0}
|
self.stats = {"yields": 0, "purges": 0, "yield_timeouts": 0, "deferred_purges": 0}
|
||||||
|
|
||||||
async def start(self):
|
async def start(self):
|
||||||
@@ -841,9 +845,19 @@ class AutoArbitrator:
|
|||||||
pass
|
pass
|
||||||
await asyncio.sleep(interval)
|
await asyncio.sleep(interval)
|
||||||
|
|
||||||
|
def suspend_oc(self, reason: str = "tuning sweep") -> None:
|
||||||
|
self.oc_suspended = True
|
||||||
|
logger.info(f"Overclock auto-switching suspended ({reason})")
|
||||||
|
|
||||||
|
def resume_oc(self, profile: Optional[str] = None) -> None:
|
||||||
|
self.oc_suspended = False
|
||||||
|
# Forget the cached profile so the next transition actually reapplies.
|
||||||
|
self.oc_profile = profile
|
||||||
|
logger.info("Overclock auto-switching resumed")
|
||||||
|
|
||||||
def _apply_oc_profile(self, profile: str):
|
def _apply_oc_profile(self, profile: str):
|
||||||
"""Apply an overclock profile in a background thread; only fire on transition."""
|
"""Apply an overclock profile in a background thread; only fire on transition."""
|
||||||
if self.oc_profile == profile:
|
if self.oc_suspended or self.oc_profile == profile:
|
||||||
return
|
return
|
||||||
self.oc_profile = profile
|
self.oc_profile = profile
|
||||||
try:
|
try:
|
||||||
|
|||||||
Reference in New Issue
Block a user