Tune profiles from measurement; add a diffusion benchmark to close the loop

The profiles were hand-written and had never been checked against the hardware. Adding
a ComfyUI benchmark alongside the existing decode one made the compute side measurable
for the first time, and most of what the profiles configured turned out to do nothing.

Measured on this card (RTX 4080 SUPER, driver 595.84):

- LLM decode is not power-bound: 73.0-73.5 tok/s flat from 222W to 370W, with the card
  never drawing more than 224W at any limit. The ollama profile's 370W did nothing.
- Diffusion is power-bound: 5.48 it/s @222W rising to 6.71 @370W, so comfy's 370W is
  worth a real +2.8% over the 320W stock default.
- Clock locks did nothing for either workload: 72.6 tok/s locked at 11251MHz vs 72.7
  unlocked; 6.77 it/s locked at 3105MHz vs 6.73 unlocked, and 6.78 at 2400MHz.
- Memory bandwidth is still the decode bottleneck (5001MHz halves throughput to 35.9
  tok/s), confirming the profile's premise -- the card just gets there unaided.
- Fans: 48,435 samples show 81C all-time max and zero thermal throttle events, while
  the ollama profile held 49.6C average by running fans at 87%. All profiles now use
  automatic fans and let the thermal governor escalate on demand.

Code changes supporting that:
- _diffusion_benchmark() queues a fixed SDXL graph via ComfyUI's API. The seed must
  vary per run: ComfyUI caches by node inputs, so a fixed seed returned in ~1ms without
  executing. Implausibly fast results are now rejected as cache hits rather than
  recorded as record scores.
- The arbitrator's automatic profile switching is suspended during a sweep. A diffusion
  benchmark trips trigger_comfy_priority, which reapplies the whole profile and would
  silently overwrite the clock being measured.
- _supported_clocks() queries the mem,gr pair; asking for a single field returned one
  column and reading index 1 yielded an empty list rather than an error. Graphics clocks
  are subsampled (the card enumerates 194 of them) and lock sweeps include an explicit
  unlocked control step.
- offsets_supported() probes once and apply_profile skips inert offset levers with an
  explanation instead of pretending they applied.
- Profiles carry a 'measured' field recording the evidence behind each setting.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
drjones
2026-08-28 11:18:35 -07:00
parent a30444ef8e
commit c689ec8711
6 changed files with 373 additions and 56 deletions

View File

@@ -33,17 +33,55 @@
* **Budgeted, Ranked Warming**: This box has 64 GB of RAM and >270 GB of model files; reading everything simply evicts whatever was warmed first. Files are ranked by recency/frequency (from the persisted event log) and warmed until a byte budget is spent, skipping anything already resident. `GET /api/warm-plan` previews the decision without executing it.
* **Memory Telemetry**: Real-time breakdown of Total Host RAM, Applications Memory, Active Model Page Cache, Free Memory, and measured Cache Residency Ratio.
### 🎛️ Dynamic Overclocking & Thermal Management
* **Workload-Aware Overclock Profiles**:
* **`ollama` Profile (Memory-Bandwidth Bound)**: Max 370W power limit, +150 MHz Core Offset, +825 MHz Memory Offset, and 100% fan speed for maximum prompt eval / generation bandwidth.
* **`comfy` Profile (Compute Bound)**: Max 370W power limit, +100 MHz Core Offset, +500 MHz Memory Offset, Core Clock locked to 2900–3105 MHz, and 75% fan speed for maximum diffusion compute.
* **`balanced` Profile (Stock/General Purpose)**: Unlocked 370W power limit with stock dynamic boost curves and automatic fan control.
* **Hardware Actuation Hierarchy**:
* Level 1: Power Limit Control (`nvidia-smi -pl 370`).
* Level 2: Core & Memory Clock Locking (`nvidia-smi -lgc` / `-lmc`).
* Level 3: Clock Offsets via headless X display (`:8`) with Coolbits support (`nvidia-settings`).
* **Hardware Fan Control**: Switch between `auto` and `manual` PWM control (30%–100%) with synchronized dual-fan actuation (`[fan:0]` and `[fan:1]`).
* **Automated Lockstep Profile Switching**: AutoArbitrator automatically switches hardware profiles in lockstep with the active workload (`comfy` on generation start, `ollama` on completion).
### 🎛️ Measured Overclock Profiles & Thermal Management
Every profile setting in this repo is now backed by a measurement from `autotune.py` on
this specific card and driver. Several long-standing settings turned out to do nothing.
**What this driver actually honours** (NVIDIA 595.84, RTX 4080 SUPER):
| Lever | Mechanism | Works? |
| :--- | :--- | :--- |
| Power limit | `nvidia-smi -pl` | ✅ Yes — and it is the only lever that changes anything measurable |
| Core / memory clock lock | `nvidia-smi -lgc` / `-lmc` | ✅ Applies correctly, but made no measurable difference to either workload |
| Core / memory clock offsets | `nvidia-settings -a ...Offset` | ❌ **Silently ignored.** The driver reports `assigned value 0` and the attribute still reads back `250`. Detected automatically by `offsets_supported()`; `apply_profile` now skips them and says so rather than pretending. |
| Fan control | `nvidia-settings GPUTargetFanSpeed` | ✅ Yes |
**Measured results** (`POST /api/autotune/sweep`):
*LLM decode is not power-bound.* Throughput is flat across the card's entire power range —
the GPU never drew more than 224 W no matter what the limit allowed:
| Power limit | 222 W | 259 W | 296 W | 320 W | 333 W | 370 W |
| :--- | :--- | :--- | :--- | :--- | :--- | :--- |
| tok/s (`qwen3.8long`) | 73.17 | 73.13 | 73.43 | 73.51 | 73.10 | 73.04 |
*Diffusion is power-bound.* Here the watts genuinely buy throughput:
| Power limit | 222 W | 259 W | 296 W | 320 W | 333 W | 370 W |
| :--- | :--- | :--- | :--- | :--- | :--- | :--- |
| it/s (SDXL 1024, 20 steps) | 5.48 | 6.22 | 6.50 | 6.52 | 6.63 | **6.71** |
*Clock locks changed nothing for either workload.* Memory clock: 72.6 tok/s locked at
11251 MHz vs 72.7 unlocked. Core clock: 6.73 it/s unlocked vs 6.77 locked at 3105 MHz —
and 6.78 at 2400 MHz, so diffusion here is not core-clock-bound at all.
*Memory bandwidth is the decode bottleneck*, confirming the profile's original premise —
dropping the memory clock to 5001 MHz halves throughput (35.9 tok/s vs 72.6). The card
simply reaches its top memory clock on its own; pinning it there adds nothing.
**Resulting profiles**:
* **`ollama`** — 320 W (stock), no locks, automatic fans. Decode draws ~224 W and is
bandwidth-bound, so the previous 370 W limit and 100% fan pinning bought nothing.
* **`comfy`** — 370 W, no locks, automatic fans. The extra power is worth a measured
**+2.8%** over the 320 W stock default.
* **`balanced`** — stock power and boost, automatic fans.
**On fans**: all three profiles previously pinned the fans to manual 100%. Across 48,435
telemetry samples this card has never exceeded **81 °C** and has logged **zero** thermal
throttle events; the `ollama` profile was holding 49.6 °C average by running the fans at
87%. Fans are now automatic in every profile, with the thermal governor escalating them
only if the card actually needs it.
### 🌡️ Thermal Governor (closed-loop de-escalation)
* Every overclock lever here is sticky: a profile locks clocks and pins the fans to a manual PWM, and nothing used to undo that. The governor watches the telemetry the sampler already collects (so it costs no extra NVML calls) and walks the overclock back through a four-step derate ladder when the card runs hot or reports a hardware throttle.
@@ -51,9 +89,12 @@
* **Guaranteed restore**: stock clocks, default power limit and automatic fans are restored by the server's shutdown hook *and* by a systemd `ExecStopPost=`, so a `SIGKILL` cannot leave the card with locked clocks and fans pinned at 100%.
### 🔬 Overclock Autotune (`autotune.py`)
* Walks a clock offset upward, running a fixed decode benchmark at each step, and reports the **fastest stable** value with its measured gain over baseline.
* Sweeps a knob (`power_limit_w`, `lock_mem_mhz`, `lock_core_max`, clock offsets) and reports the **fastest stable** value.
* **Both workloads are measurable.** `workload=ollama` benchmarks decode throughput in tok/s; `workload=comfy` queues a fixed SDXL 1024/20-step graph through ComfyUI's API and measures it/s. Without the second one there was no way to tell whether the compute-oriented `comfy` profile was doing anything at all — and it was not.
* **Refuses to sweep a knob the driver ignores.** A preflight applies a probe value and confirms the hardware moved; the probe is chosen as the candidate furthest from the current reading, since probing with the maximum proves nothing when the card already sits there. This is what caught the silently-discarded clock offsets.
* **Instability detection**: kernel `Xid`/`NVRM` messages via `journalctl -k`, benchmark failure, degenerate output, and a temperature ceiling. The sweep stops climbing the moment a step looks unstable.
* **Safety**: refuses to start while ComfyUI is executing, and restores the original profile in a `finally` block — including on exception or cancellation.
* **Honest gain reporting**: gain against the profile's *current* setting is reported separately from the spread across values tried. Conflating them turns a flat result into a headline "+102%".
* **Safety**: refuses to start while ComfyUI is executing, suspends the arbitrator's automatic profile switching for the duration (otherwise a diffusion benchmark trips `trigger_comfy_priority`, which reapplies the whole profile and overwrites the clock being measured), and restores the original profile in a `finally` block — including on exception or cancellation.
### 🗄️ Persistent Telemetry Store (`telemetry_store.py`)
* Swap history used to be an in-memory `deque(maxlen=50)` that evaporated on every restart. Telemetry and events now persist to SQLite (WAL, single writer thread, batched 1 Hz inserts, automatic retention pruning) at roughly **0.4 MB per hour**.
@@ -161,7 +202,7 @@ The HyperSwap server runs on port `9090` by default. Interactive OpenAPI/Swagger
| `/api/governor` | `GET` / `POST` | Current derate level and why; enable/disable, or clear an active derate. |
| `/api/overclock/restore` | `POST` | Drop all clock locks and offsets, restore default power limit and automatic fans. |
| `/api/autotune` | `GET` | Sweep progress, last result, and every recorded autotune step. |
| `/api/autotune/sweep` | `POST` | Walk a clock offset upward, measuring tok/s and watching for instability at each step. |
| `/api/autotune/sweep` | `POST` | Sweep a knob against a real workload (`workload`: `ollama` decode tok/s, `comfy` SDXL it/s), verifying the knob moves the hardware first. |
| `/api/autotune/cancel` | `POST` | Stop the current sweep after the step in flight; the profile is restored either way. |
---

View File

@@ -15,6 +15,7 @@ Safety properties:
"""
import asyncio
import logging
import random
import subprocess
import time
from typing import Any, Dict, List, Optional
@@ -46,28 +47,74 @@ KNOBS = {
"values": None}, # filled from the card's supported clock list
"lock_core_max": {"kind": "discrete", "verify": "clock_sm",
"values": None},
# The one lever this driver definitely honours. Worth knowing whether the extra
# watts actually buy throughput, or just heat and fan noise.
"power_limit_w": {"kind": "discrete", "verify": "power_limit",
"values": None, "no_unlocked": True},
}
def _supported_clocks(which: str = "mem") -> List[int]:
"""Discrete clock values the card will actually accept for -lmc / -lgc."""
"""Discrete clock values the card will actually accept for -lmc / -lgc.
Always queried as the mem,gr pair: asking for a single field returns one column, and
reading index 1 from it silently yields an empty list rather than an error.
"""
try:
proc = subprocess.run(
["nvidia-smi", f"--query-supported-clocks={'mem' if which == 'mem' else 'gr'}",
"--format=csv,noheader,nounits"],
capture_output=True, text=True, timeout=10)
col = 0 if which == "mem" else 1
vals = set()
["nvidia-smi", "--query-supported-clocks=mem,gr", "--format=csv,noheader,nounits"],
capture_output=True, text=True, timeout=15)
rows = []
for line in proc.stdout.splitlines():
parts = [p.strip() for p in line.split(",")]
if len(parts) > col and parts[col].isdigit():
vals.add(int(parts[col]))
return sorted(vals)
if len(parts) >= 2 and parts[0].isdigit() and parts[1].isdigit():
rows.append((int(parts[0]), int(parts[1])))
if not rows:
return []
if which == "mem":
return sorted({m for m, _ in rows})
# Graphics clocks are enumerated per memory clock; take the list for the highest
# memory clock, which is the one any real workload runs at.
top_mem = max(m for m, _ in rows)
return sorted({g for m, g in rows if m == top_mem})
except Exception as e:
logger.debug(f"supported clock query failed: {e}")
return []
def _supported_power_limits(steps: int = 5) -> List[int]:
"""Power limits between the card's minimum and maximum, in even increments."""
try:
proc = subprocess.run(
["nvidia-smi", "--query-gpu=power.min_limit,power.max_limit,power.default_limit",
"--format=csv,noheader,nounits"],
capture_output=True, text=True, timeout=10)
parts = [p.strip() for p in proc.stdout.strip().split(",")]
lo, hi, default = (int(float(parts[0])), int(float(parts[1])), int(float(parts[2])))
except Exception as e:
logger.debug(f"power limit query failed: {e}")
return []
# Start at 60% of max -- below that the card is not doing useful work for these
# workloads -- and always include the stock default as a reference point.
lo = max(lo, int(hi * 0.6))
span = hi - lo
vals = {lo + round(i * span / (steps - 1)) for i in range(steps)}
vals.add(default)
return sorted(v for v in vals if lo <= v <= hi)
def _subsample(values: List[int], max_steps: int) -> List[int]:
"""Evenly spaced subset, always keeping the endpoints.
The card enumerates ~194 graphics clocks in 15 MHz increments; benchmarking every one
would take hours and tell us nothing that a handful of well-spread points does not.
"""
if len(values) <= max_steps:
return values
idx = [round(i * (len(values) - 1) / (max_steps - 1)) for i in range(max_steps)]
return sorted({values[i] for i in idx})
def _read_hw(field: str) -> Optional[float]:
"""Read back the hardware state a knob is supposed to move."""
gpu = vram_arbitrator.get_gpu_hardware_stats()
@@ -75,6 +122,8 @@ def _read_hw(field: str) -> Optional[float]:
return gpu.get("clock_mem_mhz")
if field == "clock_sm":
return gpu.get("clock_graphics_mhz")
if field == "power_limit":
return gpu.get("power_limit_w")
if field in ("mem_offset", "core_offset"):
r = overclock_manager._nvidia_settings(
"-q", f"[gpu:0]/{'GPUMemoryTransferRateOffset' if field == 'mem_offset' else 'GPUGraphicsClockOffset'}[3]")
@@ -174,6 +223,97 @@ async def _decode_benchmark(model: str) -> Dict[str, Any]:
}
# A fixed SDXL txt2img graph. Deterministic seed/steps/resolution so every step of a
# sweep does identical work and the only variable is the clock. PreviewImage rather than
# SaveImage keeps benchmark runs out of the user's output gallery.
COMFY_BENCH_CKPT = "sd_xl_base_1.0.safetensors"
COMFY_BENCH_STEPS = 20
COMFY_BENCH_SIZE = 1024
def _comfy_workflow(ckpt: str = COMFY_BENCH_CKPT, seed: Optional[int] = None) -> Dict[str, Any]:
# The seed must vary per run. ComfyUI caches by node inputs, so a fixed seed makes the
# second and later benchmarks return in ~1ms without executing anything at all. The
# cost of the graph is identical regardless of seed, so this costs no comparability.
seed = random.randint(1, 2**31) if seed is None else seed
return {
"1": {"class_type": "CheckpointLoaderSimple", "inputs": {"ckpt_name": ckpt}},
"2": {"class_type": "CLIPTextEncode",
"inputs": {"clip": ["1", 1],
"text": "a detailed photograph of a mountain range at sunrise"}},
"3": {"class_type": "CLIPTextEncode",
"inputs": {"clip": ["1", 1], "text": "blurry, low quality"}},
"4": {"class_type": "EmptyLatentImage",
"inputs": {"width": COMFY_BENCH_SIZE, "height": COMFY_BENCH_SIZE, "batch_size": 1}},
"5": {"class_type": "KSampler",
"inputs": {"model": ["1", 0], "positive": ["2", 0], "negative": ["3", 0],
"latent_image": ["4", 0], "seed": seed, "steps": COMFY_BENCH_STEPS,
"cfg": 7.0, "sampler_name": "euler", "scheduler": "normal",
"denoise": 1.0}},
"6": {"class_type": "VAEDecode", "inputs": {"samples": ["5", 0], "vae": ["1", 2]}},
"7": {"class_type": "PreviewImage", "inputs": {"images": ["6", 0]}},
}
async def _diffusion_benchmark(ckpt: str = COMFY_BENCH_CKPT,
timeout_s: float = 300.0) -> Dict[str, Any]:
"""Queue one fixed SDXL graph and time it. This is the compute-bound counterpart to
the decode benchmark, and the only way to tell whether the 'comfy' profile helps."""
client = vram_arbitrator._client(vram_arbitrator.COMFY_API_BASE, 30.0)
t0 = time.perf_counter()
try:
resp = await client.post("/prompt", json={"prompt": _comfy_workflow(ckpt),
"client_id": "hyperswap-autotune"})
if resp.status_code != 200:
return {"ok": False, "error": f"queue failed HTTP {resp.status_code}: {resp.text[:200]}"}
prompt_id = resp.json().get("prompt_id")
except Exception as e:
return {"ok": False, "error": f"queue failed: {e}"}
while (time.perf_counter() - t0) < timeout_s:
await asyncio.sleep(0.25)
try:
h = await client.get(f"/history/{prompt_id}")
if h.status_code != 200:
continue
entry = (h.json() or {}).get(prompt_id)
if not entry:
continue
status = entry.get("status", {})
if status.get("status_str") == "error" or not status.get("completed", True):
if status.get("status_str") == "error":
return {"ok": False, "error": "ComfyUI reported an execution error",
"wall_ms": round((time.perf_counter() - t0) * 1000, 2)}
if status.get("completed"):
wall = time.perf_counter() - t0
# ComfyUI stamps execution_start/success in the status messages; the delta
# between them excludes our polling overhead and the queue wait.
stamps = {}
for msg in status.get("messages", []):
if isinstance(msg, list) and len(msg) >= 2 and isinstance(msg[1], dict):
if "timestamp" in msg[1]:
stamps[msg[0]] = msg[1]["timestamp"]
exec_ms = None
if "execution_start" in stamps and "execution_success" in stamps:
exec_ms = round(stamps["execution_success"] - stamps["execution_start"], 2)
effective_ms = exec_ms or wall * 1000
# A graph that "finished" implausibly fast was served from ComfyUI's cache
# rather than executed; treat it as an invalid sample, not a record score.
cached = effective_ms < 250
return {
"ok": not cached,
"error": "result served from ComfyUI cache, not executed" if cached else None,
"wall_ms": round(wall * 1000, 2),
"exec_ms": exec_ms,
"steps": COMFY_BENCH_STEPS,
"it_per_sec": round(COMFY_BENCH_STEPS / (effective_ms / 1000), 3),
"degenerate": cached,
}
except Exception:
continue
return {"ok": False, "error": f"diffusion benchmark timed out after {timeout_s}s"}
class SweepState:
def __init__(self) -> None:
self.running = False
@@ -187,11 +327,14 @@ state = SweepState()
async def sweep(knob: str = "mem_offset_mhz",
profile: str = "ollama",
workload: str = "auto",
model: Optional[str] = None,
start: Optional[int] = None,
stop: Optional[int] = None,
step: Optional[int] = None,
repeats: int = 1,
max_steps: int = 6,
include_unlocked: bool = True,
apply_best: bool = False) -> Dict[str, Any]:
"""Sweep one clock offset and return the fastest stable value."""
if knob not in KNOBS:
@@ -203,7 +346,23 @@ async def sweep(knob: str = "mem_offset_mhz",
if comfy.get("executing") or comfy.get("queue_remaining"):
return {"success": False, "error": "ComfyUI is busy; refusing to change clocks mid-render"}
if not model:
# 'auto': tune the workload the profile is actually for.
if workload == "auto":
workload = "comfy" if profile == "comfy" else "ollama"
if workload not in ("ollama", "comfy"):
return {"success": False, "error": "workload must be 'ollama', 'comfy' or 'auto'"}
if workload == "comfy" and not comfy.get("online"):
return {"success": False, "error": "ComfyUI is not reachable; cannot run a diffusion sweep"}
if workload == "comfy":
model = model or COMFY_BENCH_CKPT
benchmark = lambda: _diffusion_benchmark(model)
metric = "it_per_sec"
else:
benchmark = lambda: _decode_benchmark(model)
metric = "tokens_per_sec"
if workload == "ollama" and not model:
ollama = await vram_arbitrator.get_ollama_live_state()
model = ollama.get("active_model_name")
if not model:
@@ -211,9 +370,13 @@ async def sweep(knob: str = "mem_offset_mhz",
if not installed:
return {"success": False, "error": "no Ollama model available to benchmark"}
model = installed[0].get("name")
benchmark = lambda: _decode_benchmark(model)
defaults = KNOBS[knob]
if defaults.get("kind") == "discrete":
if knob == "power_limit_w":
supported = _supported_power_limits()
else:
supported = defaults.get("values") or _supported_clocks(
"mem" if knob == "lock_mem_mhz" else "gr")
if not supported:
@@ -222,6 +385,11 @@ async def sweep(knob: str = "mem_offset_mhz",
if (start is None or v >= start) and (stop is None or v <= stop)]
if not values:
return {"success": False, "error": f"no supported values in range; card offers {supported}"}
values = _subsample(values, max_steps or 6)
# 0 means "no lock at all". That is the honest control for a profile whose whole
# premise is that locking the clock beats letting the card boost on its own.
if include_unlocked and not defaults.get("no_unlocked"):
values = [0] + values
start, stop, step = values[0], values[-1], None
else:
start = defaults["default_start"] if start is None else start
@@ -235,7 +403,8 @@ async def sweep(knob: str = "mem_offset_mhz",
baseline_value = int(baseline_cfg.get(knob, 0) or 0)
# Refuse to sweep a knob the driver is going to ignore.
effectiveness = _knob_effective(knob, profile, values, baseline_value)
effectiveness = _knob_effective(knob, profile, [v for v in values if v] or values,
baseline_value)
if not effectiveness["effective"]:
return {
"success": False,
@@ -249,8 +418,9 @@ async def sweep(knob: str = "mem_offset_mhz",
t_start = time.time()
try:
# Load the model once up front so the first step does not pay the load cost.
await _decode_benchmark(model)
# Warm-up: load weights once up front so the first step does not pay the load cost.
vram_arbitrator.arbitrator.suspend_oc("autotune sweep")
await benchmark()
for value in values:
if state.cancel:
@@ -261,7 +431,7 @@ async def sweep(knob: str = "mem_offset_mhz",
samples = []
for _ in range(max(repeats, 1)):
samples.append(await _decode_benchmark(model))
samples.append(await benchmark())
if state.cancel:
break
@@ -278,11 +448,13 @@ async def sweep(knob: str = "mem_offset_mhz",
if temp >= TEMP_CEILING_C:
instability.append(f"temperature ceiling hit ({temp}°C)")
tok_s = round(max((s["tokens_per_sec"] for s in ok_samples), default=0.0), 2)
tok_s = round(max((s.get(metric, 0.0) for s in ok_samples), default=0.0), 2)
row = {
"knob": knob,
"value": value,
"profile": profile,
"workload": workload,
"metric": metric,
"model": model,
"tokens_per_sec": tok_s,
"temp_c": temp,
@@ -340,6 +512,8 @@ async def sweep(knob: str = "mem_offset_mhz",
"success": True,
"knob": knob,
"profile": profile,
"workload": workload,
"metric": metric,
"model": model,
"range": {"start": start, "stop": stop, "step": step, "values": values},
"effectiveness": effectiveness,
@@ -370,6 +544,7 @@ async def sweep(knob: str = "mem_offset_mhz",
# Always hand the card back exactly as we found it.
state.running = False
state.current = None
vram_arbitrator.arbitrator.resume_oc(profile)
try:
overclock_manager.apply_profile(profile)
logger.info(f"autotune restored profile '{profile}'")

View File

@@ -67,6 +67,46 @@ DEFAULT_PROFILES: Dict[str, Dict[str, Any]] = {
},
}
_OFFSETS_SUPPORTED: Optional[bool] = None
def offsets_supported(recheck: bool = False) -> bool:
"""Whether nvidia-settings clock offsets actually take effect on this driver.
Driver 595.84 accepts GPUGraphicsClockOffset/GPUMemoryTransferRateOffset and silently
discards them: assigning 0 returns success and the attribute still reads back its old
value. Profiles carrying core_offset_mhz/mem_offset_mhz were therefore configuring
nothing. Probed once and cached.
"""
global _OFFSETS_SUPPORTED
if _OFFSETS_SUPPORTED is not None and not recheck:
return _OFFSETS_SUPPORTED
def _read() -> Optional[int]:
q = _nvidia_settings("-q", "[gpu:0]/GPUGraphicsClockOffset[3]")
for line in (q.get("out") or "").splitlines():
if "Attribute" in line and "):" in line:
try:
return int(line.split("):")[-1].split(".")[0].strip())
except Exception:
pass
return None
before = _read()
if before is None:
_OFFSETS_SUPPORTED = False
return False
probe = before + 25
_nvidia_settings("-a", f"[gpu:0]/GPUGraphicsClockOffset[3]={probe}")
after = _read()
_nvidia_settings("-a", f"[gpu:0]/GPUGraphicsClockOffset[3]={before}")
_OFFSETS_SUPPORTED = (after is not None and after != before)
if not _OFFSETS_SUPPORTED:
logger.warning("Clock offsets are not honoured by this driver "
f"(set {probe}, read back {after}); profile offset fields are inert.")
return _OFFSETS_SUPPORTED
ACTIVE_PROFILE = "balanced"
_LAST_RESULT: Dict[str, Any] = {}
FAN_MANUAL = False
@@ -231,7 +271,11 @@ def apply_profile(name: str, overrides: Optional[Dict[str, Any]] = None) -> Dict
"power_limit": _apply_power_limit(int(cfg.get("power_limit_w", 370))),
"clock_lock": _apply_clock_lock(int(cfg.get("lock_core_min", 0)), int(cfg.get("lock_core_max", 0))),
"mem_lock": _apply_mem_lock(int(cfg.get("lock_mem_mhz", 0))),
"offsets": _apply_offsets(int(cfg.get("core_offset_mhz", 0)), int(cfg.get("mem_offset_mhz", 0))),
"offsets": (_apply_offsets(int(cfg.get("core_offset_mhz", 0)),
int(cfg.get("mem_offset_mhz", 0)))
if offsets_supported() else
{"applied": False, "supported": False,
"detail": "skipped: this driver accepts clock offsets and ignores them"}),
"fan": apply_fan_control(fan_mode, fan_speed),
}
result["gpu"] = get_gpu_state()
@@ -382,6 +426,9 @@ def get_status() -> Dict[str, Any]:
"""Full overclock status for the dashboard."""
return {
"active_profile": ACTIVE_PROFILE,
"offsets_supported": offsets_supported(),
"effective_levers": (["power_limit", "clock_lock", "mem_lock", "fan"]
+ (["offsets"] if offsets_supported() else [])),
"profiles": load_profiles(),
"gpu": get_gpu_state(),
"fan": get_fan_status(),

View File

@@ -1,29 +1,32 @@
{
"ollama": {
"label": "Ollama \u2014 LLM decode (memory-bandwidth bound)",
"power_limit_w": 370,
"core_offset_mhz": 35,
"mem_offset_mhz": 200,
"label": "Ollama — LLM decode (memory-bandwidth bound; measured insensitive to power and clocks)",
"measured": "73.0-73.5 tok/s flat from 222W to 370W (qwen3.8long, 2026-08-28). Actual draw never exceeded 224W at any limit. Memory clock lock made no difference (72.6 locked vs 72.7 unlocked).",
"power_limit_w": 320,
"core_offset_mhz": 0,
"mem_offset_mhz": 0,
"lock_core_min": 0,
"lock_core_max": 0,
"lock_mem_mhz": 0,
"fan_mode": "manual",
"fan_speed_pct": 100
"fan_mode": "auto",
"fan_speed_pct": 0
},
"comfy": {
"label": "ComfyUI \u2014 diffusion (core-compute bound)",
"label": "ComfyUI — diffusion (compute bound; genuinely power-scaling)",
"measured": "SDXL 1024/20-step: 5.48 it/s @222W, 6.22 @259W, 6.50 @296W, 6.52 @320W, 6.63 @333W, 6.71 @370W (2026-08-28). Worth +2.8% over the 320W stock default. Core clock lock made no difference across 2400-3105 MHz.",
"power_limit_w": 370,
"core_offset_mhz": 100,
"mem_offset_mhz": 150,
"lock_core_min": 2900,
"lock_core_max": 3105,
"core_offset_mhz": 0,
"mem_offset_mhz": 0,
"lock_core_min": 0,
"lock_core_max": 0,
"lock_mem_mhz": 0,
"fan_mode": "manual",
"fan_speed_pct": 100
"fan_mode": "auto",
"fan_speed_pct": 0
},
"balanced": {
"label": "Balanced \u2014 stock boost, power unlocked",
"power_limit_w": 370,
"label": "Balanced — stock power and boost, automatic fans",
"measured": "Card's own design point. 48k telemetry samples show 67.8C average under load at 39.5% auto fan, 81C all-time max, zero thermal throttle events.",
"power_limit_w": 320,
"core_offset_mhz": 0,
"mem_offset_mhz": 0,
"lock_core_min": 0,

View File

@@ -3,6 +3,7 @@ import asyncio
import contextlib
import json
import logging
import signal
import time
from contextlib import asynccontextmanager
from typing import Dict, Any, Optional, List, Set
@@ -64,6 +65,13 @@ class TelemetryBroker:
self.running = True
self.task = asyncio.create_task(self._loop())
def begin_shutdown(self) -> None:
"""Release every SSE subscriber. Safe to call from a signal handler."""
self.closing = True
for q in list(self.subscribers):
with contextlib.suppress(asyncio.QueueFull):
q.put_nowait(None)
async def stop(self) -> None:
self.running = False
self.closing = True
@@ -150,11 +158,36 @@ class TelemetryBroker:
broker = TelemetryBroker()
def _install_shutdown_hook() -> None:
"""Close SSE streams the moment a shutdown signal arrives.
uvicorn runs the lifespan shutdown only after it has finished waiting on open
connections, so releasing subscribers from there is too late: the streams keep the
server busy until the graceful timeout expires and every one of them is force
cancelled, which logs a CancelledError traceback apiece. Chaining onto the existing
signal handler lets us drain them first and leaves uvicorn's own shutdown intact.
"""
loop = asyncio.get_running_loop()
for sig in (signal.SIGTERM, signal.SIGINT):
previous = signal.getsignal(sig)
def handler(signum, frame, _prev=previous):
broker.begin_shutdown()
if callable(_prev):
_prev(signum, frame)
try:
signal.signal(sig, handler)
except (ValueError, OSError):
pass # not on the main thread; the lifespan path still cleans up
@asynccontextmanager
async def lifespan(app: FastAPI):
telemetry_store.start()
await broker.start()
await vram_arbitrator.arbitrator.start()
_install_shutdown_hook()
yield
await vram_arbitrator.arbitrator.stop()
await broker.stop()
@@ -218,13 +251,16 @@ class GovernorRequest(BaseModel):
reset: bool = Field(False, description="Clear any active derate and reapply the full profile")
class SweepRequest(BaseModel):
knob: str = Field("mem_offset_mhz", description="mem_offset_mhz | core_offset_mhz")
knob: str = Field("mem_offset_mhz", description="mem_offset_mhz | core_offset_mhz | lock_mem_mhz | lock_core_max")
profile: str = Field("ollama", description="Profile to tune")
workload: str = Field("auto", description="ollama (decode tok/s) | comfy (diffusion it/s) | auto")
model: Optional[str] = Field(None, description="Model to benchmark with; defaults to the loaded one")
start: Optional[int] = Field(None, description="First offset value")
stop: Optional[int] = Field(None, description="Last offset value")
step: Optional[int] = Field(None, description="Offset increment")
repeats: int = Field(1, description="Benchmark runs per step")
max_steps: int = Field(6, description="Cap on swept values for discrete clock knobs")
include_unlocked: bool = Field(True, description="Include an unlocked (0) control step")
apply_best: bool = Field(False, description="Write the winning value into the profile")
class RequestVramRequest(BaseModel):
@@ -263,7 +299,7 @@ async def sse_telemetry_stream(request: Request):
if await request.is_disconnected():
break
try:
snap = await asyncio.wait_for(q.get(), timeout=15.0)
snap = await asyncio.wait_for(q.get(), timeout=5.0)
if snap is None: # shutdown sentinel
break
yield f"data: {json.dumps(snap)}\n\n"
@@ -483,9 +519,10 @@ async def api_autotune_status():
async def api_autotune_sweep(req: SweepRequest):
"""Walk a clock offset upward, measuring tok/s and watching for instability at each step."""
res = await autotune.sweep(
knob=req.knob, profile=req.profile, model=req.model,
knob=req.knob, profile=req.profile, workload=req.workload, model=req.model,
start=req.start, stop=req.stop, step=req.step,
repeats=req.repeats, apply_best=req.apply_best,
repeats=req.repeats, max_steps=req.max_steps,
include_unlocked=req.include_unlocked, apply_best=req.apply_best,
)
if not res.get("success"):
raise HTTPException(status_code=400, detail=res.get("error"))

View File

@@ -666,6 +666,10 @@ class AutoArbitrator:
self.comfy_idle_since: Optional[float] = None
self.oc_profile = None
self.pending_purge = False
# While a tuning sweep is running, the arbitrator must not fight it: a ComfyUI
# benchmark would otherwise trip trigger_comfy_priority, which reapplies the whole
# 'comfy' profile and silently overwrites the clock the sweep is measuring.
self.oc_suspended = False
self.stats = {"yields": 0, "purges": 0, "yield_timeouts": 0, "deferred_purges": 0}
async def start(self):
@@ -841,9 +845,19 @@ class AutoArbitrator:
pass
await asyncio.sleep(interval)
def suspend_oc(self, reason: str = "tuning sweep") -> None:
self.oc_suspended = True
logger.info(f"Overclock auto-switching suspended ({reason})")
def resume_oc(self, profile: Optional[str] = None) -> None:
self.oc_suspended = False
# Forget the cached profile so the next transition actually reapplies.
self.oc_profile = profile
logger.info("Overclock auto-switching resumed")
def _apply_oc_profile(self, profile: str):
"""Apply an overclock profile in a background thread; only fire on transition."""
if self.oc_profile == profile:
if self.oc_suspended or self.oc_profile == profile:
return
self.oc_profile = profile
try: