Tune profiles from measurement; add a diffusion benchmark to close the loop
The profiles were hand-written and had never been checked against the hardware. Adding a ComfyUI benchmark alongside the existing decode one made the compute side measurable for the first time, and most of what the profiles configured turned out to do nothing. Measured on this card (RTX 4080 SUPER, driver 595.84): - LLM decode is not power-bound: 73.0-73.5 tok/s flat from 222W to 370W, with the card never drawing more than 224W at any limit. The ollama profile's 370W did nothing. - Diffusion is power-bound: 5.48 it/s @222W rising to 6.71 @370W, so comfy's 370W is worth a real +2.8% over the 320W stock default. - Clock locks did nothing for either workload: 72.6 tok/s locked at 11251MHz vs 72.7 unlocked; 6.77 it/s locked at 3105MHz vs 6.73 unlocked, and 6.78 at 2400MHz. - Memory bandwidth is still the decode bottleneck (5001MHz halves throughput to 35.9 tok/s), confirming the profile's premise -- the card just gets there unaided. - Fans: 48,435 samples show 81C all-time max and zero thermal throttle events, while the ollama profile held 49.6C average by running fans at 87%. All profiles now use automatic fans and let the thermal governor escalate on demand. Code changes supporting that: - _diffusion_benchmark() queues a fixed SDXL graph via ComfyUI's API. The seed must vary per run: ComfyUI caches by node inputs, so a fixed seed returned in ~1ms without executing. Implausibly fast results are now rejected as cache hits rather than recorded as record scores. - The arbitrator's automatic profile switching is suspended during a sweep. A diffusion benchmark trips trigger_comfy_priority, which reapplies the whole profile and would silently overwrite the clock being measured. - _supported_clocks() queries the mem,gr pair; asking for a single field returned one column and reading index 1 yielded an empty list rather than an error. Graphics clocks are subsampled (the card enumerates 194 of them) and lock sweeps include an explicit unlocked control step. - offsets_supported() probes once and apply_profile skips inert offset levers with an explanation instead of pretending they applied. - Profiles carry a 'measured' field recording the evidence behind each setting. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -666,6 +666,10 @@ class AutoArbitrator:
|
||||
self.comfy_idle_since: Optional[float] = None
|
||||
self.oc_profile = None
|
||||
self.pending_purge = False
|
||||
# While a tuning sweep is running, the arbitrator must not fight it: a ComfyUI
|
||||
# benchmark would otherwise trip trigger_comfy_priority, which reapplies the whole
|
||||
# 'comfy' profile and silently overwrites the clock the sweep is measuring.
|
||||
self.oc_suspended = False
|
||||
self.stats = {"yields": 0, "purges": 0, "yield_timeouts": 0, "deferred_purges": 0}
|
||||
|
||||
async def start(self):
|
||||
@@ -841,9 +845,19 @@ class AutoArbitrator:
|
||||
pass
|
||||
await asyncio.sleep(interval)
|
||||
|
||||
def suspend_oc(self, reason: str = "tuning sweep") -> None:
|
||||
self.oc_suspended = True
|
||||
logger.info(f"Overclock auto-switching suspended ({reason})")
|
||||
|
||||
def resume_oc(self, profile: Optional[str] = None) -> None:
|
||||
self.oc_suspended = False
|
||||
# Forget the cached profile so the next transition actually reapplies.
|
||||
self.oc_profile = profile
|
||||
logger.info("Overclock auto-switching resumed")
|
||||
|
||||
def _apply_oc_profile(self, profile: str):
|
||||
"""Apply an overclock profile in a background thread; only fire on transition."""
|
||||
if self.oc_profile == profile:
|
||||
if self.oc_suspended or self.oc_profile == profile:
|
||||
return
|
||||
self.oc_profile = profile
|
||||
try:
|
||||
|
||||
Reference in New Issue
Block a user