Report the GPU state that is real, and reconcile when it drifts

The API claimed the card was at 320W while nvidia-smi reported 370W. Three
separate defects, all introduced by me in this branch.

The readback was stale. get_gpu_state()/get_fan_status() gained a 2s cache so the
dashboard's polling would stop forking sudo every few seconds, but apply_profile
read back through that cache and its invalidation ran afterwards. A profile that
had just moved the card 370W -> 320W therefore returned a payload whose detail
string said "set to 320.00 W from 370.00 W" next to a power_limit_w of 370.0.
Caches are now cleared before the readback, which is forced.

Fan control could fail for an entire session. On boot this unit can start before
the headless X server on :8 that owns the GPU accepts connections, and the fan
assignment fails with "Error resolving target specification 'gpu:0'". Nothing
retried and nothing surfaced it, so the fans were left unconfigured with the
failure visible only inside one log line. apply_fan_control now recognises that
specific error and retries up to 5 times.

Nothing verified the result. ACTIVE_PROFILE defaults to "balanced" at import,
which is indistinguishable from "balanced was successfully applied" -- so a failed
startup apply left the app confidently reporting a profile it had never put on the
hardware. apply_profile now returns a `verified` block comparing intent against
readback and logs a warning on mismatch; profile_drift() exposes the comparison
plus whether any profile has actually been applied since startup; and the 1Hz
sampler calls reconcile_profile() once a minute to re-apply on drift.

Verified by setting 370W externally behind the service's back: the drift was
reported immediately and corrected automatically 40s later.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
drjones
2026-09-02 19:51:23 -07:00
parent 868d82794d
commit f5917a0464
2 changed files with 109 additions and 6 deletions

View File

@@ -49,6 +49,8 @@ class TelemetryBroker:
layer, so neither needs to poll the GPU on its own.
"""
RECONCILE_EVERY_N = 60 # once a minute at 1 Hz
def __init__(self, interval_s: float = 1.0) -> None:
self.interval_s = interval_s
self.snapshot: Dict[str, Any] = {}
@@ -107,6 +109,13 @@ class TelemetryBroker:
# Feed the governor and the durable store from the sample we already have.
thermal_governor.governor.observe(snap.get("gpu", {}),
overclock_manager.ACTIVE_PROFILE)
# Cheap, infrequent check that the card still matches the active profile.
# A startup apply can fail silently (the headless X server may not be up
# yet), and an external tool can move the power limit underneath us.
if self.samples % self.RECONCILE_EVERY_N == 0:
await asyncio.get_running_loop().run_in_executor(
None, overclock_manager.reconcile_profile)
telemetry_store.record_telemetry(
snap.get("gpu", {}), snap.get("ram", {}),
profile=overclock_manager.ACTIVE_PROFILE,