Report the GPU state that is real, and reconcile when it drifts
The API claimed the card was at 320W while nvidia-smi reported 370W. Three separate defects, all introduced by me in this branch. The readback was stale. get_gpu_state()/get_fan_status() gained a 2s cache so the dashboard's polling would stop forking sudo every few seconds, but apply_profile read back through that cache and its invalidation ran afterwards. A profile that had just moved the card 370W -> 320W therefore returned a payload whose detail string said "set to 320.00 W from 370.00 W" next to a power_limit_w of 370.0. Caches are now cleared before the readback, which is forced. Fan control could fail for an entire session. On boot this unit can start before the headless X server on :8 that owns the GPU accepts connections, and the fan assignment fails with "Error resolving target specification 'gpu:0'". Nothing retried and nothing surfaced it, so the fans were left unconfigured with the failure visible only inside one log line. apply_fan_control now recognises that specific error and retries up to 5 times. Nothing verified the result. ACTIVE_PROFILE defaults to "balanced" at import, which is indistinguishable from "balanced was successfully applied" -- so a failed startup apply left the app confidently reporting a profile it had never put on the hardware. apply_profile now returns a `verified` block comparing intent against readback and logs a warning on mismatch; profile_drift() exposes the comparison plus whether any profile has actually been applied since startup; and the 1Hz sampler calls reconcile_profile() once a minute to re-apply on drift. Verified by setting 370W externally behind the service's back: the drift was reported immediately and corrected automatically 40s later. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -49,6 +49,8 @@ class TelemetryBroker:
|
||||
layer, so neither needs to poll the GPU on its own.
|
||||
"""
|
||||
|
||||
RECONCILE_EVERY_N = 60 # once a minute at 1 Hz
|
||||
|
||||
def __init__(self, interval_s: float = 1.0) -> None:
|
||||
self.interval_s = interval_s
|
||||
self.snapshot: Dict[str, Any] = {}
|
||||
@@ -107,6 +109,13 @@ class TelemetryBroker:
|
||||
# Feed the governor and the durable store from the sample we already have.
|
||||
thermal_governor.governor.observe(snap.get("gpu", {}),
|
||||
overclock_manager.ACTIVE_PROFILE)
|
||||
|
||||
# Cheap, infrequent check that the card still matches the active profile.
|
||||
# A startup apply can fail silently (the headless X server may not be up
|
||||
# yet), and an external tool can move the power limit underneath us.
|
||||
if self.samples % self.RECONCILE_EVERY_N == 0:
|
||||
await asyncio.get_running_loop().run_in_executor(
|
||||
None, overclock_manager.reconcile_profile)
|
||||
telemetry_store.record_telemetry(
|
||||
snap.get("gpu", {}), snap.get("ram", {}),
|
||||
profile=overclock_manager.ACTIVE_PROFILE,
|
||||
|
||||
Reference in New Issue
Block a user