The API claimed the card was at 320W while nvidia-smi reported 370W. Three separate defects, all introduced by me in this branch. The readback was stale. get_gpu_state()/get_fan_status() gained a 2s cache so the dashboard's polling would stop forking sudo every few seconds, but apply_profile read back through that cache and its invalidation ran afterwards. A profile that had just moved the card 370W -> 320W therefore returned a payload whose detail string said "set to 320.00 W from 370.00 W" next to a power_limit_w of 370.0. Caches are now cleared before the readback, which is forced. Fan control could fail for an entire session. On boot this unit can start before the headless X server on :8 that owns the GPU accepts connections, and the fan assignment fails with "Error resolving target specification 'gpu:0'". Nothing retried and nothing surfaced it, so the fans were left unconfigured with the failure visible only inside one log line. apply_fan_control now recognises that specific error and retries up to 5 times. Nothing verified the result. ACTIVE_PROFILE defaults to "balanced" at import, which is indistinguishable from "balanced was successfully applied" -- so a failed startup apply left the app confidently reporting a profile it had never put on the hardware. apply_profile now returns a `verified` block comparing intent against readback and logs a warning on mismatch; profile_drift() exposes the comparison plus whether any profile has actually been applied since startup; and the 1Hz sampler calls reconcile_profile() once a minute to re-apply on drift. Verified by setting 370W externally behind the service's back: the drift was reported immediately and corrected automatically 40s later. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
26 KiB
26 KiB