Files
gpu-program-swapper/server.py
drjones f5917a0464 Report the GPU state that is real, and reconcile when it drifts
The API claimed the card was at 320W while nvidia-smi reported 370W. Three
separate defects, all introduced by me in this branch.

The readback was stale. get_gpu_state()/get_fan_status() gained a 2s cache so the
dashboard's polling would stop forking sudo every few seconds, but apply_profile
read back through that cache and its invalidation ran afterwards. A profile that
had just moved the card 370W -> 320W therefore returned a payload whose detail
string said "set to 320.00 W from 370.00 W" next to a power_limit_w of 370.0.
Caches are now cleared before the readback, which is forced.

Fan control could fail for an entire session. On boot this unit can start before
the headless X server on :8 that owns the GPU accepts connections, and the fan
assignment fails with "Error resolving target specification 'gpu:0'". Nothing
retried and nothing surfaced it, so the fans were left unconfigured with the
failure visible only inside one log line. apply_fan_control now recognises that
specific error and retries up to 5 times.

Nothing verified the result. ACTIVE_PROFILE defaults to "balanced" at import,
which is indistinguishable from "balanced was successfully applied" -- so a failed
startup apply left the app confidently reporting a profile it had never put on the
hardware. apply_profile now returns a `verified` block comparing intent against
readback and logs a warning on mismatch; profile_drift() exposes the comparison
plus whether any profile has actually been applied since startup; and the 1Hz
sampler calls reconcile_profile() once a minute to re-apply on drift.

Verified by setting 370W externally behind the service's back: the drift was
reported immediately and corrected automatically 40s later.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-02 19:51:23 -07:00

26 KiB