Fix two more stale-cache bugs; account for VRAM this service cannot reclaim
The stale-readback bug fixed in f5917a0 was a class, not an instance. Two more:
- ram_optimizer's residency report cached for 15s and was never invalidated when
anything warmed a file, so warming a model and then looking at residency showed
the state from before the warm. warm_file_to_ram now invalidates it.
- _PID_KIND_CACHE was keyed on pid alone and never expired. Linux recycles PIDs, so
a stale entry could attribute a new process's VRAM to Ollama or ComfyUI -- inside
the very snapshot the yield barrier trusts to decide whether VRAM was released.
Now keyed by (pid, process start time) and bounded.
Unmanaged VRAM. Investigating a persistence-mode warning turned up a third GPU
consumer this service does not model: stt_relay.py, holding 842 MB for nearly three
days. It was bucketed as "system" alongside gnome-shell's 3.9 MB. That conflation
matters, because ComfyUI's memory can be reclaimed and a third party's cannot, and
the reclaim path assumed ComfyUI was always to blame for missing headroom.
Processes are now bucketed ollama | comfy | desktop | unmanaged. The breakdown
reports desktop_gb and unmanaged_gb separately and names the unmanaged processes;
when a reclaim-and-retry still fails, the error identifies them rather than
implying ComfyUI was at fault; and the dashboard shows the unreclaimable total, so
headroom the arbitrator can never give back is visible rather than inferred.
Checked and deliberately not changed: persistence mode reads Disabled, but
nvidia-persistenced is active and two clients hold the GPU open continuously, so
the driver never unloads. The nvidia-smi warning is legacy noise here and is not a
source of the profile drift.
Tests: 169 (was 164). The new ones cover the bucketing, and one existing test used
Xorg as its "unknown process" fixture -- correct before a display server had its own
bucket, wrong after.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -417,6 +417,16 @@ _report_cache: Dict[str, Any] = {"ts": 0.0, "report": None}
|
||||
REPORT_TTL_S = 15.0
|
||||
|
||||
|
||||
def invalidate_cache_report() -> None:
|
||||
"""Drop the cached residency report.
|
||||
|
||||
Anything that changes what is resident must call this, or the report keeps serving
|
||||
pre-change numbers for up to REPORT_TTL_S -- so warming a model and then looking at
|
||||
residency showed the state from before the warm.
|
||||
"""
|
||||
_report_cache["report"] = None
|
||||
|
||||
|
||||
def get_cache_report(include_files: bool = True, force_refresh: bool = False) -> Dict[str, Any]:
|
||||
"""Measured page-cache residency across the whole model catalog.
|
||||
|
||||
@@ -512,6 +522,7 @@ def warm_file_to_ram(filepath: str, chunk_size: int = 16 * 1024 * 1024,
|
||||
|
||||
duration = time.perf_counter() - t0
|
||||
after = page_residency(filepath, probe_windows=32)
|
||||
invalidate_cache_report()
|
||||
return {
|
||||
"success": True,
|
||||
"filepath": filepath,
|
||||
|
||||
Reference in New Issue
Block a user