Recalibrate cache-hit thresholds against measured loads; bring MCP to parity

Calibration. The same 12.87GB model loaded through Ollama on this box:

  3.1% resident (FADV_DONTNEED) -> 34.3s -> 0.38 GB/s
  100% resident (force-warmed)  ->  4.9s -> 2.63 GB/s

The thresholds had been guessed from PCIe bus bandwidth: cache hit at >=5 GB/s. A fully
warm load only reaches 2.63 GB/s, because load_duration covers host-to-device transfer
and model init as well as the file read -- the page cache itself reads at 6.4 GB/s. The
5 GB/s bar was therefore unreachable, and every warm load was being reported as a
partial hit. Now 2.0 / 0.8 GB/s, either side of the measured 6.9x separation.

Warm-skip was also unsafe. A 12.87GB blob was skipped as already resident on the
strength of twelve 2MB probe windows, then loaded at 2.44 GB/s. Skipping now requires
warm_confident: an exact cachestat reading, or a probe finding every one of 32 denser
samples resident. warm_file_to_ram/warm_ollama_blob take force=True, exposed on the
warm-model endpoint, whose Pydantic model was missing the field entirely.

MCP parity: the server had drifted well behind the REST API. Adds tools for measured
residency, warm planning, VRAM requests, per-profile analytics, thermal governor
control, overclock status/apply/restore, and autotune sweeps plus status -- 23 tools
and 6 resources, up from 12 and 3. The telemetry store now starts in __main__ rather
than at import scope, since server.py imports this module for the benchmark tool.

README: replaced the remaining theoretical claims (31.5 GB/s bus rate, sub-1.5s loads,
15ms yields) with the measured numbers.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
drjones
2026-08-28 14:40:05 -07:00
parent c689ec8711
commit 01d2f4cfdd
5 changed files with 204 additions and 32 deletions

View File

@@ -29,10 +29,20 @@ COMFY_API_BASE = "http://127.0.0.1:8188"
# Circular buffer for transition events (the durable log lives in telemetry_store)
SWITCH_HISTORY = deque(maxlen=50)
# Bandwidth thresholds used to classify how a model actually got into VRAM.
# PCIe 4.0 x16 tops out near 31.5 GB/s; this NVMe sustains well under 2 GB/s.
RAM_HIT_GBPS = 5.0
PARTIAL_HIT_GBPS = 1.5
# Bandwidth thresholds for classifying how a model reached VRAM, calibrated by measuring
# the same 12.87 GB model loaded cold and warm on this box (2026-08-28):
#
# 3.1% resident -> 34.3 s -> 0.38 GB/s
# 100% resident -> 4.9 s -> 2.63 GB/s
#
# The first cut at these numbers assumed a page-cache-fed load would approach the bus
# rate and set the cache-hit bar at 5 GB/s. It does not: Ollama's load_duration covers
# host-to-device transfer and model initialisation as well as the file read, so a fully
# resident model still reports ~2.6 GB/s while the page cache itself reads at 6.4 GB/s.
# A 5 GB/s bar could therefore never be met, and every warm load was being reported as
# a partial hit. Thresholds now sit either side of the measured 6.9x separation.
RAM_HIT_GBPS = 2.0
PARTIAL_HIT_GBPS = 0.8
# How long Ollama's VRAM may take to actually drain before we stop waiting.
# Ollama will not unload a model while a generation is in flight, so a short ceiling
@@ -557,11 +567,14 @@ def classify_load(size_bytes: int, load_duration_ms: float) -> Dict[str, Any]:
return {"cache_status": "Already in VRAM", "load_gbps": None, "is_ram_hit": True}
if not size_bytes:
# No size on record — fall back to the old heuristic, but say so.
# Without a size we cannot compute bandwidth at all; this is a guess and is
# labelled as one. 8s roughly splits the measured warm (4.9s) and cold (34.3s)
# loads for a mid-size model, but it is meaningless for very small or large ones.
return {
"cache_status": "RAM Cache Hit ⚡" if load_duration_ms < 2500 else "Cold Disk Load 💾",
"cache_status": "RAM Cache Hit ⚡" if load_duration_ms < 8000 else "Cold Disk Load 💾",
"load_gbps": None,
"is_ram_hit": load_duration_ms < 2500,
"detail": "size unknown, fell back to duration heuristic",
"is_ram_hit": load_duration_ms < 8000,
"detail": "size unknown, fell back to a duration guess",
}
gbps = (size_bytes / (1024**3)) / (load_duration_ms / 1000.0)
if gbps >= RAM_HIT_GBPS: