Recalibrate cache-hit thresholds against measured loads; bring MCP to parity
Calibration. The same 12.87GB model loaded through Ollama on this box: 3.1% resident (FADV_DONTNEED) -> 34.3s -> 0.38 GB/s 100% resident (force-warmed) -> 4.9s -> 2.63 GB/s The thresholds had been guessed from PCIe bus bandwidth: cache hit at >=5 GB/s. A fully warm load only reaches 2.63 GB/s, because load_duration covers host-to-device transfer and model init as well as the file read -- the page cache itself reads at 6.4 GB/s. The 5 GB/s bar was therefore unreachable, and every warm load was being reported as a partial hit. Now 2.0 / 0.8 GB/s, either side of the measured 6.9x separation. Warm-skip was also unsafe. A 12.87GB blob was skipped as already resident on the strength of twelve 2MB probe windows, then loaded at 2.44 GB/s. Skipping now requires warm_confident: an exact cachestat reading, or a probe finding every one of 32 denser samples resident. warm_file_to_ram/warm_ollama_blob take force=True, exposed on the warm-model endpoint, whose Pydantic model was missing the field entirely. MCP parity: the server had drifted well behind the REST API. Adds tools for measured residency, warm planning, VRAM requests, per-profile analytics, thermal governor control, overclock status/apply/restore, and autotune sweeps plus status -- 23 tools and 6 resources, up from 12 and 3. The telemetry store now starts in __main__ rather than at import scope, since server.py imports this module for the benchmark tool. README: replaced the remaining theoretical claims (31.5 GB/s bus rate, sub-1.5s loads, 15ms yields) with the measured numbers. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -29,10 +29,20 @@ COMFY_API_BASE = "http://127.0.0.1:8188"
|
||||
# Circular buffer for transition events (the durable log lives in telemetry_store)
|
||||
SWITCH_HISTORY = deque(maxlen=50)
|
||||
|
||||
# Bandwidth thresholds used to classify how a model actually got into VRAM.
|
||||
# PCIe 4.0 x16 tops out near 31.5 GB/s; this NVMe sustains well under 2 GB/s.
|
||||
RAM_HIT_GBPS = 5.0
|
||||
PARTIAL_HIT_GBPS = 1.5
|
||||
# Bandwidth thresholds for classifying how a model reached VRAM, calibrated by measuring
|
||||
# the same 12.87 GB model loaded cold and warm on this box (2026-08-28):
|
||||
#
|
||||
# 3.1% resident -> 34.3 s -> 0.38 GB/s
|
||||
# 100% resident -> 4.9 s -> 2.63 GB/s
|
||||
#
|
||||
# The first cut at these numbers assumed a page-cache-fed load would approach the bus
|
||||
# rate and set the cache-hit bar at 5 GB/s. It does not: Ollama's load_duration covers
|
||||
# host-to-device transfer and model initialisation as well as the file read, so a fully
|
||||
# resident model still reports ~2.6 GB/s while the page cache itself reads at 6.4 GB/s.
|
||||
# A 5 GB/s bar could therefore never be met, and every warm load was being reported as
|
||||
# a partial hit. Thresholds now sit either side of the measured 6.9x separation.
|
||||
RAM_HIT_GBPS = 2.0
|
||||
PARTIAL_HIT_GBPS = 0.8
|
||||
|
||||
# How long Ollama's VRAM may take to actually drain before we stop waiting.
|
||||
# Ollama will not unload a model while a generation is in flight, so a short ceiling
|
||||
@@ -557,11 +567,14 @@ def classify_load(size_bytes: int, load_duration_ms: float) -> Dict[str, Any]:
|
||||
return {"cache_status": "Already in VRAM", "load_gbps": None, "is_ram_hit": True}
|
||||
if not size_bytes:
|
||||
# No size on record — fall back to the old heuristic, but say so.
|
||||
# Without a size we cannot compute bandwidth at all; this is a guess and is
|
||||
# labelled as one. 8s roughly splits the measured warm (4.9s) and cold (34.3s)
|
||||
# loads for a mid-size model, but it is meaningless for very small or large ones.
|
||||
return {
|
||||
"cache_status": "RAM Cache Hit ⚡" if load_duration_ms < 2500 else "Cold Disk Load 💾",
|
||||
"cache_status": "RAM Cache Hit ⚡" if load_duration_ms < 8000 else "Cold Disk Load 💾",
|
||||
"load_gbps": None,
|
||||
"is_ram_hit": load_duration_ms < 2500,
|
||||
"detail": "size unknown, fell back to duration heuristic",
|
||||
"is_ram_hit": load_duration_ms < 8000,
|
||||
"detail": "size unknown, fell back to a duration guess",
|
||||
}
|
||||
gbps = (size_bytes / (1024**3)) / (load_duration_ms / 1000.0)
|
||||
if gbps >= RAM_HIT_GBPS:
|
||||
|
||||
Reference in New Issue
Block a user