Recalibrate cache-hit thresholds against measured loads; bring MCP to parity
Calibration. The same 12.87GB model loaded through Ollama on this box: 3.1% resident (FADV_DONTNEED) -> 34.3s -> 0.38 GB/s 100% resident (force-warmed) -> 4.9s -> 2.63 GB/s The thresholds had been guessed from PCIe bus bandwidth: cache hit at >=5 GB/s. A fully warm load only reaches 2.63 GB/s, because load_duration covers host-to-device transfer and model init as well as the file read -- the page cache itself reads at 6.4 GB/s. The 5 GB/s bar was therefore unreachable, and every warm load was being reported as a partial hit. Now 2.0 / 0.8 GB/s, either side of the measured 6.9x separation. Warm-skip was also unsafe. A 12.87GB blob was skipped as already resident on the strength of twelve 2MB probe windows, then loaded at 2.44 GB/s. Skipping now requires warm_confident: an exact cachestat reading, or a probe finding every one of 32 denser samples resident. warm_file_to_ram/warm_ollama_blob take force=True, exposed on the warm-model endpoint, whose Pydantic model was missing the field entirely. MCP parity: the server had drifted well behind the REST API. Adds tools for measured residency, warm planning, VRAM requests, per-profile analytics, thermal governor control, overclock status/apply/restore, and autotune sweeps plus status -- 23 tools and 6 resources, up from 12 and 3. The telemetry store now starts in __main__ rather than at import scope, since server.py imports this module for the benchmark tool. README: replaced the remaining theoretical claims (31.5 GB/s bus rate, sub-1.5s loads, 15ms yields) with the measured numbers. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -227,6 +227,7 @@ class WarmRequest(BaseModel):
|
||||
model_name: Optional[str] = Field(None, description="Ollama model name to warm", example="gemma4:26b")
|
||||
filepath: Optional[str] = Field(None, description="Absolute file path of Safetensors/GGUF to warm into RAM")
|
||||
blob_only: bool = Field(False, description="Warm the model's weights into page cache without loading VRAM")
|
||||
force: bool = Field(False, description="Warm even if residency sampling thinks it is already resident")
|
||||
|
||||
class WarmAllRequest(BaseModel):
|
||||
budget_gb: Optional[float] = Field(None, description="Byte budget for warming; defaults to 70% of MemAvailable", example=24.0)
|
||||
@@ -362,11 +363,11 @@ async def api_warm_plan(budget_gb: Optional[float] = Query(None, description="Ov
|
||||
async def api_warm_model(req: WarmRequest):
|
||||
"""Pre-warm a specific Ollama model or file path into the Linux page cache."""
|
||||
if req.model_name and req.blob_only:
|
||||
return ram_optimizer.warm_ollama_blob(req.model_name)
|
||||
return ram_optimizer.warm_ollama_blob(req.model_name, force=req.force)
|
||||
if req.model_name:
|
||||
return await ram_optimizer.warm_ollama_model(req.model_name, keep_alive="1m")
|
||||
if req.filepath:
|
||||
return ram_optimizer.warm_file_to_ram(req.filepath)
|
||||
return ram_optimizer.warm_file_to_ram(req.filepath, force=req.force)
|
||||
raise HTTPException(status_code=400, detail="model_name or filepath required")
|
||||
|
||||
@app.get("/api/cache/report", summary="Measured Page-Cache Residency", tags=["Memory Optimization"])
|
||||
|
||||
Reference in New Issue
Block a user