Recalibrate cache-hit thresholds against measured loads; bring MCP to parity

Calibration. The same 12.87GB model loaded through Ollama on this box:

  3.1% resident (FADV_DONTNEED) -> 34.3s -> 0.38 GB/s
  100% resident (force-warmed)  ->  4.9s -> 2.63 GB/s

The thresholds had been guessed from PCIe bus bandwidth: cache hit at >=5 GB/s. A fully
warm load only reaches 2.63 GB/s, because load_duration covers host-to-device transfer
and model init as well as the file read -- the page cache itself reads at 6.4 GB/s. The
5 GB/s bar was therefore unreachable, and every warm load was being reported as a
partial hit. Now 2.0 / 0.8 GB/s, either side of the measured 6.9x separation.

Warm-skip was also unsafe. A 12.87GB blob was skipped as already resident on the
strength of twelve 2MB probe windows, then loaded at 2.44 GB/s. Skipping now requires
warm_confident: an exact cachestat reading, or a probe finding every one of 32 denser
samples resident. warm_file_to_ram/warm_ollama_blob take force=True, exposed on the
warm-model endpoint, whose Pydantic model was missing the field entirely.

MCP parity: the server had drifted well behind the REST API. Adds tools for measured
residency, warm planning, VRAM requests, per-profile analytics, thermal governor
control, overclock status/apply/restore, and autotune sweeps plus status -- 23 tools
and 6 resources, up from 12 and 3. The telemetry store now starts in __main__ rather
than at import scope, since server.py imports this module for the benchmark tool.

README: replaced the remaining theoretical claims (31.5 GB/s bus rate, sub-1.5s loads,
15ms yields) with the measured numbers.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
drjones
2026-08-28 14:40:05 -07:00
parent c689ec8711
commit 01d2f4cfdd
5 changed files with 204 additions and 32 deletions

View File

@@ -107,8 +107,8 @@ only if the card actually needs it.
* **Server-Sent Events (SSE)**: A single background sampler produces one 1 Hz snapshot and fans it out to every subscriber via `GET /api/stream`. Previously each connected client independently re-ran the whole snapshot — NVML, `/proc/meminfo`, an HTTP round-trip each to Ollama and ComfyUI, and a recursive walk of the ComfyUI models tree with a `stat()` per checkpoint — once per second, so opening the dashboard in three tabs tripled the load on the thing it was measuring. Slow clients drop stale frames instead of stalling the sampler. * **Server-Sent Events (SSE)**: A single background sampler produces one 1 Hz snapshot and fans it out to every subscriber via `GET /api/stream`. Previously each connected client independently re-ran the whole snapshot — NVML, `/proc/meminfo`, an HTTP round-trip each to Ollama and ComfyUI, and a recursive walk of the ComfyUI models tree with a `stat()` per checkpoint — once per second, so opening the dashboard in three tabs tripled the load on the thing it was measuring. Slow clients drop stale frames instead of stalling the sampler.
### 🤖 Model Context Protocol (MCP 2.0) Server ### 🤖 Model Context Protocol (MCP 2.0) Server
* **12 Native Agentic Tools**: Allows AI agents (Antigravity CLI, Claude Desktop, Cursor) to manage GPU resources, trigger model hot-swaps, tune fan curves, and inspect telemetry. * **23 Native Agentic Tools**: Allows AI agents (Antigravity CLI, Claude Desktop, Cursor) to manage GPU resources, trigger model hot-swaps, measure page-cache residency, read persisted performance analytics, drive the thermal governor, and run overclock sweeps.
* **3 Live MCP Resources**: Exposes live metrics, model catalogs, and switch logs as streamable resources (`gpu://metrics/live`, `gpu://models/catalog`, `gpu://history/switches`). * **6 Live MCP Resources**: Live metrics, model catalog, switch log, measured cache residency, per-profile analytics, and the overclock profiles with the evidence behind each setting.
* **Dual Transport Support**: Run via standard input/output (`--stdio`) or network Server-Sent Events (`--sse --port 8001`). * **Dual Transport Support**: Run via standard input/output (`--stdio`) or network Server-Sent Events (`--sse --port 8001`).
### ⏱️ Automated Latency & Throughput Benchmark Engine ### ⏱️ Automated Latency & Throughput Benchmark Engine
@@ -135,19 +135,32 @@ flowchart TD
REST["REST API & OpenAPI Docs"] REST["REST API & OpenAPI Docs"]
MCP["Model Context Protocol (MCP 2.0)"] MCP["Model Context Protocol (MCP 2.0)"]
SSE["1Hz Real-Time SSE Stream"] SSE["1Hz Real-Time SSE Stream"]
Arbitrator["VRAM Arbitrator (15ms Soft-Yield)"] Arbitrator["VRAM Arbitrator (confirmed yield)"]
Overclock["Overclock & Fan Manager"] Overclock["Overclock & Fan Manager"]
Warmer["Page Cache Pre-Warmer"] Warmer["Page Cache Pre-Warmer"]
end end
HostRAM <== "PCIe 4.0 x16 Bus (~31.5 GB/s Hot-Swap)" ==> GPU HostRAM <== "PCIe 4.0 x16 Bus (measured 2.6 GB/s warm model load)" ==> GPU
Orchestrator --> GPU Orchestrator --> GPU
Orchestrator --> HostRAM Orchestrator --> HostRAM
``` ```
### The Physics of Sub-Second Switching ### The Physics of Sub-Second Switching
* **Host RAM as Staging**: Active LLMs and diffusion checkpoints remain resident in the 64GB Linux Page Cache. * **Host RAM as Staging**: Active LLMs and diffusion checkpoints remain resident in the 64GB Linux Page Cache.
* **PCIe 4.0 x16 Hot-Swapping**: Transferring weights across PCIe 4.0 x16 achieves **~31.5 GB/s** bandwidth, reducing model loads from 30+ seconds (disk) to **under 1.5 seconds**. * **Warm vs cold model loads, measured.** The same 12.87 GB model, loaded through Ollama on this box:
| Page-cache residency | Load time | Effective rate |
| :--- | :--- | :--- |
| 3.1% (dropped with `FADV_DONTNEED`) | 34.3 s | 0.38 GB/s |
| 100% (force-warmed) | 4.9 s | 2.63 GB/s |
A **6.9× speedup**, and the reason the page cache matters. Note the effective rate is
well below the PCIe 4.0 x16 bus rate and below the 6.4 GB/s the page cache itself
reads at: Ollama's `load_duration` also covers host-to-device transfer and model
initialisation, not just the file read. Classification thresholds are calibrated
against these measured numbers rather than the theoretical bus bandwidth — an earlier
5 GB/s cache-hit bar sat above what a fully warm load can even achieve, so every warm
load was misreported as a partial hit.
* **Soft-Yielding**: Dropping Ollama's VRAM allocation via `keep_alive: 0` preserves the weights in host RAM. Measured on this box: the HTTP request returns in **~63 ms**, and the driver finishes releasing 14.9 GB **~77 ms after that**. HyperSwap waits for the second number before handing VRAM to ComfyUI — the earlier "~15 ms" figure timed the request, not the release. * **Soft-Yielding**: Dropping Ollama's VRAM allocation via `keep_alive: 0` preserves the weights in host RAM. Measured on this box: the HTTP request returns in **~63 ms**, and the driver finishes releasing 14.9 GB **~77 ms after that**. HyperSwap waits for the second number before handing VRAM to ComfyUI — the earlier "~15 ms" figure timed the request, not the release.
--- ---
@@ -220,19 +233,33 @@ HyperSwap includes a native **MCP 2.0 server** (`mcp_server.py`) exposing orches
| **`set_gpu_fan_speed`** | `mode` (str), `percent` (optional int) | Sets fan speed mode (`auto`\|`manual`) and target PWM % (30–100%). | | **`set_gpu_fan_speed`** | `mode` (str), `percent` (optional int) | Sets fan speed mode (`auto`\|`manual`) and target PWM % (30–100%). |
| **`get_host_memory_status`** | *None* | 64GB host RAM breakdown, active page cache size, and cache ratio. | | **`get_host_memory_status`** | *None* | 64GB host RAM breakdown, active page cache size, and cache ratio. |
| **`switch_ollama_model`** | `model_name` (str), `keep_alive` (str) | Hot-swaps active LLM in VRAM, measures latency (ms) and tokens/sec. | | **`switch_ollama_model`** | `model_name` (str), `keep_alive` (str) | Hot-swaps active LLM in VRAM, measures latency (ms) and tokens/sec. |
| **`soft_yield_ollama_vram`** | `model_name` (optional str) | Yields Ollama VRAM to 0 MB in ~15ms while keeping model weights in RAM cache. | | **`soft_yield_ollama_vram`** | `model_name` (optional str) | Yields Ollama VRAM to 0 MB and waits for NVML to confirm the driver actually released it. Returns the request/confirm split. |
| **`purge_comfyui_vram`** | *None* | Purges loaded diffusion models from ComfyUI pipeline VRAM. | | **`purge_comfyui_vram`** | *None* | Purges loaded diffusion models from ComfyUI pipeline VRAM. |
| **`prewarm_all_models_to_ram`** | *None* | Faults all local LLM and diffusion checkpoints into Linux OS page cache. | | **`prewarm_all_models_to_ram`** | *None* | Warms the highest-value models into page cache within a byte budget, skipping what is already resident. |
| **`prewarm_single_model`** | `model_name` (optional str), `filepath` (optional str) | Pre-warms a single GGUF or Safetensors file into RAM. | | **`prewarm_single_model`** | `model_name` (optional str), `filepath` (optional str) | Pre-warms a single GGUF or Safetensors file into RAM. |
| **`list_available_models`** | *None* | Lists all installed Ollama models and discovered ComfyUI Safetensors on disk. | | **`list_available_models`** | *None* | Lists all installed Ollama models and discovered ComfyUI Safetensors on disk. |
| **`get_switch_history`** | `limit` (int, default 20) | Retrieves recent switch events, millisecond latencies, and RAM hit status. | | **`get_switch_history`** | `limit` (int, default 20) | Retrieves recent switch events, millisecond latencies, and RAM hit status. |
| **`run_model_switch_benchmark`**| `iterations` (int, default 2) | Automated round-trip latency benchmark between installed models. | | **`run_model_switch_benchmark`**| `iterations` (int, default 2) | Automated round-trip latency benchmark between installed models. |
| **`get_page_cache_residency`** | `include_files` (bool) | Measured page-cache residency per model file, with the measurement method used for each. |
| **`get_warm_plan`** | `budget_gb` (optional float) | Previews what warming would read and skip, ranked by recency/frequency. Does not warm. |
| **`request_vram_for_ollama`** | `needed_gb` (float) | Purges ComfyUI's checkpoints immediately if VRAM headroom is short, bypassing the idle timer. |
| **`get_profile_performance`** | `days` (float, default 7) | Measured tok/s and thermals per overclock profile, from persisted history. |
| **`get_thermal_governor_status`** | *None* | Current derate level, the reason for it, and escalation history. |
| **`set_thermal_governor`** | `enabled` (optional bool), `reset` (bool) | Enable/disable the governor, or clear an active derate. |
| **`get_overclock_status`** | *None* | Active profile, all profiles with their evidence, and which levers this driver honours. |
| **`apply_overclock_profile`** | `profile` (str) | Apply `ollama` \| `comfy` \| `balanced`. |
| **`restore_stock_gpu_state`** | *None* | Drop clock locks and offsets, restore default power limit, return fans to automatic. |
| **`run_overclock_sweep`** | `knob`, `profile`, `workload`, `start`, `stop`, `repeats`, `apply_best` | Sweep a knob against a real workload and report the fastest stable value. Verifies the knob moves the hardware first. Takes minutes. |
| **`get_autotune_status`** | *None* | Sweep progress, the last result table, and all recorded autotune steps. |
### MCP Resources List ### MCP Resources List
* `gpu://metrics/live`: Real-time snapshot of GPU sensors and RAM page cache. * `gpu://metrics/live`: Real-time snapshot of GPU sensors and RAM page cache.
* `gpu://models/catalog`: Catalog of all discovered GGUF and Safetensors models. * `gpu://models/catalog`: Catalog of all discovered GGUF and Safetensors models.
* `gpu://history/switches`: Event log of recent model transitions and swap speeds. * `gpu://history/switches`: Event log of recent model transitions and swap speeds.
* `gpu://cache/residency`: Measured page-cache residency across every model on disk.
* `gpu://analytics/profiles`: Measured throughput and thermals per overclock profile.
* `gpu://overclock/profiles`: Overclock profiles including the measurement behind each setting.
--- ---

View File

@@ -7,16 +7,19 @@ import logging
from typing import Dict, List, Any, Optional from typing import Dict, List, Any, Optional
from mcp.server import MCPServer from mcp.server import MCPServer
import ram_optimizer import autotune
import vram_arbitrator
import overclock_manager import overclock_manager
import ram_optimizer
import telemetry_store
import thermal_governor
import vram_arbitrator
logging.basicConfig(level=logging.INFO, format="%(asctime)s [%(levelname)s] %(name)s: %(message)s") logging.basicConfig(level=logging.INFO, format="%(asctime)s [%(levelname)s] %(name)s: %(message)s")
logger = logging.getLogger("gpu_swapper_mcp") logger = logging.getLogger("gpu_swapper_mcp")
mcp = MCPServer( mcp = MCPServer(
name="gpu-program-swapper", name="gpu-program-swapper",
version="1.0.0", version="2.0.0",
description="Orchestrates high-speed GPU VRAM hot-swaps between Ollama LLMs and ComfyUI with 64GB RAM cache telemetry." description="Orchestrates high-speed GPU VRAM hot-swaps between Ollama LLMs and ComfyUI with 64GB RAM cache telemetry."
) )
@@ -133,6 +136,105 @@ def set_gpu_fan_speed(mode: str = "auto", percent: Optional[int] = None) -> str:
res = overclock_manager.set_fan_auto() res = overclock_manager.set_fan_auto()
return json.dumps(res, indent=2) return json.dumps(res, indent=2)
@mcp.tool()
def get_page_cache_residency(include_files: bool = True) -> str:
"""Measure how much of each model on disk is genuinely resident in the Linux page cache.
Uses cachestat(2) where the kernel permits it and a read-rate probe where it does not
(Ollama blobs are owned by another user). Reports which method was used per file, and
marks anything it cannot measure rather than guessing."""
report = ram_optimizer.get_cache_report(include_files=include_files)
report["capability"] = ram_optimizer.residency_capability()
return json.dumps(report, indent=2, default=str)
@mcp.tool()
def get_warm_plan(budget_gb: Optional[float] = None) -> str:
"""Preview which models pre-warming would load into RAM, in what order, and what it
would skip — ranked by recency/frequency and capped by a byte budget. Does not warm."""
return json.dumps(ram_optimizer.build_warm_plan(budget_gb), indent=2, default=str)
@mcp.tool()
async def request_vram_for_ollama(needed_gb: float = 0.0) -> str:
"""Free VRAM for an LLM right now: purges ComfyUI's cached checkpoints immediately if
there is not enough headroom, instead of waiting for the normal idle timer."""
res = await vram_arbitrator.arbitrator.request_vram_for_ollama(needed_gb)
return json.dumps(res, indent=2, default=str)
@mcp.tool()
def get_profile_performance(days: float = 7.0) -> str:
"""Compare measured decode throughput and thermals per overclock profile, from
persisted history. Answers whether a given profile is actually delivering more tok/s."""
return json.dumps({
"window_days": days,
"profiles": telemetry_store.profile_comparison(days),
"swaps": telemetry_store.swap_stats(days),
}, indent=2, default=str)
@mcp.tool()
def get_thermal_governor_status() -> str:
"""Current thermal derate level, why it was applied, and the escalation history."""
return json.dumps(thermal_governor.governor.get_status(), indent=2, default=str)
@mcp.tool()
def set_thermal_governor(enabled: Optional[bool] = None, reset: bool = False) -> str:
"""Enable or disable the thermal governor, or clear an active derate and reapply the
full profile."""
if enabled is not None:
thermal_governor.governor.set_enabled(enabled)
if reset:
thermal_governor.governor.reset()
return json.dumps(thermal_governor.governor.get_status(), indent=2, default=str)
@mcp.tool()
def get_overclock_status() -> str:
"""Active overclock profile, all profiles with the evidence behind their settings, and
which hardware levers this driver actually honours (clock offsets are ignored on some)."""
return json.dumps(overclock_manager.get_status(), indent=2, default=str)
@mcp.tool()
def apply_overclock_profile(profile: str) -> str:
"""Apply an overclock profile by name: ollama | comfy | balanced."""
return json.dumps(overclock_manager.apply_profile(profile), indent=2, default=str)
@mcp.tool()
def restore_stock_gpu_state() -> str:
"""Drop all clock locks and offsets, restore the default power limit, and return the
fans to automatic control."""
return json.dumps(overclock_manager.restore_safe("MCP request"), indent=2, default=str)
@mcp.tool()
async def run_overclock_sweep(knob: str = "power_limit_w", profile: str = "ollama",
workload: str = "auto", start: Optional[int] = None,
stop: Optional[int] = None, repeats: int = 1,
apply_best: bool = False) -> str:
"""Sweep one GPU knob against a real workload and report the fastest stable value.
knob: power_limit_w | lock_mem_mhz | lock_core_max | mem_offset_mhz | core_offset_mhz
workload: 'ollama' (decode tok/s), 'comfy' (SDXL it/s), or 'auto' to match the profile.
Verifies the knob actually moves the hardware before sweeping, refuses to run while
ComfyUI is busy, and always restores the original profile. Takes minutes."""
res = await autotune.sweep(knob=knob, profile=profile, workload=workload,
start=start, stop=stop, repeats=repeats,
apply_best=apply_best)
return json.dumps(res, indent=2, default=str)
@mcp.tool()
def get_autotune_status() -> str:
"""Sweep progress, the last sweep's full result table, and every recorded autotune step."""
return json.dumps(autotune.get_status(), indent=2, default=str)
# ========================================== # ==========================================
# MCP RESOURCES # MCP RESOURCES
# ========================================== # ==========================================
@@ -154,6 +256,21 @@ def get_switch_history_resource() -> str:
"""Recent model switch events and latencies.""" """Recent model switch events and latencies."""
return json.dumps(vram_arbitrator.get_switch_history(), indent=2) return json.dumps(vram_arbitrator.get_switch_history(), indent=2)
@mcp.resource("gpu://cache/residency")
def get_cache_residency_resource() -> str:
"""Measured page-cache residency across every model on disk."""
return json.dumps(ram_optimizer.get_cache_report(include_files=True), indent=2, default=str)
@mcp.resource("gpu://analytics/profiles")
def get_profile_analytics_resource() -> str:
"""Measured throughput and thermals per overclock profile, from persisted history."""
return json.dumps(telemetry_store.profile_comparison(7.0), indent=2, default=str)
@mcp.resource("gpu://overclock/profiles")
def get_overclock_profiles_resource() -> str:
"""Overclock profiles, including the measurement recorded behind each setting."""
return json.dumps(overclock_manager.get_status(), indent=2, default=str)
if __name__ == "__main__": if __name__ == "__main__":
import argparse import argparse
@@ -163,6 +280,11 @@ if __name__ == "__main__":
parser.add_argument("--port", type=int, default=8001, help="Port for SSE transport") parser.add_argument("--port", type=int, default=8001, help="Port for SSE transport")
args = parser.parse_args() args = parser.parse_args()
# Only when run as a standalone server. server.py imports this module for the
# benchmark tool, and starting the store at import scope would spin up a writer as a
# side effect of that import.
telemetry_store.start()
if args.sse: if args.sse:
mcp.run(transport="sse", host="0.0.0.0", port=args.port) mcp.run(transport="sse", host="0.0.0.0", port=args.port)
else: else:

View File

@@ -143,7 +143,7 @@ PROBE_WINDOW_BYTES = 2 * 1024 * 1024
PROBE_CACHED_GBPS = 1.5 PROBE_CACHED_GBPS = 1.5
def _throughput_probe(fd: int, size: int) -> Dict[str, Any]: def _throughput_probe(fd: int, size: int, windows_override: Optional[int] = None) -> Dict[str, Any]:
"""Infer residency by timing reads of small windows spread across the file. """Infer residency by timing reads of small windows spread across the file.
Used only where cachestat is not permitted (Ollama's blobs are owned by uid `ollama`). Used only where cachestat is not permitted (Ollama's blobs are owned by uid `ollama`).
@@ -158,7 +158,7 @@ def _throughput_probe(fd: int, size: int) -> Dict[str, Any]:
pollution the probe itself created, and leaving them behind would slowly warm the pollution the probe itself created, and leaving them behind would slowly warm the
cache with data nobody asked for. cache with data nobody asked for.
""" """
windows = min(PROBE_WINDOWS, max(int(size // PROBE_WINDOW_BYTES), 1)) windows = min(windows_override or PROBE_WINDOWS, max(int(size // PROBE_WINDOW_BYTES), 1))
if windows <= 0: if windows <= 0:
return {"resident_pct": 0.0, "windows": 0} return {"resident_pct": 0.0, "windows": 0}
@@ -194,7 +194,8 @@ def _throughput_probe(fd: int, size: int) -> Dict[str, Any]:
} }
def page_residency(filepath: str, allow_probe: bool = True) -> Dict[str, Any]: def page_residency(filepath: str, allow_probe: bool = True,
probe_windows: Optional[int] = None) -> Dict[str, Any]:
"""Measure what fraction of a file is resident in the Linux page cache.""" """Measure what fraction of a file is resident in the Linux page cache."""
try: try:
size = os.path.getsize(filepath) size = os.path.getsize(filepath)
@@ -216,7 +217,7 @@ def page_residency(filepath: str, allow_probe: bool = True) -> Dict[str, Any]:
method, measurable = "cachestat", True method, measurable = "cachestat", True
extra = {"dirty_pages": cs.nr_dirty, "evicted_pages": cs.nr_evicted} extra = {"dirty_pages": cs.nr_dirty, "evicted_pages": cs.nr_evicted}
elif allow_probe: elif allow_probe:
probe = _throughput_probe(fd, size) probe = _throughput_probe(fd, size, probe_windows)
pct = probe["resident_pct"] pct = probe["resident_pct"]
method, measurable = "probe", True method, measurable = "probe", True
extra = {"probe_windows": probe["windows"], "probe_median_gbps": probe.get("median_gbps")} extra = {"probe_windows": probe["windows"], "probe_median_gbps": probe.get("median_gbps")}
@@ -236,6 +237,11 @@ def page_residency(filepath: str, allow_probe: bool = True) -> Dict[str, Any]:
"method": method, "method": method,
"measurable": measurable, "measurable": measurable,
"warm": pct >= WARM_SKIP_THRESHOLD_PCT, "warm": pct >= WARM_SKIP_THRESHOLD_PCT,
# Only an exact measurement is trustworthy enough to skip work on. A probe of a
# dozen 2 MB windows can clear 90% on a file that is mostly cold -- observed
# here as a 12.87 GB "already resident" blob that then loaded at 2.44 GB/s.
"warm_confident": (method == "cachestat" and pct >= WARM_SKIP_THRESHOLD_PCT)
or (method == "probe" and pct >= 100.0),
**extra, **extra,
} }
except Exception as e: except Exception as e:
@@ -468,16 +474,19 @@ def get_cache_report(include_files: bool = True, force_refresh: bool = False) ->
# ---------------------------------------------------------------- warming # ---------------------------------------------------------------- warming
def warm_file_to_ram(filepath: str, chunk_size: int = 16 * 1024 * 1024, def warm_file_to_ram(filepath: str, chunk_size: int = 16 * 1024 * 1024,
skip_if_warm: bool = True) -> Dict[str, Any]: skip_if_warm: bool = True, force: bool = False) -> Dict[str, Any]:
"""Pre-fault a file into the Linux page cache, skipping it if already resident.""" """Pre-fault a file into the Linux page cache, skipping it only if confidently resident."""
if not os.path.exists(filepath): if not os.path.exists(filepath):
return {"success": False, "error": f"File not found: {filepath}", "duration_ms": 0} return {"success": False, "error": f"File not found: {filepath}", "duration_ms": 0}
before = page_residency(filepath) # Probe densely here: this decision skips real work, so it is worth 32 samples
if skip_if_warm and before.get("warm"): # rather than 12.
before = page_residency(filepath, probe_windows=32)
if skip_if_warm and not force and before.get("warm_confident"):
return { return {
"success": True, "filepath": filepath, "skipped": True, "success": True, "filepath": filepath, "skipped": True,
"reason": "already resident", "resident_pct": before.get("resident_pct"), "reason": "already resident", "resident_pct": before.get("resident_pct"),
"method": before.get("method"),
"size_mb": round(before.get("size_bytes", 0) / (1024**2), 2), "size_mb": round(before.get("size_bytes", 0) / (1024**2), 2),
"duration_ms": 0.0, "bytes_read": 0, "duration_ms": 0.0, "bytes_read": 0,
} }
@@ -499,7 +508,7 @@ def warm_file_to_ram(filepath: str, chunk_size: int = 16 * 1024 * 1024,
bytes_read += n bytes_read += n
duration = time.perf_counter() - t0 duration = time.perf_counter() - t0
after = page_residency(filepath) after = page_residency(filepath, probe_windows=32)
return { return {
"success": True, "success": True,
"filepath": filepath, "filepath": filepath,
@@ -551,11 +560,11 @@ async def warm_ollama_model(model_name: str, keep_alive: str = "5m") -> Dict[str
"duration_ms": round((time.perf_counter() - t0) * 1000, 2)} "duration_ms": round((time.perf_counter() - t0) * 1000, 2)}
def warm_ollama_blob(model_name: str) -> Dict[str, Any]: def warm_ollama_blob(model_name: str, force: bool = False) -> Dict[str, Any]:
"""Warm a specific Ollama model's GGUF into page cache without touching VRAM.""" """Warm a specific Ollama model's GGUF into page cache without touching VRAM."""
for f in find_ollama_model_files(): for f in find_ollama_model_files():
if f["model"] == model_name: if f["model"] == model_name:
res = warm_file_to_ram(f["full_path"]) res = warm_file_to_ram(f["full_path"], force=force)
res["model"] = model_name res["model"] = model_name
return res return res
return {"success": False, "error": f"no blob found for model '{model_name}'"} return {"success": False, "error": f"no blob found for model '{model_name}'"}
@@ -603,13 +612,13 @@ def build_warm_plan(budget_gb: Optional[float] = None) -> Dict[str, Any]:
if c["full_path"] in seen_paths: if c["full_path"] in seen_paths:
continue continue
seen_paths.add(c["full_path"]) seen_paths.add(c["full_path"])
res = page_residency(c["full_path"]) res = page_residency(c["full_path"], probe_windows=32)
entry = { entry = {
"name": c["name"], "kind": c["kind"], "full_path": c["full_path"], "name": c["name"], "kind": c["kind"], "full_path": c["full_path"],
"size_gb": c.get("size_gb", 0), "score": round(c["score"], 4), "size_gb": c.get("size_gb", 0), "score": round(c["score"], 4),
"resident_pct": res.get("resident_pct", 0.0), "resident_pct": res.get("resident_pct", 0.0),
} }
if res.get("warm"): if res.get("warm_confident"):
entry["action"] = "already-warm" entry["action"] = "already-warm"
skipped.append(entry) skipped.append(entry)
continue continue

View File

@@ -227,6 +227,7 @@ class WarmRequest(BaseModel):
model_name: Optional[str] = Field(None, description="Ollama model name to warm", example="gemma4:26b") model_name: Optional[str] = Field(None, description="Ollama model name to warm", example="gemma4:26b")
filepath: Optional[str] = Field(None, description="Absolute file path of Safetensors/GGUF to warm into RAM") filepath: Optional[str] = Field(None, description="Absolute file path of Safetensors/GGUF to warm into RAM")
blob_only: bool = Field(False, description="Warm the model's weights into page cache without loading VRAM") blob_only: bool = Field(False, description="Warm the model's weights into page cache without loading VRAM")
force: bool = Field(False, description="Warm even if residency sampling thinks it is already resident")
class WarmAllRequest(BaseModel): class WarmAllRequest(BaseModel):
budget_gb: Optional[float] = Field(None, description="Byte budget for warming; defaults to 70% of MemAvailable", example=24.0) budget_gb: Optional[float] = Field(None, description="Byte budget for warming; defaults to 70% of MemAvailable", example=24.0)
@@ -362,11 +363,11 @@ async def api_warm_plan(budget_gb: Optional[float] = Query(None, description="Ov
async def api_warm_model(req: WarmRequest): async def api_warm_model(req: WarmRequest):
"""Pre-warm a specific Ollama model or file path into the Linux page cache.""" """Pre-warm a specific Ollama model or file path into the Linux page cache."""
if req.model_name and req.blob_only: if req.model_name and req.blob_only:
return ram_optimizer.warm_ollama_blob(req.model_name) return ram_optimizer.warm_ollama_blob(req.model_name, force=req.force)
if req.model_name: if req.model_name:
return await ram_optimizer.warm_ollama_model(req.model_name, keep_alive="1m") return await ram_optimizer.warm_ollama_model(req.model_name, keep_alive="1m")
if req.filepath: if req.filepath:
return ram_optimizer.warm_file_to_ram(req.filepath) return ram_optimizer.warm_file_to_ram(req.filepath, force=req.force)
raise HTTPException(status_code=400, detail="model_name or filepath required") raise HTTPException(status_code=400, detail="model_name or filepath required")
@app.get("/api/cache/report", summary="Measured Page-Cache Residency", tags=["Memory Optimization"]) @app.get("/api/cache/report", summary="Measured Page-Cache Residency", tags=["Memory Optimization"])

View File

@@ -29,10 +29,20 @@ COMFY_API_BASE = "http://127.0.0.1:8188"
# Circular buffer for transition events (the durable log lives in telemetry_store) # Circular buffer for transition events (the durable log lives in telemetry_store)
SWITCH_HISTORY = deque(maxlen=50) SWITCH_HISTORY = deque(maxlen=50)
# Bandwidth thresholds used to classify how a model actually got into VRAM. # Bandwidth thresholds for classifying how a model reached VRAM, calibrated by measuring
# PCIe 4.0 x16 tops out near 31.5 GB/s; this NVMe sustains well under 2 GB/s. # the same 12.87 GB model loaded cold and warm on this box (2026-08-28):
RAM_HIT_GBPS = 5.0 #
PARTIAL_HIT_GBPS = 1.5 # 3.1% resident -> 34.3 s -> 0.38 GB/s
# 100% resident -> 4.9 s -> 2.63 GB/s
#
# The first cut at these numbers assumed a page-cache-fed load would approach the bus
# rate and set the cache-hit bar at 5 GB/s. It does not: Ollama's load_duration covers
# host-to-device transfer and model initialisation as well as the file read, so a fully
# resident model still reports ~2.6 GB/s while the page cache itself reads at 6.4 GB/s.
# A 5 GB/s bar could therefore never be met, and every warm load was being reported as
# a partial hit. Thresholds now sit either side of the measured 6.9x separation.
RAM_HIT_GBPS = 2.0
PARTIAL_HIT_GBPS = 0.8
# How long Ollama's VRAM may take to actually drain before we stop waiting. # How long Ollama's VRAM may take to actually drain before we stop waiting.
# Ollama will not unload a model while a generation is in flight, so a short ceiling # Ollama will not unload a model while a generation is in flight, so a short ceiling
@@ -557,11 +567,14 @@ def classify_load(size_bytes: int, load_duration_ms: float) -> Dict[str, Any]:
return {"cache_status": "Already in VRAM", "load_gbps": None, "is_ram_hit": True} return {"cache_status": "Already in VRAM", "load_gbps": None, "is_ram_hit": True}
if not size_bytes: if not size_bytes:
# No size on record — fall back to the old heuristic, but say so. # No size on record — fall back to the old heuristic, but say so.
# Without a size we cannot compute bandwidth at all; this is a guess and is
# labelled as one. 8s roughly splits the measured warm (4.9s) and cold (34.3s)
# loads for a mid-size model, but it is meaningless for very small or large ones.
return { return {
"cache_status": "RAM Cache Hit ⚡" if load_duration_ms < 2500 else "Cold Disk Load 💾", "cache_status": "RAM Cache Hit ⚡" if load_duration_ms < 8000 else "Cold Disk Load 💾",
"load_gbps": None, "load_gbps": None,
"is_ram_hit": load_duration_ms < 2500, "is_ram_hit": load_duration_ms < 8000,
"detail": "size unknown, fell back to duration heuristic", "detail": "size unknown, fell back to a duration guess",
} }
gbps = (size_bytes / (1024**3)) / (load_duration_ms / 1000.0) gbps = (size_bytes / (1024**3)) / (load_duration_ms / 1000.0)
if gbps >= RAM_HIT_GBPS: if gbps >= RAM_HIT_GBPS: