Recalibrate cache-hit thresholds against measured loads; bring MCP to parity
Calibration. The same 12.87GB model loaded through Ollama on this box: 3.1% resident (FADV_DONTNEED) -> 34.3s -> 0.38 GB/s 100% resident (force-warmed) -> 4.9s -> 2.63 GB/s The thresholds had been guessed from PCIe bus bandwidth: cache hit at >=5 GB/s. A fully warm load only reaches 2.63 GB/s, because load_duration covers host-to-device transfer and model init as well as the file read -- the page cache itself reads at 6.4 GB/s. The 5 GB/s bar was therefore unreachable, and every warm load was being reported as a partial hit. Now 2.0 / 0.8 GB/s, either side of the measured 6.9x separation. Warm-skip was also unsafe. A 12.87GB blob was skipped as already resident on the strength of twelve 2MB probe windows, then loaded at 2.44 GB/s. Skipping now requires warm_confident: an exact cachestat reading, or a probe finding every one of 32 denser samples resident. warm_file_to_ram/warm_ollama_blob take force=True, exposed on the warm-model endpoint, whose Pydantic model was missing the field entirely. MCP parity: the server had drifted well behind the REST API. Adds tools for measured residency, warm planning, VRAM requests, per-profile analytics, thermal governor control, overclock status/apply/restore, and autotune sweeps plus status -- 23 tools and 6 resources, up from 12 and 3. The telemetry store now starts in __main__ rather than at import scope, since server.py imports this module for the benchmark tool. README: replaced the remaining theoretical claims (31.5 GB/s bus rate, sub-1.5s loads, 15ms yields) with the measured numbers. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
41
README.md
41
README.md
@@ -107,8 +107,8 @@ only if the card actually needs it.
|
||||
* **Server-Sent Events (SSE)**: A single background sampler produces one 1 Hz snapshot and fans it out to every subscriber via `GET /api/stream`. Previously each connected client independently re-ran the whole snapshot — NVML, `/proc/meminfo`, an HTTP round-trip each to Ollama and ComfyUI, and a recursive walk of the ComfyUI models tree with a `stat()` per checkpoint — once per second, so opening the dashboard in three tabs tripled the load on the thing it was measuring. Slow clients drop stale frames instead of stalling the sampler.
|
||||
|
||||
### 🤖 Model Context Protocol (MCP 2.0) Server
|
||||
* **12 Native Agentic Tools**: Allows AI agents (Antigravity CLI, Claude Desktop, Cursor) to manage GPU resources, trigger model hot-swaps, tune fan curves, and inspect telemetry.
|
||||
* **3 Live MCP Resources**: Exposes live metrics, model catalogs, and switch logs as streamable resources (`gpu://metrics/live`, `gpu://models/catalog`, `gpu://history/switches`).
|
||||
* **23 Native Agentic Tools**: Allows AI agents (Antigravity CLI, Claude Desktop, Cursor) to manage GPU resources, trigger model hot-swaps, measure page-cache residency, read persisted performance analytics, drive the thermal governor, and run overclock sweeps.
|
||||
* **6 Live MCP Resources**: Live metrics, model catalog, switch log, measured cache residency, per-profile analytics, and the overclock profiles with the evidence behind each setting.
|
||||
* **Dual Transport Support**: Run via standard input/output (`--stdio`) or network Server-Sent Events (`--sse --port 8001`).
|
||||
|
||||
### ⏱️ Automated Latency & Throughput Benchmark Engine
|
||||
@@ -135,19 +135,32 @@ flowchart TD
|
||||
REST["REST API & OpenAPI Docs"]
|
||||
MCP["Model Context Protocol (MCP 2.0)"]
|
||||
SSE["1Hz Real-Time SSE Stream"]
|
||||
Arbitrator["VRAM Arbitrator (15ms Soft-Yield)"]
|
||||
Arbitrator["VRAM Arbitrator (confirmed yield)"]
|
||||
Overclock["Overclock & Fan Manager"]
|
||||
Warmer["Page Cache Pre-Warmer"]
|
||||
end
|
||||
|
||||
HostRAM <== "PCIe 4.0 x16 Bus (~31.5 GB/s Hot-Swap)" ==> GPU
|
||||
HostRAM <== "PCIe 4.0 x16 Bus (measured 2.6 GB/s warm model load)" ==> GPU
|
||||
Orchestrator --> GPU
|
||||
Orchestrator --> HostRAM
|
||||
```
|
||||
|
||||
### The Physics of Sub-Second Switching
|
||||
* **Host RAM as Staging**: Active LLMs and diffusion checkpoints remain resident in the 64GB Linux Page Cache.
|
||||
* **PCIe 4.0 x16 Hot-Swapping**: Transferring weights across PCIe 4.0 x16 achieves **~31.5 GB/s** bandwidth, reducing model loads from 30+ seconds (disk) to **under 1.5 seconds**.
|
||||
* **Warm vs cold model loads, measured.** The same 12.87 GB model, loaded through Ollama on this box:
|
||||
|
||||
| Page-cache residency | Load time | Effective rate |
|
||||
| :--- | :--- | :--- |
|
||||
| 3.1% (dropped with `FADV_DONTNEED`) | 34.3 s | 0.38 GB/s |
|
||||
| 100% (force-warmed) | 4.9 s | 2.63 GB/s |
|
||||
|
||||
A **6.9× speedup**, and the reason the page cache matters. Note the effective rate is
|
||||
well below the PCIe 4.0 x16 bus rate and below the 6.4 GB/s the page cache itself
|
||||
reads at: Ollama's `load_duration` also covers host-to-device transfer and model
|
||||
initialisation, not just the file read. Classification thresholds are calibrated
|
||||
against these measured numbers rather than the theoretical bus bandwidth — an earlier
|
||||
5 GB/s cache-hit bar sat above what a fully warm load can even achieve, so every warm
|
||||
load was misreported as a partial hit.
|
||||
* **Soft-Yielding**: Dropping Ollama's VRAM allocation via `keep_alive: 0` preserves the weights in host RAM. Measured on this box: the HTTP request returns in **~63 ms**, and the driver finishes releasing 14.9 GB **~77 ms after that**. HyperSwap waits for the second number before handing VRAM to ComfyUI — the earlier "~15 ms" figure timed the request, not the release.
|
||||
|
||||
---
|
||||
@@ -220,19 +233,33 @@ HyperSwap includes a native **MCP 2.0 server** (`mcp_server.py`) exposing orches
|
||||
| **`set_gpu_fan_speed`** | `mode` (str), `percent` (optional int) | Sets fan speed mode (`auto`\|`manual`) and target PWM % (30–100%). |
|
||||
| **`get_host_memory_status`** | *None* | 64GB host RAM breakdown, active page cache size, and cache ratio. |
|
||||
| **`switch_ollama_model`** | `model_name` (str), `keep_alive` (str) | Hot-swaps active LLM in VRAM, measures latency (ms) and tokens/sec. |
|
||||
| **`soft_yield_ollama_vram`** | `model_name` (optional str) | Yields Ollama VRAM to 0 MB in ~15ms while keeping model weights in RAM cache. |
|
||||
| **`soft_yield_ollama_vram`** | `model_name` (optional str) | Yields Ollama VRAM to 0 MB and waits for NVML to confirm the driver actually released it. Returns the request/confirm split. |
|
||||
| **`purge_comfyui_vram`** | *None* | Purges loaded diffusion models from ComfyUI pipeline VRAM. |
|
||||
| **`prewarm_all_models_to_ram`** | *None* | Faults all local LLM and diffusion checkpoints into Linux OS page cache. |
|
||||
| **`prewarm_all_models_to_ram`** | *None* | Warms the highest-value models into page cache within a byte budget, skipping what is already resident. |
|
||||
| **`prewarm_single_model`** | `model_name` (optional str), `filepath` (optional str) | Pre-warms a single GGUF or Safetensors file into RAM. |
|
||||
| **`list_available_models`** | *None* | Lists all installed Ollama models and discovered ComfyUI Safetensors on disk. |
|
||||
| **`get_switch_history`** | `limit` (int, default 20) | Retrieves recent switch events, millisecond latencies, and RAM hit status. |
|
||||
| **`run_model_switch_benchmark`**| `iterations` (int, default 2) | Automated round-trip latency benchmark between installed models. |
|
||||
| **`get_page_cache_residency`** | `include_files` (bool) | Measured page-cache residency per model file, with the measurement method used for each. |
|
||||
| **`get_warm_plan`** | `budget_gb` (optional float) | Previews what warming would read and skip, ranked by recency/frequency. Does not warm. |
|
||||
| **`request_vram_for_ollama`** | `needed_gb` (float) | Purges ComfyUI's checkpoints immediately if VRAM headroom is short, bypassing the idle timer. |
|
||||
| **`get_profile_performance`** | `days` (float, default 7) | Measured tok/s and thermals per overclock profile, from persisted history. |
|
||||
| **`get_thermal_governor_status`** | *None* | Current derate level, the reason for it, and escalation history. |
|
||||
| **`set_thermal_governor`** | `enabled` (optional bool), `reset` (bool) | Enable/disable the governor, or clear an active derate. |
|
||||
| **`get_overclock_status`** | *None* | Active profile, all profiles with their evidence, and which levers this driver honours. |
|
||||
| **`apply_overclock_profile`** | `profile` (str) | Apply `ollama` \| `comfy` \| `balanced`. |
|
||||
| **`restore_stock_gpu_state`** | *None* | Drop clock locks and offsets, restore default power limit, return fans to automatic. |
|
||||
| **`run_overclock_sweep`** | `knob`, `profile`, `workload`, `start`, `stop`, `repeats`, `apply_best` | Sweep a knob against a real workload and report the fastest stable value. Verifies the knob moves the hardware first. Takes minutes. |
|
||||
| **`get_autotune_status`** | *None* | Sweep progress, the last result table, and all recorded autotune steps. |
|
||||
|
||||
### MCP Resources List
|
||||
|
||||
* `gpu://metrics/live`: Real-time snapshot of GPU sensors and RAM page cache.
|
||||
* `gpu://models/catalog`: Catalog of all discovered GGUF and Safetensors models.
|
||||
* `gpu://history/switches`: Event log of recent model transitions and swap speeds.
|
||||
* `gpu://cache/residency`: Measured page-cache residency across every model on disk.
|
||||
* `gpu://analytics/profiles`: Measured throughput and thermals per overclock profile.
|
||||
* `gpu://overclock/profiles`: Overclock profiles including the measurement behind each setting.
|
||||
|
||||
---
|
||||
|
||||
|
||||
Reference in New Issue
Block a user