diff --git a/.gitignore b/.gitignore index bee8803..aa003ec 100644 --- a/.gitignore +++ b/.gitignore @@ -5,3 +5,8 @@ __pycache__/ .venv/ venv/ .DS_Store + +# persistent telemetry store +hyperswap.db +hyperswap.db-wal +hyperswap.db-shm diff --git a/README.md b/README.md index 02390db..67df288 100644 --- a/README.md +++ b/README.md @@ -20,17 +20,18 @@ ## 1. Feature Matrix ### โก Bidirectional VRAM Hot-Swapping & Arbitration -* **Sub-25ms Soft-Yield**: Instantly releases Ollama VRAM allocations (`keep_alive: 0`) down to 0 MB when ComfyUI needs to run diffusion workloads without evicting weights from system RAM. -* **Auto-Purge for ComfyUI**: Automatically purges diffusion pipeline checkpoints and VRAM buffers when an image/video generation job finishes, releasing 100% of VRAM back to Ollama. -* **Real-Time ComfyUI WebSocket & Watchdog Listener**: Subscribes directly to `ws://127.0.0.1:8188/ws` and runs a 300ms watchdog loop to detect prompt queueing and node execution in real time. +* **Confirmed Soft-Yield (barrier, not fire-and-forget)**: Releases Ollama VRAM allocations (`keep_alive: 0`) down to 0 MB, then **waits on NVML until the driver has actually freed the allocation** before letting ComfyUI proceed. Posting `keep_alive: 0` only *asks* Ollama to unload; on this box the HTTP call returns in ~63 ms while the driver takes a further ~77 ms to release 14.9 GB. Returning during that window is how diffusion ends up allocating into VRAM that is still occupied. +* **Idle-Aware ComfyUI Purge**: Diffusion checkpoints are held for `COMFY_IDLE_PURGE_S` (30 s) of genuinely empty queue rather than purged 1.5 s after every prompt โ iterating on a workflow no longer pays a full checkpoint reload per run. An immediate purge still happens the moment Ollama actually asks for VRAM (`POST /api/request-vram`). +* **Real-Time ComfyUI WebSocket & Watchdog Listener**: Subscribes directly to `ws://127.0.0.1:8188/ws`. The WebSocket is the primary signal; a connection-pooled watchdog polls `/queue` at 1 Hz purely as a fallback, backing off to 3 s while the socket is healthy. * **Process-Level VRAM Attribution**: Live NVML process inspection attributes exact GPU memory usage across Ollama (`llama-server`), ComfyUI (`python`), and Desktop display servers (`gnome-shell`, `Xorg`). -* **Hot-Swap Transition History**: Circular buffer logs all model switch events, swap durations (in ms), tokens/sec throughput, and RAM cache hit status (`RAM Cache Hit โก` vs `Cold Disk Load ๐พ`). +* **Bandwidth-Classified Transition History**: Every switch is classified by the bandwidth it actually achieved (`model size รท load duration`) rather than a fixed duration threshold: `RAM Cache Hit โก` (โฅ5 GB/s), `Partial Cache ๐ค` (โฅ1.5 GB/s), `Cold Disk Load ๐พ` (below that). The previous `load_duration < 2500 ms` rule called a 12.9 GB model read at 2.9 GB/s a "cold disk load" and a 0.5 GB model read from NVMe a "cache hit". ### ๐ง 64GB Host RAM Cache & Page Pre-warmer * **Zero-Latency Model Discovery**: Automatic cataloging of all local Ollama models (`/usr/share/ollama/.ollama/models`, `~/.ollama/models`) and ComfyUI model directories (`checkpoints`, `diffusion_models`, `unet`, `vae`, `clip`, `loras`, `controlnet`). -* **POSIX `fadvise` & Pinned Pre-warmer**: Pre-faults multi-gigabyte GGUFs and Safetensors into the Linux OS Page Cache so that reloading models across PCIe 4.0 x16 runs at ~31.5 GB/s (sub-second VRAM loads). -* **Granular Pre-warming Controls**: Pre-warm all discovered models in bulk or target individual models/safetensors on demand. -* **Memory Telemetry**: Real-time breakdown of Total Host RAM, Applications Memory, Active Model Page Cache, Free Memory, and Cache Residency Ratio. +* **POSIX `fadvise` & Pinned Pre-warmer**: Pre-faults multi-gigabyte GGUFs and Safetensors into the Linux OS Page Cache so that reloading models runs at page-cache speed rather than disk speed. +* **Measured Residency via `cachestat(2)`**: Residency is measured, not assumed. `cachestat(2)` gives exact cached-page counts per file. Where the kernel refuses it โ it only permits introspection of files you own, and Ollama's blobs are owned by uid `ollama` โ HyperSwap falls back to a randomised read-rate probe and labels the result as such. Files it cannot measure are reported as unmeasurable rather than guessed at. +* **Budgeted, Ranked Warming**: This box has 64 GB of RAM and >270 GB of model files; reading everything simply evicts whatever was warmed first. Files are ranked by recency/frequency (from the persisted event log) and warmed until a byte budget is spent, skipping anything already resident. `GET /api/warm-plan` previews the decision without executing it. +* **Memory Telemetry**: Real-time breakdown of Total Host RAM, Applications Memory, Active Model Page Cache, Free Memory, and measured Cache Residency Ratio. ### ๐๏ธ Dynamic Overclocking & Thermal Management * **Workload-Aware Overclock Profiles**: @@ -44,11 +45,25 @@ * **Hardware Fan Control**: Switch between `auto` and `manual` PWM control (30%โ100%) with synchronized dual-fan actuation (`[fan:0]` and `[fan:1]`). * **Automated Lockstep Profile Switching**: AutoArbitrator automatically switches hardware profiles in lockstep with the active workload (`comfy` on generation start, `ollama` on completion). +### ๐ก๏ธ Thermal Governor (closed-loop de-escalation) +* Every overclock lever here is sticky: a profile locks clocks and pins the fans to a manual PWM, and nothing used to undo that. The governor watches the telemetry the sampler already collects (so it costs no extra NVML calls) and walks the overclock back through a four-step derate ladder when the card runs hot or reports a hardware throttle. +* **Hysteresis by design**: escalation needs 5 consecutive bad samples, recovery needs 30 consecutive good ones, with a 20 s cooldown between changes โ a single spike during a diffusion step will not cause profile thrash. +* **Guaranteed restore**: stock clocks, default power limit and automatic fans are restored by the server's shutdown hook *and* by a systemd `ExecStopPost=`, so a `SIGKILL` cannot leave the card with locked clocks and fans pinned at 100%. + +### ๐ฌ Overclock Autotune (`autotune.py`) +* Walks a clock offset upward, running a fixed decode benchmark at each step, and reports the **fastest stable** value with its measured gain over baseline. +* **Instability detection**: kernel `Xid`/`NVRM` messages via `journalctl -k`, benchmark failure, degenerate output, and a temperature ceiling. The sweep stops climbing the moment a step looks unstable. +* **Safety**: refuses to start while ComfyUI is executing, and restores the original profile in a `finally` block โ including on exception or cancellation. + +### ๐๏ธ Persistent Telemetry Store (`telemetry_store.py`) +* Swap history used to be an in-memory `deque(maxlen=50)` that evaporated on every restart. Telemetry and events now persist to SQLite (WAL, single writer thread, batched 1 Hz inserts, automatic retention pruning) at roughly **0.4 MB per hour**. +* This is what makes the app's central question answerable: **`GET /api/analytics/profiles` compares decode throughput per overclock profile**, joined against the thermals recorded while that profile was active. + ### ๐ Real-Time Web Telemetry Dashboard (`:9090`) * **Live Hardware Telemetry**: GPU utilization %, GPU temperature (ยฐC), power draw (W), fan speeds (%), and graphics/memory clock frequencies (MHz). * **Live Dual-Axis Time-Series Chart**: Real-time graphical visualization of VRAM usage (GB) and Host RAM Cache (GB) with zero frontend polling overhead. * **Interactive Control Center**: Trigger model hot-swaps, soft-yields, cache pre-warms, fan adjustments, and benchmarks directly from the web interface. -* **Server-Sent Events (SSE)**: Pushes unified 1Hz telemetry updates via `GET /api/stream`. +* **Server-Sent Events (SSE)**: A single background sampler produces one 1 Hz snapshot and fans it out to every subscriber via `GET /api/stream`. Previously each connected client independently re-ran the whole snapshot โ NVML, `/proc/meminfo`, an HTTP round-trip each to Ollama and ComfyUI, and a recursive walk of the ComfyUI models tree with a `stat()` per checkpoint โ once per second, so opening the dashboard in three tabs tripled the load on the thing it was measuring. Slow clients drop stale frames instead of stalling the sampler. ### ๐ค Model Context Protocol (MCP 2.0) Server * **12 Native Agentic Tools**: Allows AI agents (Antigravity CLI, Claude Desktop, Cursor) to manage GPU resources, trigger model hot-swaps, tune fan curves, and inspect telemetry. @@ -92,7 +107,7 @@ flowchart TD ### The Physics of Sub-Second Switching * **Host RAM as Staging**: Active LLMs and diffusion checkpoints remain resident in the 64GB Linux Page Cache. * **PCIe 4.0 x16 Hot-Swapping**: Transferring weights across PCIe 4.0 x16 achieves **~31.5 GB/s** bandwidth, reducing model loads from 30+ seconds (disk) to **under 1.5 seconds**. -* **Soft-Yielding**: Dropping Ollama's VRAM allocation via `keep_alive: 0` takes **~15ms** while preserving the weights in host RAM. +* **Soft-Yielding**: Dropping Ollama's VRAM allocation via `keep_alive: 0` preserves the weights in host RAM. Measured on this box: the HTTP request returns in **~63 ms**, and the driver finishes releasing 14.9 GB **~77 ms after that**. HyperSwap waits for the second number before handing VRAM to ComfyUI โ the earlier "~15 ms" figure timed the request, not the release. --- @@ -119,12 +134,36 @@ The HyperSwap server runs on port `9090` by default. Interactive OpenAPI/Swagger | Endpoint | Method | Description | | :--- | :--- | :--- | | `/api/switch-model` | `POST` | Hot-swaps the active Ollama LLM in VRAM and tracks transition timing. | -| `/api/free-vram` | `POST` | Instructs Ollama to soft-yield VRAM down to 0 MB in ~15ms while retaining RAM cache. | +| `/api/free-vram` | `POST` | Soft-yields Ollama VRAM to 0 MB and **waits for NVML to confirm the release** (`?confirm=false` to skip). Returns `request_ms`, `confirm_ms` and the GB actually freed. | | `/api/comfy-free` | `POST` | Instructs ComfyUI to purge loaded diffusion weights and VRAM cache. | -| `/api/warm-all` | `POST` | Pre-faults all installed Ollama models and ComfyUI Safetensors into the Linux page cache. | -| `/api/warm-model` | `POST` | Pre-warms a specific model or file into RAM. | +| `/api/request-vram` | `POST` | Ollama-priority path: purges ComfyUI immediately if there is not enough free VRAM. | +| `/api/warm-all` | `POST` | Warms the highest-value models into page cache within a byte budget (`budget_gb`). | +| `/api/warm-plan` | `GET` | Previews what warming would read, in what order, and what it would skip โ without doing it. | +| `/api/warm-model` | `POST` | Pre-warms a specific model or file into RAM (`blob_only` warms weights without touching VRAM). | +| `/api/cache/report` | `GET` | Measured page-cache residency per model file, with the measurement method used for each. | | `/api/benchmark` | `POST` | Runs an automated back-and-forth model swap benchmark and calculates average latency. | +### Analytics Endpoints (persisted) + +| Endpoint | Method | Description | +| :--- | :--- | :--- | +| `/api/analytics/profiles` | `GET` | **Decode throughput per overclock profile**, joined with the thermals recorded under it. | +| `/api/analytics/swaps` | `GET` | Aggregated swap/yield/purge latencies, cache-hit split, and per-model throughput. | +| `/api/analytics/timeseries` | `GET` | Downsampled telemetry history for charts that outlive a page refresh. | +| `/api/analytics/models` | `GET` | Recency/frequency model ranking used to prioritise the warm budget. | +| `/api/history?durable=true` | `GET` | Swap history from the persistent store rather than the in-memory ring. | +| `/api/db` | `GET` | Store location, row counts and how many hours of history are held. | + +### Governor & Autotune Endpoints + +| Endpoint | Method | Description | +| :--- | :--- | :--- | +| `/api/governor` | `GET` / `POST` | Current derate level and why; enable/disable, or clear an active derate. | +| `/api/overclock/restore` | `POST` | Drop all clock locks and offsets, restore default power limit and automatic fans. | +| `/api/autotune` | `GET` | Sweep progress, last result, and every recorded autotune step. | +| `/api/autotune/sweep` | `POST` | Walk a clock offset upward, measuring tok/s and watching for instability at each step. | +| `/api/autotune/cancel` | `POST` | Stop the current sweep after the step in flight; the profile is restored either way. | + --- ## 4. Model Context Protocol (MCP 2.0) Reference diff --git a/autotune.py b/autotune.py new file mode 100644 index 0000000..c819096 --- /dev/null +++ b/autotune.py @@ -0,0 +1,263 @@ +"""Closed-loop overclock autotuner. + +The profiles in this repo were hand-tuned and had already drifted apart from the defaults +in overclock_manager.py, with no record of which numbers were actually faster. This module +answers that empirically: it walks a clock offset upward, measures real decode throughput +at each step, watches for instability, and reports the highest setting that was both +stable and fastest. + +Safety properties: + * The original profile is always restored, including on exception or cancellation. + * A sweep refuses to start while ComfyUI is executing, so it cannot corrupt someone's + render by yanking clocks mid-graph. + * Every step is bounded by a temperature ceiling and checked for kernel Xid messages, + and the sweep stops climbing the moment a step looks unstable. +""" +import asyncio +import logging +import subprocess +import time +from typing import Any, Dict, List, Optional + +import overclock_manager +import telemetry_store +import vram_arbitrator + +logger = logging.getLogger("autotune") + +BENCH_PROMPT = ("Write a detailed technical explanation of how virtual memory paging " + "works in a modern operating system kernel.") +BENCH_TOKENS = 160 +SETTLE_S = 2.5 +TEMP_CEILING_C = 84.0 + +KNOBS = { + "mem_offset_mhz": {"default_start": 0, "default_stop": 1000, "default_step": 100}, + "core_offset_mhz": {"default_start": 0, "default_stop": 300, "default_step": 25}, +} + + +def _xid_since(since_ts: float) -> List[str]: + """Look for NVIDIA Xid errors in the kernel log โ the clearest instability signal.""" + try: + since = time.strftime("%Y-%m-%d %H:%M:%S", time.localtime(since_ts)) + proc = subprocess.run( + ["journalctl", "-k", "--since", since, "--no-pager", "-q"], + capture_output=True, text=True, timeout=10, + ) + return [ln.strip() for ln in proc.stdout.splitlines() + if "Xid" in ln or "NVRM:" in ln] + except Exception as e: + logger.debug(f"Xid check unavailable: {e}") + return [] + + +async def _decode_benchmark(model: str) -> Dict[str, Any]: + """One fixed decode run. Throughput here is the thing being optimised.""" + client = vram_arbitrator._client(vram_arbitrator.OLLAMA_API_BASE, 300.0) + t0 = time.perf_counter() + resp = await client.post("/api/generate", json={ + "model": model, + "prompt": BENCH_PROMPT, + "stream": False, + "keep_alive": "10m", + "options": {"num_predict": BENCH_TOKENS, "temperature": 0.0, "seed": 42}, + }) + wall_ms = round((time.perf_counter() - t0) * 1000, 2) + if resp.status_code != 200: + return {"ok": False, "error": f"HTTP {resp.status_code}: {resp.text[:200]}", + "wall_ms": wall_ms} + data = resp.json() + eval_ms = data.get("eval_duration", 0) / 1e6 + eval_count = data.get("eval_count", 0) + text = data.get("response", "") or "" + return { + "ok": True, + "tokens_per_sec": round(eval_count / (eval_ms / 1000), 2) if eval_ms > 0 else 0.0, + "eval_count": eval_count, + "eval_ms": round(eval_ms, 2), + "prompt_eval_ms": round(data.get("prompt_eval_duration", 0) / 1e6, 2), + "wall_ms": wall_ms, + "response_chars": len(text), + # A model producing almost nothing, or pure repetition, is a corruption signal. + "degenerate": eval_count < BENCH_TOKENS * 0.5 or len(set(text.split())) < 8, + } + + +class SweepState: + def __init__(self) -> None: + self.running = False + self.cancel = False + self.current: Optional[Dict[str, Any]] = None + self.last_result: Optional[Dict[str, Any]] = None + + +state = SweepState() + + +async def sweep(knob: str = "mem_offset_mhz", + profile: str = "ollama", + model: Optional[str] = None, + start: Optional[int] = None, + stop: Optional[int] = None, + step: Optional[int] = None, + repeats: int = 1, + apply_best: bool = False) -> Dict[str, Any]: + """Sweep one clock offset and return the fastest stable value.""" + if knob not in KNOBS: + return {"success": False, "error": f"unknown knob '{knob}'; try {list(KNOBS)}"} + if state.running: + return {"success": False, "error": "a sweep is already running"} + + comfy = await vram_arbitrator.get_comfyui_live_state() + if comfy.get("executing") or comfy.get("queue_remaining"): + return {"success": False, "error": "ComfyUI is busy; refusing to change clocks mid-render"} + + if not model: + ollama = await vram_arbitrator.get_ollama_live_state() + model = ollama.get("active_model_name") + if not model: + installed = ollama.get("installed_models") or [] + if not installed: + return {"success": False, "error": "no Ollama model available to benchmark"} + model = installed[0].get("name") + + defaults = KNOBS[knob] + start = defaults["default_start"] if start is None else start + stop = defaults["default_stop"] if stop is None else stop + step = defaults["default_step"] if step is None else step + if step <= 0 or stop < start: + return {"success": False, "error": "invalid sweep range"} + + baseline_cfg = overclock_manager.load_profiles().get(profile, {}) + state.running = True + state.cancel = False + results: List[Dict[str, Any]] = [] + t_start = time.time() + + try: + # Load the model once up front so the first step does not pay the load cost. + await _decode_benchmark(model) + + value = start + while value <= stop and not state.cancel: + overclock_manager.apply_profile(profile, overrides={knob: value}) + await asyncio.sleep(SETTLE_S) + step_started = time.time() + + samples = [] + for _ in range(max(repeats, 1)): + samples.append(await _decode_benchmark(model)) + if state.cancel: + break + + gpu = vram_arbitrator.get_gpu_hardware_stats() + xids = _xid_since(step_started) + ok_samples = [s for s in samples if s.get("ok") and not s.get("degenerate")] + temp = gpu.get("temperature_c", 0) or 0 + + instability = [] + if xids: + instability.append(f"kernel Xid: {xids[0][:120]}") + if len(ok_samples) < len(samples): + instability.append("benchmark failed or produced degenerate output") + if temp >= TEMP_CEILING_C: + instability.append(f"temperature ceiling hit ({temp}ยฐC)") + + tok_s = round(max((s["tokens_per_sec"] for s in ok_samples), default=0.0), 2) + row = { + "knob": knob, + "value": value, + "profile": profile, + "model": model, + "tokens_per_sec": tok_s, + "temp_c": temp, + "power_w": gpu.get("power_w"), + "clock_sm_mhz": gpu.get("clock_graphics_mhz"), + "clock_mem_mhz": gpu.get("clock_mem_mhz"), + "throttle_reasons": gpu.get("throttle_reasons"), + "stable": not instability, + "instability": "; ".join(instability) or None, + "samples": samples, + } + results.append(row) + telemetry_store.record_autotune({ + "profile": profile, "knob": knob, + "core_offset_mhz": value if knob == "core_offset_mhz" else baseline_cfg.get("core_offset_mhz"), + "mem_offset_mhz": value if knob == "mem_offset_mhz" else baseline_cfg.get("mem_offset_mhz"), + "tokens_per_sec": tok_s, "temp_c": temp, "power_w": gpu.get("power_w"), + "stable": row["stable"], "instability": row["instability"], + "note": f"sweep {knob} {start}..{stop} step {step}", + }) + state.current = {"knob": knob, "value": value, "stop": stop, + "tokens_per_sec": tok_s, "stable": row["stable"]} + logger.info(f"autotune {knob}={value}: {tok_s} tok/s, {temp}ยฐC, " + f"stable={row['stable']} {row['instability'] or ''}") + + if not row["stable"]: + logger.warning(f"autotune stopping climb at {knob}={value}: {row['instability']}") + break + value += step + + stable = [r for r in results if r["stable"] and r["tokens_per_sec"] > 0] + best = max(stable, key=lambda r: r["tokens_per_sec"]) if stable else None + baseline = next((r for r in results if r["value"] == start), None) + gain_pct = None + if best and baseline and baseline["tokens_per_sec"] > 0: + gain_pct = round((best["tokens_per_sec"] / baseline["tokens_per_sec"] - 1) * 100, 2) + + applied = None + if apply_best and best: + overclock_manager.set_profile(profile, {knob: best["value"]}) + applied = {knob: best["value"], "profile": profile} + logger.info(f"autotune wrote {knob}={best['value']} into profile '{profile}'") + + result = { + "success": True, + "knob": knob, + "profile": profile, + "model": model, + "range": {"start": start, "stop": stop, "step": step}, + "steps_run": len(results), + "duration_s": round(time.time() - t_start, 1), + "cancelled": state.cancel, + "best": {k: best[k] for k in ("value", "tokens_per_sec", "temp_c", "clock_mem_mhz", + "clock_sm_mhz")} if best else None, + "baseline_tokens_per_sec": baseline["tokens_per_sec"] if baseline else None, + "gain_pct": gain_pct, + "applied_to_profile": applied, + "first_unstable": next(({"value": r["value"], "why": r["instability"]} + for r in results if not r["stable"]), None), + "table": [{k: r[k] for k in ("value", "tokens_per_sec", "temp_c", "power_w", + "clock_mem_mhz", "clock_sm_mhz", "stable", + "instability")} for r in results], + } + state.last_result = result + return result + finally: + # Always hand the card back exactly as we found it. + state.running = False + state.current = None + try: + overclock_manager.apply_profile(profile) + logger.info(f"autotune restored profile '{profile}'") + except Exception as e: + logger.error(f"autotune failed to restore profile, forcing stock: {e}") + overclock_manager.restore_safe("autotune restore failed") + + +def get_status() -> Dict[str, Any]: + return { + "running": state.running, + "current": state.current, + "last_result": state.last_result, + "knobs": KNOBS, + "history": telemetry_store.autotune_history(100), + } + + +def cancel() -> Dict[str, Any]: + if not state.running: + return {"cancelled": False, "reason": "no sweep running"} + state.cancel = True + return {"cancelled": True} diff --git a/overclock_manager.py b/overclock_manager.py index 5f5d3b0..1fb7368 100644 --- a/overclock_manager.py +++ b/overclock_manager.py @@ -207,14 +207,20 @@ def _apply_offsets(core_mhz: int, mem_mhz: int) -> Dict[str, Any]: } -def apply_profile(name: str) -> Dict[str, Any]: - """Apply a named overclock profile to the GPU. Returns a full result report.""" +def apply_profile(name: str, overrides: Optional[Dict[str, Any]] = None) -> Dict[str, Any]: + """Apply a named overclock profile to the GPU. Returns a full result report. + + `overrides` lets the thermal governor and the autotuner apply a modified version of a + profile (a derated offset, a probe clock) without mutating what is stored on disk. + """ global ACTIVE_PROFILE, _LAST_RESULT profiles = load_profiles() if name not in profiles: return {"success": False, "error": f"unknown profile '{name}'", "profile": name} - cfg = profiles[name] + cfg = dict(profiles[name]) + if overrides: + cfg.update(overrides) fan_mode = cfg.get("fan_mode", "auto") fan_speed = int(cfg.get("fan_speed_pct", 0)) @@ -230,15 +236,31 @@ def apply_profile(name: str) -> Dict[str, Any]: } result["gpu"] = get_gpu_state() result["fan_status"] = get_fan_status() + result["overrides"] = overrides or {} + _STATE_CACHE["value"] = None + _FAN_CACHE["value"] = None ACTIVE_PROFILE = name _LAST_RESULT = result logger.info(f"Overclock profile applied: {name} -> {json.dumps(result, default=str)}") return result -def get_gpu_state() -> Dict[str, Any]: - """Read back live GPU clocks/power/limits via nvidia-smi.""" +_STATE_CACHE: Dict[str, Any] = {"ts": 0.0, "value": None} +_FAN_CACHE: Dict[str, Any] = {"ts": 0.0, "value": None} +STATE_TTL_S = 2.0 + + +def get_gpu_state(force: bool = False) -> Dict[str, Any]: + """Read back live GPU clocks/power/limits via nvidia-smi. + + Cached for STATE_TTL_S: this forks `sudo nvidia-smi`, and the dashboard polls the + status endpoint every few seconds. NVML already covers the live 1 Hz telemetry. + """ + import time as _time + if not force and _STATE_CACHE["value"] is not None and \ + (_time.time() - _STATE_CACHE["ts"]) < STATE_TTL_S: + return _STATE_CACHE["value"] state: Dict[str, Any] = {} r = _smi( "--query-gpu=driver_version,name,memory.total,power.limit,power.max_limit,power.default_limit," @@ -257,6 +279,7 @@ def get_gpu_state() -> Dict[str, Any]: state[k] = float(parts[i]) except ValueError: state[k] = parts[i] + _STATE_CACHE.update({"ts": __import__("time").time(), "value": state}) return state @@ -298,11 +321,16 @@ def set_fan_auto() -> Dict[str, Any]: ok = r["rc"] == 0 if ok: FAN_MANUAL = False + _FAN_CACHE["value"] = None return {"success": ok, "manual": False, "fan_speed_pct": None, "detail": r.get("out") or r.get("err")} -def get_fan_status() -> Dict[str, Any]: - """Read current fan control mode + target speed.""" +def get_fan_status(force: bool = False) -> Dict[str, Any]: + """Read current fan control mode + target speed (cached; forks nvidia-settings).""" + import time as _time + if not force and _FAN_CACHE["value"] is not None and \ + (_time.time() - _FAN_CACHE["ts"]) < STATE_TTL_S: + return _FAN_CACHE["value"] global FAN_MANUAL target = None manual = FAN_MANUAL @@ -320,7 +348,34 @@ def get_fan_status() -> Dict[str, Any]: target = int(line.split("):")[-1].split(".")[0].strip()) except Exception: pass - return {"manual": manual, "mode": "manual" if manual else "auto", "target_speed_pct": target} + result = {"manual": manual, "mode": "manual" if manual else "auto", "target_speed_pct": target} + _FAN_CACHE.update({"ts": __import__("time").time(), "value": result}) + return result + + +def restore_safe(reason: str = "shutdown") -> Dict[str, Any]: + """Return the card to stock: no clock locks, no offsets, default power, automatic fans. + + This matters because every lever here is sticky. If the service dies while a profile is + applied, the GPU keeps the locked clocks and, worse, keeps the fans pinned at whatever + manual PWM was last set. Nothing was undoing that. + """ + logger.warning(f"Restoring GPU to safe stock state ({reason})") + result = { + "reason": reason, + "clock_lock": _apply_clock_lock(0, 0), + "mem_lock": _apply_mem_lock(0), + "offsets": _apply_offsets(0, 0), + "fan": set_fan_auto(), + } + # Hand the power limit back to the card's own default rather than assuming 370 W. + state = get_gpu_state() + default_w = state.get("power_default_w") + if isinstance(default_w, (int, float)) and default_w > 0: + result["power_limit"] = _apply_power_limit(int(default_w)) + global ACTIVE_PROFILE + ACTIVE_PROFILE = "stock" + return result def get_status() -> Dict[str, Any]: diff --git a/overclock_profiles.json b/overclock_profiles.json index 573a6b7..d871404 100644 --- a/overclock_profiles.json +++ b/overclock_profiles.json @@ -2,8 +2,8 @@ "ollama": { "label": "Ollama \u2014 LLM decode (memory-bandwidth bound)", "power_limit_w": 370, - "core_offset_mhz": 150, - "mem_offset_mhz": 825, + "core_offset_mhz": 35, + "mem_offset_mhz": 200, "lock_core_min": 0, "lock_core_max": 0, "lock_mem_mhz": 0, @@ -14,12 +14,12 @@ "label": "ComfyUI \u2014 diffusion (core-compute bound)", "power_limit_w": 370, "core_offset_mhz": 100, - "mem_offset_mhz": 500, + "mem_offset_mhz": 150, "lock_core_min": 2900, "lock_core_max": 3105, "lock_mem_mhz": 0, "fan_mode": "manual", - "fan_speed_pct": 75 + "fan_speed_pct": 100 }, "balanced": { "label": "Balanced \u2014 stock boost, power unlocked", diff --git a/ram_optimizer.py b/ram_optimizer.py index a8306ee..bf0bff2 100644 --- a/ram_optimizer.py +++ b/ram_optimizer.py @@ -1,16 +1,48 @@ -"""RAM Optimizer and Model Pre-warmer for High-Speed Switching.""" -import os -import glob -import time -import httpx +"""RAM Optimizer and Model Pre-warmer for High-Speed Switching. + +Two things changed here versus the naive version: + + 1. Residency is *measured*, not assumed. mincore(2) tells us exactly what fraction of + each model file is resident in the Linux page cache, so "RAM Cache Hit" stops being + a guess based on how long a load took. + 2. Warming is *budgeted*. This box has 64 GB of RAM and >33 GB of models; reading every + file top-to-bottom simply evicts whatever was warmed first. Files are now scored by + recency/frequency (from the telemetry store) and warmed until a byte budget is hit, + skipping anything already resident. +""" +import ctypes +import ctypes.util +import json import logging -from typing import Dict, List, Any +import os +import random +import time +from typing import Dict, List, Any, Optional, Tuple + +import httpx + +import telemetry_store logger = logging.getLogger("ram_optimizer") OLLAMA_API_BASE = "http://localhost:11434" COMFY_API_BASE = "http://127.0.0.1:8188" -COMFY_MODELS_DIR = "/home/drjones/ComfyUI/models" +COMFY_MODELS_DIR = os.environ.get("HYPERSWAP_COMFY_MODELS", "/home/drjones/ComfyUI/models") +OLLAMA_MODEL_DIRS = [ + "/usr/share/ollama/.ollama/models", + os.path.expanduser("~/.ollama/models"), +] + +PAGE_SIZE = os.sysconf("SC_PAGE_SIZE") +# Files bigger than this are sampled rather than fully mapped for residency. +RESIDENCY_FULL_MAP_LIMIT = 2 * 1024 ** 3 +RESIDENCY_SAMPLE_WINDOWS = 64 +RESIDENCY_WINDOW_BYTES = 16 * 1024 * 1024 +# A file at/above this residency is considered warm and is skipped by the warmer. +WARM_SKIP_THRESHOLD_PCT = 90.0 + +CATALOG_TTL_S = 30.0 + def get_detailed_meminfo() -> Dict[str, Any]: """Parse /proc/meminfo for precise page cache and RAM stats.""" @@ -25,7 +57,7 @@ def get_detailed_meminfo() -> Dict[str, Any]: info[key] = int(val) * 1024 # Convert kB to bytes except Exception as e: logger.error(f"Failed to read /proc/meminfo: {e}") - + total = info.get("MemTotal", 0) free = info.get("MemFree", 0) available = info.get("MemAvailable", 0) @@ -51,43 +83,439 @@ def get_detailed_meminfo() -> Dict[str, Any]: "cache_ratio_pct": round((cached / total * 100) if total > 0 else 0, 1), } -def warm_file_to_ram(filepath: str, chunk_size: int = 16 * 1024 * 1024) -> Dict[str, Any]: - """Pre-fault/read file into Linux OS Page Cache at maximum disk read speed.""" + +# ---------------------------------------------------------------- page residency +# +# Measuring page-cache residency turned out to be the subtle part. +# +# * cachestat(2) (Linux 6.5+) is the right tool: exact cached-page counts for an fd, +# no mmap, microseconds per call. But the kernel only permits it on files you own +# or can write -- the Ollama blobs are owned by uid `ollama`, so it returns EPERM. +# * mincore(2) does NOT fail closed for those files on this kernel: it reports every +# page as resident, which produced 128 GB of "resident" model weights on a box with +# 46 GB of page cache. It is therefore not used at all. +# +# So: cachestat where permitted, and an explicit read-throughput probe where it is not. +# Anything we cannot measure is reported as unmeasurable rather than guessed at. + +_libc = None +_SYS_cachestat = 451 # x86_64 + + +class _CachestatRange(ctypes.Structure): + _fields_ = [("off", ctypes.c_uint64), ("len", ctypes.c_uint64)] + + +class _Cachestat(ctypes.Structure): + _fields_ = [ + ("nr_cache", ctypes.c_uint64), + ("nr_dirty", ctypes.c_uint64), + ("nr_writeback", ctypes.c_uint64), + ("nr_evicted", ctypes.c_uint64), + ("nr_recently_evicted", ctypes.c_uint64), + ] + + +def _get_libc(): + global _libc + if _libc is None: + _libc = ctypes.CDLL(ctypes.util.find_library("c") or "libc.so.6", use_errno=True) + return _libc + + +def _cachestat(fd: int, offset: int, length: int) -> Optional[_Cachestat]: + """Raw cachestat(2). Returns None if the kernel refuses (EPERM/ENOSYS).""" + libc = _get_libc() + rng = _CachestatRange(offset, length) + cs = _Cachestat() + ctypes.set_errno(0) + rc = libc.syscall(ctypes.c_long(_SYS_cachestat), ctypes.c_int(fd), + ctypes.byref(rng), ctypes.byref(cs), ctypes.c_uint(0)) + if rc != 0: + return None + return cs + + +PROBE_WINDOWS = 12 +PROBE_WINDOW_BYTES = 2 * 1024 * 1024 +# Measured on this box: cold NVMe reads land around 0.35-0.5 GB/s, page-cache reads at +# 3.2-13 GB/s. 1.5 GB/s sits in the empty middle of that gap. +PROBE_CACHED_GBPS = 1.5 + + +def _throughput_probe(fd: int, size: int) -> Dict[str, Any]: + """Infer residency by timing reads of small windows spread across the file. + + Used only where cachestat is not permitted (Ollama's blobs are owned by uid `ollama`). + + Two details matter for correctness: + + * Offsets are random per call. A fixed stride made the probe self-fulfilling: the + first pass faulted its 24 MB of sample windows into the page cache, and every pass + after that re-read exactly those windows and reported 100% resident for a file that + was almost entirely cold. + * Windows that read cold are handed straight back with FADV_DONTNEED. Those pages are + pollution the probe itself created, and leaving them behind would slowly warm the + cache with data nobody asked for. + """ + windows = min(PROBE_WINDOWS, max(int(size // PROBE_WINDOW_BYTES), 1)) + if windows <= 0: + return {"resident_pct": 0.0, "windows": 0} + + max_off = max(size - PROBE_WINDOW_BYTES, 0) + offsets = sorted(random.randint(0, max_off) for _ in range(windows)) if max_off else [0] + buf = bytearray(PROBE_WINDOW_BYTES) + cached = 0 + rates = [] + for off in offsets: + length = min(PROBE_WINDOW_BYTES, size - off) + if length <= 0: + continue + view = memoryview(buf)[:length] + t0 = time.perf_counter() + os.preadv(fd, [view], off) + dt = time.perf_counter() - t0 + gbps = (length / (1024 ** 3)) / dt if dt > 0 else 0.0 + rates.append(gbps) + if gbps >= PROBE_CACHED_GBPS: + cached += 1 + else: + # We just pulled this off disk; put it back the way we found it. + try: + os.posix_fadvise(fd, off, length, os.POSIX_FADV_DONTNEED) + except Exception: + pass + n = len(rates) + return { + "resident_pct": round((cached / n * 100) if n else 0.0, 1), + "windows": n, + "median_gbps": round(sorted(rates)[n // 2], 2) if n else 0.0, + "sampled_gb": round(n * PROBE_WINDOW_BYTES / (1024 ** 3), 3), + } + + +def page_residency(filepath: str, allow_probe: bool = True) -> Dict[str, Any]: + """Measure what fraction of a file is resident in the Linux page cache.""" + try: + size = os.path.getsize(filepath) + except OSError as e: + return {"success": False, "error": str(e), "resident_pct": 0.0, "measurable": False} + if size == 0: + return {"success": True, "resident_pct": 0.0, "size_bytes": 0, "measurable": True, + "method": "empty"} + + try: + fd = os.open(filepath, os.O_RDONLY) + except OSError as e: + return {"success": False, "error": str(e), "resident_pct": 0.0, "measurable": False} + try: + cs = _cachestat(fd, 0, size) + if cs is not None: + total_pages = (size + PAGE_SIZE - 1) // PAGE_SIZE + pct = round((cs.nr_cache / total_pages * 100) if total_pages else 0.0, 1) + method, measurable = "cachestat", True + extra = {"dirty_pages": cs.nr_dirty, "evicted_pages": cs.nr_evicted} + elif allow_probe: + probe = _throughput_probe(fd, size) + pct = probe["resident_pct"] + method, measurable = "probe", True + extra = {"probe_windows": probe["windows"], "probe_median_gbps": probe.get("median_gbps")} + else: + return {"success": True, "filepath": filepath, "size_bytes": size, + "size_gb": round(size / (1024**3), 3), "resident_pct": None, + "measurable": False, "method": "unavailable", "warm": None, + "reason": "cachestat not permitted for this file (not owned by us)"} + + return { + "success": True, + "filepath": filepath, + "size_bytes": size, + "size_gb": round(size / (1024**3), 3), + "resident_pct": pct, + "resident_bytes": int(size * pct / 100.0), + "method": method, + "measurable": measurable, + "warm": pct >= WARM_SKIP_THRESHOLD_PCT, + **extra, + } + except Exception as e: + return {"success": False, "error": str(e), "resident_pct": 0.0, + "size_bytes": size, "measurable": False} + finally: + os.close(fd) + + +def residency_capability() -> Dict[str, Any]: + """Report whether exact residency is available, and how to enable it if not.""" + catalog = get_model_catalog() + blocked = [] + for f in catalog["ollama"]: + try: + fd = os.open(f["full_path"], os.O_RDONLY) + except OSError: + continue + try: + if _cachestat(fd, 0, 4096) is None: + blocked.append(f["full_path"]) + finally: + os.close(fd) + break # one probe is enough; blobs share a directory and owner + if not blocked: + return {"exact_everywhere": True} + owner = "" + try: + import pwd + owner = pwd.getpwuid(os.stat(blocked[0]).st_uid).pw_name + except Exception: + owner = str(os.stat(blocked[0]).st_uid) + return { + "exact_everywhere": False, + "method_for_blocked": "probe", + "reason": f"cachestat(2) is only permitted on files you own or can write; " + f"Ollama blobs are owned by '{owner}'", + "hint": f"exact numbers for Ollama weights need read/write access, e.g. " + f"'sudo usermod -aG {owner} $USER' plus group-write on the blobs directory", + } + + +# ---------------------------------------------------------------- catalogs + +_catalog_cache: Dict[str, Any] = {"ts": 0.0, "sig": None, "comfy": [], "ollama": []} + + +def _dir_signature(root: str) -> Tuple: + """Cheap fingerprint of a model tree: (mtime, entry count) per subdirectory.""" + sig = [] + if not os.path.isdir(root): + return tuple(sig) + for dirpath, dirnames, filenames in os.walk(root): + try: + sig.append((dirpath, os.stat(dirpath).st_mtime_ns, len(filenames))) + except OSError: + continue + return tuple(sig) + + +def find_ollama_model_files() -> List[Dict[str, Any]]: + """Map installed Ollama models to their on-disk GGUF blobs via the manifest tree. + + Knowing the blob path is what lets us warm (or measure) a specific model's weights + without pulling them into VRAM. + """ + results: List[Dict[str, Any]] = [] + seen = set() + for root in OLLAMA_MODEL_DIRS: + manifests = os.path.join(root, "manifests") + blobs = os.path.join(root, "blobs") + if not os.path.isdir(manifests): + continue + for dirpath, _, filenames in os.walk(manifests): + for tag in filenames: + manifest_path = os.path.join(dirpath, tag) + try: + with open(manifest_path) as f: + manifest = json.load(f) + except Exception: + continue + rel = os.path.relpath(dirpath, manifests) + parts = rel.split(os.sep) + # registry/namespace/name -> "name:tag", keeping non-library namespaces + name = parts[-1] if parts else rel + namespace = parts[-2] if len(parts) >= 2 else "library" + model_name = f"{name}:{tag}" if namespace == "library" else f"{namespace}/{name}:{tag}" + for layer in manifest.get("layers", []): + if layer.get("mediaType") != "application/vnd.ollama.image.model": + continue + digest = (layer.get("digest") or "").replace(":", "-") + blob_path = os.path.join(blobs, digest) + if not os.path.exists(blob_path): + continue + key = (model_name, blob_path) + if key in seen: + continue + seen.add(key) + size = layer.get("size") or os.path.getsize(blob_path) + results.append({ + "model": model_name, + "filename": digest, + "full_path": blob_path, + "size_bytes": size, + "size_gb": round(size / (1024**3), 3), + "kind": "ollama", + }) + return results + + +def find_comfy_model_files(force_refresh: bool = False) -> List[Dict[str, Any]]: + """Discover all model files under ComfyUI models (cached). + + This used to run inside the 1Hz telemetry snapshot, meaning a full recursive walk plus + a stat() of every checkpoint once per second per connected dashboard. It is now cached + behind a directory-mtime fingerprint. + """ + _refresh_catalog(force_refresh) + return _catalog_cache["comfy"] + + +def get_model_catalog(force_refresh: bool = False) -> Dict[str, Any]: + _refresh_catalog(force_refresh) + return { + "comfy": _catalog_cache["comfy"], + "ollama": _catalog_cache["ollama"], + "cached_at": _catalog_cache["ts"], + } + + +def _refresh_catalog(force: bool = False) -> None: + now = time.time() + if not force and (now - _catalog_cache["ts"]) < CATALOG_TTL_S: + return + sig = _dir_signature(COMFY_MODELS_DIR) + if not force and sig == _catalog_cache["sig"] and _catalog_cache["comfy"]: + _catalog_cache["ts"] = now + return + + extensions = (".safetensors", ".ckpt", ".pt", ".bin", ".gguf", ".sft") + results = [] + if os.path.exists(COMFY_MODELS_DIR): + for root, _, files in os.walk(COMFY_MODELS_DIR): + for file in files: + if not file.endswith(extensions): + continue + full_path = os.path.join(root, file) + try: + st = os.stat(full_path) + except OSError: + continue + results.append({ + "filename": file, + "rel_path": os.path.relpath(full_path, COMFY_MODELS_DIR), + "full_path": full_path, + "category": os.path.relpath(root, COMFY_MODELS_DIR).split(os.sep)[0], + "size_bytes": st.st_size, + "size_mb": round(st.st_size / (1024**2), 2), + "size_gb": round(st.st_size / (1024**3), 3), + "mtime": st.st_mtime, + "kind": "comfy", + }) + _catalog_cache.update({"ts": now, "sig": sig, "comfy": results, + "ollama": find_ollama_model_files()}) + + +# ---------------------------------------------------------------- residency report + +_report_cache: Dict[str, Any] = {"ts": 0.0, "report": None} +REPORT_TTL_S = 15.0 + + +def get_cache_report(include_files: bool = True, force_refresh: bool = False) -> Dict[str, Any]: + """Measured page-cache residency across the whole model catalog. + + Deduplicated by blob path: several Ollama tags routinely point at the same GGUF, and + counting each tag separately produced more "resident" bytes than the box has RAM. + """ + now = time.time() + cached = _report_cache["report"] + if cached and not force_refresh and (now - _report_cache["ts"]) < REPORT_TTL_S: + return cached if include_files else {**cached, "files": []} + + t0 = time.perf_counter() + catalog = get_model_catalog() + by_path: Dict[str, Dict[str, Any]] = {} + for f in list(catalog["ollama"]) + list(catalog["comfy"]): + path = f["full_path"] + name = f.get("model") or f.get("rel_path") or f.get("filename") + if path in by_path: + by_path[path]["aliases"].append(name) + continue + by_path[path] = {"entry": f, "name": name, "aliases": []} + + entries = [] + total_bytes = resident_bytes = 0 + for path, meta in by_path.items(): + f = meta["entry"] + res = page_residency(path) + size = f.get("size_bytes") or res.get("size_bytes") or 0 + rb = res.get("resident_bytes", 0) + total_bytes += size + resident_bytes += rb + entries.append({ + "name": meta["name"], + "aliases": meta["aliases"], + "kind": f.get("kind"), + "full_path": path, + "size_gb": round(size / (1024**3), 3), + "resident_pct": res.get("resident_pct", 0.0), + "resident_gb": round(rb / (1024**3), 3), + "warm": res.get("warm", False), + }) + entries.sort(key=lambda e: e["resident_gb"], reverse=True) + + report = { + "scan_ms": round((time.perf_counter() - t0) * 1000, 1), + "files_scanned": len(entries), + "unique_blobs": len(by_path), + "catalog_total_gb": round(total_bytes / (1024**3), 2), + "resident_total_gb": round(resident_bytes / (1024**3), 2), + "residency_pct": round((resident_bytes / total_bytes * 100) if total_bytes else 0, 1), + "warm_files": sum(1 for e in entries if e["warm"]), + "files": entries, + } + _report_cache.update({"ts": now, "report": report}) + return report if include_files else {**report, "files": []} + + +# ---------------------------------------------------------------- warming + +def warm_file_to_ram(filepath: str, chunk_size: int = 16 * 1024 * 1024, + skip_if_warm: bool = True) -> Dict[str, Any]: + """Pre-fault a file into the Linux page cache, skipping it if already resident.""" if not os.path.exists(filepath): return {"success": False, "error": f"File not found: {filepath}", "duration_ms": 0} - + + before = page_residency(filepath) + if skip_if_warm and before.get("warm"): + return { + "success": True, "filepath": filepath, "skipped": True, + "reason": "already resident", "resident_pct": before.get("resident_pct"), + "size_mb": round(before.get("size_bytes", 0) / (1024**2), 2), + "duration_ms": 0.0, "bytes_read": 0, + } + t0 = time.perf_counter() file_size = os.path.getsize(filepath) bytes_read = 0 try: with open(filepath, "rb") as f: - # Hint kernel that we will read this sequentially try: os.posix_fadvise(f.fileno(), 0, file_size, os.POSIX_FADV_WILLNEED) except Exception: pass - buf = bytearray(chunk_size) while True: n = f.readinto(buf) if not n: break bytes_read += n - + duration = time.perf_counter() - t0 - duration_ms = round(duration * 1000, 2) - speed_mb_s = round((bytes_read / (1024**2)) / duration if duration > 0 else 0, 2) + after = page_residency(filepath) return { "success": True, "filepath": filepath, + "skipped": False, "size_bytes": file_size, "size_mb": round(file_size / (1024**2), 2), "bytes_read": bytes_read, - "duration_ms": duration_ms, - "speed_mb_s": speed_mb_s, + "duration_ms": round(duration * 1000, 2), + "speed_mb_s": round((bytes_read / (1024**2)) / duration if duration > 0 else 0, 2), + "resident_pct_before": before.get("resident_pct", 0.0), + "resident_pct_after": after.get("resident_pct", 0.0), } except Exception as e: - return {"success": False, "error": str(e), "duration_ms": round((time.perf_counter() - t0) * 1000, 2)} + return {"success": False, "error": str(e), + "duration_ms": round((time.perf_counter() - t0) * 1000, 2)} + async def warm_ollama_model(model_name: str, keep_alive: str = "5m") -> Dict[str, Any]: """Warm an Ollama model into memory and measure time.""" @@ -101,76 +529,148 @@ async def warm_ollama_model(model_name: str, keep_alive: str = "5m") -> Dict[str duration = time.perf_counter() - t0 if resp.status_code == 200: data = resp.json() - return { + res = { "success": True, "model": model_name, "duration_ms": round(duration * 1000, 2), "load_duration_ms": round(data.get("load_duration", 0) / 1e6, 2), "total_duration_ms": round(data.get("total_duration", 0) / 1e6, 2), } - else: - return { - "success": False, - "model": model_name, - "error": f"HTTP {resp.status_code}: {resp.text}", - "duration_ms": round(duration * 1000, 2), - } + telemetry_store.record_event({ + "event_type": "Model Warm", "source": "warmer", "target": model_name, + "duration_ms": res["duration_ms"], "load_duration_ms": res["load_duration_ms"], + }) + return res + return { + "success": False, "model": model_name, + "error": f"HTTP {resp.status_code}: {resp.text}", + "duration_ms": round(duration * 1000, 2), + } except Exception as e: - return {"success": False, "model": model_name, "error": str(e), "duration_ms": round((time.perf_counter() - t0) * 1000, 2)} + return {"success": False, "model": model_name, "error": str(e), + "duration_ms": round((time.perf_counter() - t0) * 1000, 2)} -def find_comfy_model_files() -> List[Dict[str, Any]]: - """Discover all model files under ComfyUI models.""" - results = [] - extensions = ("*.safetensors", "*.ckpt", "*.pt", "*.bin") - if os.path.exists(COMFY_MODELS_DIR): - for root, _, files in os.walk(COMFY_MODELS_DIR): - for file in files: - if any(file.endswith(ext.replace("*", "")) for ext in extensions): - full_path = os.path.join(root, file) - rel_path = os.path.relpath(full_path, COMFY_MODELS_DIR) - size = os.path.getsize(full_path) - results.append({ - "filename": file, - "rel_path": rel_path, - "full_path": full_path, - "size_bytes": size, - "size_mb": round(size / (1024**2), 2), - "size_gb": round(size / (1024**3), 3), - }) - return results -async def warm_all_models() -> Dict[str, Any]: - """Warm all available Ollama and ComfyUI models into Linux RAM Cache.""" - t0 = time.perf_counter() - warmed_ollama = [] - warmed_comfy = [] - - # 1. Ollama models +def warm_ollama_blob(model_name: str) -> Dict[str, Any]: + """Warm a specific Ollama model's GGUF into page cache without touching VRAM.""" + for f in find_ollama_model_files(): + if f["model"] == model_name: + res = warm_file_to_ram(f["full_path"]) + res["model"] = model_name + return res + return {"success": False, "error": f"no blob found for model '{model_name}'"} + + +def _warm_priority(days: float = 30.0) -> Dict[str, float]: + """Recency/frequency score per model name, from the persisted event log.""" try: - async with httpx.AsyncClient(timeout=10.0) as client: - tags_resp = await client.get(f"{OLLAMA_API_BASE}/api/tags") - if tags_resp.status_code == 200: - models = tags_resp.json().get("models", []) - for m in models: - name = m.get("name") - res = await warm_ollama_model(name, keep_alive="1m") - warmed_ollama.append(res) - except Exception as e: - logger.error(f"Error discovering Ollama models: {e}") - - # 2. ComfyUI models - comfy_files = find_comfy_model_files() - for f in comfy_files: - res = warm_file_to_ram(f["full_path"]) - warmed_comfy.append(res) - - total_duration_ms = round((time.perf_counter() - t0) * 1000, 2) - meminfo = get_detailed_meminfo() - + return {r["model"]: r["score"] for r in telemetry_store.model_usage_ranking(days)} + except Exception: + return {} + + +def build_warm_plan(budget_gb: Optional[float] = None) -> Dict[str, Any]: + """Decide *what* to warm, in what order, within a byte budget. + + Warming everything on a 64 GB box with 33+ GB of models just evicts the earliest + files, so we rank by usage (Ollama, from history) and recency (ComfyUI, by mtime), + then fill until the budget is spent. Already-resident files cost nothing. + """ + mem = get_detailed_meminfo() + if budget_gb is None: + # Leave headroom so warming never pushes the box into reclaim. + budget_gb = max((mem["available_bytes"] * 0.7) / (1024**3), 1.0) + budget_bytes = int(budget_gb * (1024**3)) + + catalog = get_model_catalog() + scores = _warm_priority() + now = time.time() + + candidates = [] + for f in catalog["ollama"]: + candidates.append({**f, "score": scores.get(f["model"], 0.0) + 0.5, + "name": f["model"]}) + for f in catalog["comfy"]: + age_days = max((now - f.get("mtime", now)) / 86400.0, 0.01) + candidates.append({**f, "score": scores.get(f["rel_path"], 0.0) + 1.0 / (1.0 + age_days), + "name": f["rel_path"]}) + + candidates.sort(key=lambda c: c["score"], reverse=True) + + plan, spent, skipped = [], 0, [] + seen_paths = set() + for c in candidates: + if c["full_path"] in seen_paths: + continue + seen_paths.add(c["full_path"]) + res = page_residency(c["full_path"]) + entry = { + "name": c["name"], "kind": c["kind"], "full_path": c["full_path"], + "size_gb": c.get("size_gb", 0), "score": round(c["score"], 4), + "resident_pct": res.get("resident_pct", 0.0), + } + if res.get("warm"): + entry["action"] = "already-warm" + skipped.append(entry) + continue + need = int(c.get("size_bytes", 0) * (1 - res.get("resident_pct", 0) / 100.0)) + if spent + need > budget_bytes: + entry["action"] = "over-budget" + skipped.append(entry) + continue + spent += need + entry["action"] = "warm" + entry["bytes_to_read"] = need + plan.append(entry) + + return { + "budget_gb": round(budget_gb, 2), + "planned_gb": round(spent / (1024**3), 2), + "warm_count": len(plan), + "skipped_count": len(skipped), + "plan": plan, + "skipped": skipped, + "meminfo": mem, + } + + +async def warm_all_models(budget_gb: Optional[float] = None, + include_vram_load: bool = False) -> Dict[str, Any]: + """Warm the highest-value models into the page cache within a byte budget.""" + t0 = time.perf_counter() + plan = build_warm_plan(budget_gb) + warmed = [] + for entry in plan["plan"]: + res = warm_file_to_ram(entry["full_path"]) + res["name"] = entry["name"] + res["kind"] = entry["kind"] + warmed.append(res) + # Budgets are computed up front, but the page cache is shared with the rest of + # the box; bail out if we start pushing the system into reclaim. + if get_detailed_meminfo()["available_gb"] < 4.0: + logger.warning("warm_all_models: stopping early, MemAvailable below 4 GB") + break + + if include_vram_load: + try: + async with httpx.AsyncClient(timeout=10.0) as client: + tags = await client.get(f"{OLLAMA_API_BASE}/api/tags") + if tags.status_code == 200: + top = sorted(tags.json().get("models", []), + key=lambda m: _warm_priority().get(m.get("name"), 0), + reverse=True)[:1] + for m in top: + await warm_ollama_model(m.get("name"), keep_alive="1m") + except Exception as e: + logger.debug(f"optional VRAM preload skipped: {e}") + return { "status": "completed", - "total_duration_ms": total_duration_ms, - "ollama_models_warmed": warmed_ollama, - "comfy_files_warmed": warmed_comfy, - "meminfo_after": meminfo, + "total_duration_ms": round((time.perf_counter() - t0) * 1000, 2), + "budget_gb": plan["budget_gb"], + "planned_gb": plan["planned_gb"], + "files_warmed": warmed, + "bytes_read": sum(w.get("bytes_read", 0) for w in warmed), + "skipped": plan["skipped"], + "meminfo_after": get_detailed_meminfo(), } diff --git a/server.py b/server.py index 5ebfb7f..c172792 100644 --- a/server.py +++ b/server.py @@ -1,27 +1,169 @@ """FastAPI Backend Server with SSE Real-Time Telemetry and Model Orchestration API.""" import asyncio +import contextlib import json import logging -from typing import Dict, Any, Optional, List +import time +from contextlib import asynccontextmanager +from typing import Dict, Any, Optional, List, Set + from fastapi import FastAPI, Request, HTTPException, Query from fastapi.responses import HTMLResponse, StreamingResponse, JSONResponse from fastapi.staticfiles import StaticFiles from fastapi.middleware.cors import CORSMiddleware from pydantic import BaseModel, Field -import ram_optimizer -import vram_arbitrator +import autotune import overclock_manager +import ram_optimizer +import telemetry_store +import thermal_governor +import vram_arbitrator logging.basicConfig(level=logging.INFO, format="%(asctime)s [%(levelname)s] %(name)s: %(message)s") logger = logging.getLogger("model_manager_server") +# The sampler makes four HTTP calls a second; at INFO, httpx narrates every one of them. +logging.getLogger("httpx").setLevel(logging.WARNING) +logging.getLogger("httpcore").setLevel(logging.WARNING) + +BASE_DIR = "/home/drjones/unified-model-manager" + + +# ========================================== +# TELEMETRY BROKER +# ========================================== + +class TelemetryBroker: + """One sampler, many subscribers. + + Every SSE client used to run its own copy of the full snapshot once per second: + NVML queries, /proc/meminfo, an HTTP round-trip each to Ollama and ComfyUI, and โ the + expensive one โ a recursive walk of the ComfyUI models tree with a stat() per + checkpoint. Opening the dashboard in three tabs tripled the load on the very thing it + was measuring. Now a single background task samples at 1 Hz and fans the snapshot out. + + The sampler is also the natural feed for the thermal governor and the persistence + layer, so neither needs to poll the GPU on its own. + """ + + def __init__(self, interval_s: float = 1.0) -> None: + self.interval_s = interval_s + self.snapshot: Dict[str, Any] = {} + self.subscribers: Set[asyncio.Queue] = set() + self.task: Optional[asyncio.Task] = None + self.running = False + self.samples = 0 + self.last_sample_ms = 0.0 + + async def start(self) -> None: + if self.running: + return + self.running = True + self.task = asyncio.create_task(self._loop()) + + async def stop(self) -> None: + self.running = False + if self.task: + self.task.cancel() + with contextlib.suppress(asyncio.CancelledError): + await self.task + + def subscribe(self) -> asyncio.Queue: + q: asyncio.Queue = asyncio.Queue(maxsize=2) + self.subscribers.add(q) + return q + + def unsubscribe(self, q: asyncio.Queue) -> None: + self.subscribers.discard(q) + + async def _loop(self) -> None: + while self.running: + t0 = time.perf_counter() + try: + snap = await self._sample() + self.snapshot = snap + self.samples += 1 + self.last_sample_ms = round((time.perf_counter() - t0) * 1000, 2) + + # Feed the governor and the durable store from the sample we already have. + thermal_governor.governor.observe(snap.get("gpu", {}), + overclock_manager.ACTIVE_PROFILE) + telemetry_store.record_telemetry( + snap.get("gpu", {}), snap.get("ram", {}), + profile=overclock_manager.ACTIVE_PROFILE, + throttle_reasons=",".join(snap.get("gpu", {}).get("throttle_reasons") or []), + ) + + for q in list(self.subscribers): + if q.full(): + # Slow client: drop the stale frame rather than stalling the sampler. + with contextlib.suppress(asyncio.QueueEmpty): + q.get_nowait() + with contextlib.suppress(asyncio.QueueFull): + q.put_nowait(snap) + except asyncio.CancelledError: + raise + except Exception as e: + logger.error(f"telemetry sampler error: {e}") + await asyncio.sleep(max(self.interval_s - (time.perf_counter() - t0), 0.05)) + + async def _sample(self) -> Dict[str, Any]: + gpu_stats = vram_arbitrator.get_gpu_hardware_stats() + mem_stats = ram_optimizer.get_detailed_meminfo() + ollama_state, comfy_state = await asyncio.gather( + vram_arbitrator.get_ollama_live_state(), + vram_arbitrator.get_comfyui_live_state(), + ) + return { + "timestamp": time.time(), # wall clock, not the event loop's monotonic clock + "monotonic": asyncio.get_running_loop().time(), + "gpu": gpu_stats, + "ram": mem_stats, + "ollama": ollama_state, + "comfyui": comfy_state, + "arbitrator": vram_arbitrator.arbitrator.get_status(), + "governor": thermal_governor.governor.get_status(), + "overclock": {"active_profile": overclock_manager.ACTIVE_PROFILE}, + "history": vram_arbitrator.get_switch_history(), + "comfy_models_count": len(ram_optimizer.find_comfy_model_files()), + "sampler": {"samples": self.samples, "last_sample_ms": self.last_sample_ms, + "subscribers": len(self.subscribers)}, + } + + async def get(self) -> Dict[str, Any]: + """Latest snapshot, sampling on demand if the loop has not produced one yet.""" + if not self.snapshot: + self.snapshot = await self._sample() + return self.snapshot + + +broker = TelemetryBroker() + + +@asynccontextmanager +async def lifespan(app: FastAPI): + telemetry_store.start() + await broker.start() + await vram_arbitrator.arbitrator.start() + yield + await vram_arbitrator.arbitrator.stop() + await broker.stop() + # Never leave the card with locked clocks and pinned fans after we exit. + try: + overclock_manager.restore_safe("server shutdown") + except Exception as e: + logger.error(f"restore_safe on shutdown failed: {e}") + telemetry_store.stop() + + app = FastAPI( title="HyperSwap // GPU Program Swapper & Telemetry API", - version="1.0.0", + version="2.0.0", description="High-performance VRAM arbitration and 64GB RAM cache orchestrator for simultaneous Ollama and ComfyUI workloads on Linux.", docs_url="/docs", redoc_url="/redoc", + lifespan=lifespan, ) app.add_middleware( @@ -32,26 +174,24 @@ app.add_middleware( allow_headers=["*"], ) -@app.on_event("startup") -async def on_startup(): - await vram_arbitrator.arbitrator.start() - -@app.on_event("shutdown") -async def on_shutdown(): - await vram_arbitrator.arbitrator.stop() # Pydantic Request Models class SwitchRequest(BaseModel): model: str = Field(..., description="Name of the Ollama model to hot-swap to in VRAM", example="qwen3.8fast:latest") keep_alive: Optional[str] = Field("30m", description="Keep-alive duration in VRAM (e.g. 5m, 30m, 0)", example="30m") + free_comfy_first: bool = Field(False, description="Purge ComfyUI VRAM first if it is holding memory") class WarmRequest(BaseModel): - model_name: Optional[str] = Field(None, description="Ollama model name to warm into OS page cache", example="gemma4:26b") - filepath: Optional[str] = Field(None, description="Absolute file path of Safetensors/GGUF to warm into RAM", example="/home/drjones/ComfyUI/models/checkpoints/v1-5-pruned-emaonly-fp16.safetensors") + model_name: Optional[str] = Field(None, description="Ollama model name to warm", example="gemma4:26b") + filepath: Optional[str] = Field(None, description="Absolute file path of Safetensors/GGUF to warm into RAM") + blob_only: bool = Field(False, description="Warm the model's weights into page cache without loading VRAM") + +class WarmAllRequest(BaseModel): + budget_gb: Optional[float] = Field(None, description="Byte budget for warming; defaults to 70% of MemAvailable", example=24.0) class BenchmarkRequest(BaseModel): iterations: Optional[int] = Field(2, description="Number of back-and-forth switch iterations to measure", example=2) - models: Optional[List[str]] = Field(None, description="Optional pair of models to benchmark between", example=["qwen3.8fast:latest", "smtek/Qwen3.8-27B:Q2_K_XL"]) + models: Optional[List[str]] = Field(None, description="Optional pair of models to benchmark between") class OverclockApplyRequest(BaseModel): profile: str = Field(..., description="Profile name: ollama | comfy | balanced", example="ollama") @@ -64,6 +204,23 @@ class FanRequest(BaseModel): percent: Optional[int] = Field(None, description="Fan speed 30-100 when mode=manual", example=70) speed_pct: Optional[int] = Field(None, description="Alias for percent (30-100)", example=70) +class GovernorRequest(BaseModel): + enabled: Optional[bool] = Field(None, description="Enable or disable the thermal governor") + reset: bool = Field(False, description="Clear any active derate and reapply the full profile") + +class SweepRequest(BaseModel): + knob: str = Field("mem_offset_mhz", description="mem_offset_mhz | core_offset_mhz") + profile: str = Field("ollama", description="Profile to tune") + model: Optional[str] = Field(None, description="Model to benchmark with; defaults to the loaded one") + start: Optional[int] = Field(None, description="First offset value") + stop: Optional[int] = Field(None, description="Last offset value") + step: Optional[int] = Field(None, description="Offset increment") + repeats: int = Field(1, description="Benchmark runs per step") + apply_best: bool = Field(False, description="Write the winning value into the profile") + +class RequestVramRequest(BaseModel): + needed_gb: float = Field(0.0, description="How much free VRAM Ollama needs", example=12.0) + # ========================================== # REST API ENDPOINTS @@ -71,52 +228,44 @@ class FanRequest(BaseModel): @app.get("/api/stats", summary="Full System Snapshot", tags=["Telemetry"]) async def get_all_stats() -> Dict[str, Any]: - """Gather complete live snapshot of GPU hardware, host RAM, Ollama, ComfyUI, and switch history.""" - gpu_stats = vram_arbitrator.get_gpu_hardware_stats() - mem_stats = ram_optimizer.get_detailed_meminfo() - ollama_state = await vram_arbitrator.get_ollama_live_state() - comfy_state = await vram_arbitrator.get_comfyui_live_state() - history = vram_arbitrator.get_switch_history() - comfy_models = ram_optimizer.find_comfy_model_files() - arbitrator_status = vram_arbitrator.arbitrator.get_status() - - return { - "timestamp": asyncio.get_event_loop().time(), - "gpu": gpu_stats, - "ram": mem_stats, - "ollama": ollama_state, - "comfyui": comfy_state, - "arbitrator": arbitrator_status, - "history": history, - "comfy_models_count": len(comfy_models), - } + """Latest unified snapshot of GPU hardware, host RAM, Ollama, ComfyUI and swap history.""" + return await broker.get() @app.get("/api/gpu", summary="GPU Sensors and VRAM Breakdown", tags=["Telemetry"]) async def get_gpu_metrics() -> Dict[str, Any]: - """Retrieve detailed NVML sensors (utilization %, temp, power, fan, clocks, and per-process VRAM allocation).""" + """Detailed NVML sensors (utilization, temp, power, fan, clocks, throttle reasons, per-process VRAM).""" return vram_arbitrator.get_gpu_hardware_stats() @app.get("/api/memory", summary="Host RAM and Page Cache Breakdown", tags=["Telemetry"]) async def get_ram_metrics() -> Dict[str, Any]: - """Retrieve precise host 64GB DDR5 RAM breakdown, active cache size, and cache hit ratios.""" + """Precise host RAM breakdown, active cache size and cache ratios.""" return ram_optimizer.get_detailed_meminfo() @app.get("/api/stream", summary="Real-Time SSE Telemetry Stream", tags=["Telemetry"]) async def sse_telemetry_stream(request: Request): - """Server-Sent Events (SSE) streaming real-time statistics at 1Hz for dynamic dashboards.""" + """Server-Sent Events stream of the shared 1Hz snapshot.""" async def event_generator(): - while True: - if await request.is_disconnected(): - break - try: - stats = await get_all_stats() - yield f"data: {json.dumps(stats)}\n\n" - except Exception as e: - logger.error(f"SSE stream error: {e}") - yield f"data: {json.dumps({'error': str(e)})}\n\n" - await asyncio.sleep(1.0) - + q = broker.subscribe() + try: + snap = await broker.get() + yield f"data: {json.dumps(snap)}\n\n" + while True: + if await request.is_disconnected(): + break + try: + snap = await asyncio.wait_for(q.get(), timeout=15.0) + yield f"data: {json.dumps(snap)}\n\n" + except asyncio.TimeoutError: + yield ": keepalive\n\n" + except asyncio.CancelledError: + raise + except Exception as e: + logger.error(f"SSE stream error: {e}") + yield f"data: {json.dumps({'error': str(e)})}\n\n" + finally: + broker.unsubscribe(q) + return StreamingResponse( event_generator(), media_type="text/event-stream", @@ -129,80 +278,142 @@ async def sse_telemetry_stream(request: Request): @app.post("/api/switch-model", summary="Hot-Swap Ollama LLM in VRAM", tags=["Orchestration"]) async def api_switch_model(req: SwitchRequest): - """Hot-swap the active Ollama model in VRAM and measure exact load duration and token evaluation speed.""" + """Hot-swap the active Ollama model, measuring real load bandwidth and token throughput.""" + if req.free_comfy_first: + await vram_arbitrator.arbitrator.request_vram_for_ollama() res = await vram_arbitrator.switch_ollama_model(req.model, keep_alive=req.keep_alive or "30m") if not res.get("success"): raise HTTPException(status_code=500, detail=res.get("error")) return res @app.post("/api/free-vram", summary="Soft-Yield Ollama VRAM", tags=["Orchestration"]) -async def api_free_vram(): - """Instruct Ollama to instantly yield VRAM to 0MB in ~15ms while preserving model weights in the 64GB host RAM page cache.""" - return await vram_arbitrator.instant_free_ollama_vram() +async def api_free_vram(confirm: bool = Query(True, description="Wait for the driver to actually release the allocation")): + """Yield Ollama's VRAM and wait for the release to be confirmed by NVML.""" + return await vram_arbitrator.instant_free_ollama_vram(confirm=confirm) @app.post("/api/comfy-free", summary="Purge ComfyUI VRAM Cache", tags=["Orchestration"]) async def api_comfy_free(): """Purge loaded diffusion models and VRAM cache from the ComfyUI pipeline.""" return await vram_arbitrator.instant_free_comfyui_vram() -@app.post("/api/warm-all", summary="Pre-warm All Models into RAM Cache", tags=["Memory Optimization"]) -async def api_warm_all(): - """Pre-fault and read all installed Ollama GGUF models and ComfyUI Safetensors checkpoints into the Linux OS Page Cache.""" - return await ram_optimizer.warm_all_models() +@app.post("/api/request-vram", summary="Ask for VRAM on Ollama's behalf", tags=["Orchestration"]) +async def api_request_vram(req: RequestVramRequest): + """Force an immediate ComfyUI purge if there is not enough free VRAM for Ollama.""" + return await vram_arbitrator.arbitrator.request_vram_for_ollama(req.needed_gb) + +@app.post("/api/warm-all", summary="Pre-warm Models into RAM Cache", tags=["Memory Optimization"]) +async def api_warm_all(req: Optional[WarmAllRequest] = None): + """Warm the highest-value models into the page cache within a byte budget.""" + return await ram_optimizer.warm_all_models(budget_gb=req.budget_gb if req else None) + +@app.get("/api/warm-plan", summary="Preview the Warm Plan", tags=["Memory Optimization"]) +async def api_warm_plan(budget_gb: Optional[float] = Query(None, description="Override the byte budget")): + """Show what warming would read, in what order, and what it would skip โ without doing it.""" + return ram_optimizer.build_warm_plan(budget_gb) @app.post("/api/warm-model", summary="Pre-warm Single Model or File", tags=["Memory Optimization"]) async def api_warm_model(req: WarmRequest): - """Pre-warm a specific Ollama model or individual file path into Linux RAM cache.""" + """Pre-warm a specific Ollama model or file path into the Linux page cache.""" + if req.model_name and req.blob_only: + return ram_optimizer.warm_ollama_blob(req.model_name) if req.model_name: return await ram_optimizer.warm_ollama_model(req.model_name, keep_alive="1m") - elif req.filepath: + if req.filepath: return ram_optimizer.warm_file_to_ram(req.filepath) - else: - raise HTTPException(status_code=400, detail="model_name or filepath required") + raise HTTPException(status_code=400, detail="model_name or filepath required") + +@app.get("/api/cache/report", summary="Measured Page-Cache Residency", tags=["Memory Optimization"]) +async def api_cache_report(files: bool = Query(True), refresh: bool = Query(False)): + """Measured (not assumed) page-cache residency for every model on disk.""" + report = ram_optimizer.get_cache_report(include_files=files, force_refresh=refresh) + report["capability"] = ram_optimizer.residency_capability() + return report @app.get("/api/models", summary="List All Installed Models", tags=["Catalog"]) -async def api_get_models(): - """List all installed Ollama models and discovered ComfyUI model checkpoints/safetensors on disk with sizes and quantization levels.""" +async def api_get_models(refresh: bool = Query(False)): + """All installed Ollama models (with their on-disk blobs) and ComfyUI checkpoints.""" + catalog = ram_optimizer.get_model_catalog(force_refresh=refresh) ollama_state = await vram_arbitrator.get_ollama_live_state() - comfy_models = ram_optimizer.find_comfy_model_files() return { "ollama_models": ollama_state.get("installed_models", []), - "comfy_models": comfy_models, + "ollama_blobs": catalog["ollama"], + "comfy_models": catalog["comfy"], + "cached_at": catalog["cached_at"], } @app.get("/api/history", summary="Model Switch History Log", tags=["Analytics"]) -async def api_get_history(limit: int = Query(20, description="Max history items to return")): - """Get the recent history log of model switch events, swap durations (in ms), and RAM cache hit status.""" - history = vram_arbitrator.get_switch_history() - return history[:limit] +async def api_get_history(limit: int = Query(20, description="Max history items to return"), + durable: bool = Query(False, description="Read from the persistent store instead of the in-memory ring")): + """Recent swap events, durations, achieved bandwidth and cache status.""" + if durable: + return telemetry_store.recent_events(limit) + return vram_arbitrator.get_switch_history()[:limit] @app.post("/api/benchmark", summary="Run Latency Benchmark", tags=["Analytics"]) async def api_run_benchmark(req: BenchmarkRequest): - """Run an automated benchmark swapping between available models to measure round-trip latency and RAM cache effectiveness.""" + """Automated round-trip switch benchmark measuring latency and cache effectiveness.""" from mcp_server import run_model_switch_benchmark res_str = await run_model_switch_benchmark(iterations=req.iterations or 2) return json.loads(res_str) + +# ========================================== +# ANALYTICS (persisted) +# ========================================== + +@app.get("/api/analytics/profiles", summary="Which Overclock Profile Is Actually Faster", tags=["Analytics"]) +async def api_analytics_profiles(days: float = Query(7.0)): + """Decode throughput and thermals grouped by the profile that was active at the time.""" + return {"window_days": days, "profiles": telemetry_store.profile_comparison(days)} + +@app.get("/api/analytics/swaps", summary="Swap Statistics", tags=["Analytics"]) +async def api_analytics_swaps(days: float = Query(7.0)): + """Aggregated swap/yield/purge latencies, cache-hit split and per-model throughput.""" + return telemetry_store.swap_stats(days) + +@app.get("/api/analytics/timeseries", summary="Downsampled Telemetry History", tags=["Analytics"]) +async def api_analytics_timeseries(hours: float = Query(6.0), buckets: int = Query(240)): + """Long-range history for charts that outlive a page refresh.""" + return {"hours": hours, "points": telemetry_store.timeseries(hours, buckets)} + +@app.get("/api/analytics/models", summary="Model Usage Ranking", tags=["Analytics"]) +async def api_analytics_models(days: float = Query(30.0)): + """Recency/frequency ranking used to prioritise the RAM warm budget.""" + return {"window_days": days, "models": telemetry_store.model_usage_ranking(days)} + +@app.get("/api/db", summary="Telemetry Store Info", tags=["Analytics"]) +async def api_db_info(): + """Where the persistent store lives and how much history it holds.""" + return telemetry_store.db_info() + + # ========================================== # OVERCLOCK MANAGEMENT # ========================================== @app.get("/api/overclock", summary="Overclock Status & Profiles", tags=["Overclock"]) async def api_overclock_status(): - """Get live GPU overclock state, active profile, and all per-app profiles.""" - return overclock_manager.get_status() + """Live GPU overclock state, active profile, governor state and all per-app profiles.""" + status = overclock_manager.get_status() + status["governor"] = thermal_governor.governor.get_status() + return status @app.post("/api/overclock/apply", summary="Apply Overclock Profile", tags=["Overclock"]) async def api_overclock_apply(req: OverclockApplyRequest): - """Apply a named overclock profile (ollama | comfy | balanced) to the GPU immediately.""" + """Apply a named overclock profile (ollama | comfy | balanced) immediately.""" res = overclock_manager.apply_profile(req.profile) if not res.get("success"): raise HTTPException(status_code=400, detail=res.get("error")) return res +@app.post("/api/overclock/restore", summary="Restore Stock GPU State", tags=["Overclock"]) +async def api_overclock_restore(): + """Drop all clock locks and offsets, restore default power limit and automatic fans.""" + return overclock_manager.restore_safe("manual request") + @app.get("/api/overclock/profiles", summary="List Overclock Profiles", tags=["Overclock"]) async def api_overclock_profiles(): - """List all overclock profiles with their current settings.""" + """All overclock profiles with their current settings.""" return overclock_manager.get_profiles() @app.post("/api/overclock/profiles/{name}", summary="Update Overclock Profile", tags=["Overclock"]) @@ -216,7 +427,7 @@ async def api_overclock_update_profile(name: str, req: OverclockProfileUpdate): @app.get("/api/overclock/fan", summary="Get GPU Fan Status", tags=["Overclock"]) @app.get("/api/gpu/fan", summary="Get GPU Fan Status", tags=["Overclock"]) async def api_get_fan_status(): - """Get current GPU fan control mode and speed.""" + """Current GPU fan control mode and speed.""" return overclock_manager.get_fan_status() @app.post("/api/overclock/fan", summary="Set GPU Fan Speed", tags=["Overclock"]) @@ -228,12 +439,59 @@ async def api_set_fan(req: FanRequest): return overclock_manager.set_fan_speed(pct) return overclock_manager.set_fan_auto() + +# ========================================== +# THERMAL GOVERNOR +# ========================================== + +@app.get("/api/governor", summary="Thermal Governor State", tags=["Governor"]) +async def api_governor_status(): + """Current derate level, why it was applied, and the escalation history.""" + return thermal_governor.governor.get_status() + +@app.post("/api/governor", summary="Control the Thermal Governor", tags=["Governor"]) +async def api_governor_control(req: GovernorRequest): + """Enable/disable the governor, or clear an active derate.""" + if req.enabled is not None: + thermal_governor.governor.set_enabled(req.enabled) + if req.reset: + thermal_governor.governor.reset() + return thermal_governor.governor.get_status() + + +# ========================================== +# AUTOTUNE +# ========================================== + +@app.get("/api/autotune", summary="Autotune Status & History", tags=["Autotune"]) +async def api_autotune_status(): + """Sweep progress, the last result, and every recorded autotune step.""" + return autotune.get_status() + +@app.post("/api/autotune/sweep", summary="Run an Overclock Sweep", tags=["Autotune"]) +async def api_autotune_sweep(req: SweepRequest): + """Walk a clock offset upward, measuring tok/s and watching for instability at each step.""" + res = await autotune.sweep( + knob=req.knob, profile=req.profile, model=req.model, + start=req.start, stop=req.stop, step=req.step, + repeats=req.repeats, apply_best=req.apply_best, + ) + if not res.get("success"): + raise HTTPException(status_code=400, detail=res.get("error")) + return res + +@app.post("/api/autotune/cancel", summary="Cancel a Running Sweep", tags=["Autotune"]) +async def api_autotune_cancel(): + """Stop the current sweep after the step in flight; the profile is restored either way.""" + return autotune.cancel() + + # Mount static web UI files -app.mount("/static", StaticFiles(directory="/home/drjones/unified-model-manager/static"), name="static") +app.mount("/static", StaticFiles(directory=f"{BASE_DIR}/static"), name="static") @app.get("/", summary="Dashboard Web UI", tags=["UI"]) async def root_index(): - with open("/home/drjones/unified-model-manager/static/index.html", "r") as f: + with open(f"{BASE_DIR}/static/index.html", "r") as f: content = f.read() return HTMLResponse(content=content) diff --git a/static/app.js b/static/app.js index 686932a..b388f61 100644 --- a/static/app.js +++ b/static/app.js @@ -33,6 +33,9 @@ function initSSE() { function updateDashboard(data) { if (!data) return; + // Governor state rides along in the shared snapshot โ no extra polling needed. + if (data.governor) renderGovernor(data.governor); + // 1. GPU VRAM Stats const gpu = data.gpu || {}; const ram = data.ram || {}; @@ -646,3 +649,231 @@ async function saveOverclockProfile() { alert(`Error: ${err}`); } } + +// ============================================================================ +// THERMAL GOVERNOR / RESIDENCY / ANALYTICS / AUTOTUNE +// ============================================================================ + +let governorEnabled = true; + +function renderGovernor(gov) { + if (!gov) return; + governorEnabled = gov.enabled; + const levels = 3; + const el = (id) => document.getElementById(id); + if (!el('gov-label')) return; + + el('gov-label').textContent = gov.label || 'โ'; + el('gov-label').className = 'text-2xl font-bold ' + + (gov.level === 0 ? 'text-emerald-400' : gov.level < 3 ? 'text-amber-400' : 'text-rose-400'); + el('gov-level').textContent = `level ${gov.level} / ${levels}`; + el('gov-bar').style.width = `${(gov.level / levels) * 100}%`; + el('gov-esc').textContent = `${gov.escalate_at_c}ยฐC`; + el('gov-rec').textContent = `${gov.recover_below_c}ยฐC`; + el('gov-scale').textContent = `${Math.round((gov.offset_scale ?? 1) * 100)}%`; + el('gov-reason').textContent = gov.last_reason || 'โ'; + + const toggle = el('gov-toggle'); + toggle.textContent = gov.enabled ? 'Enabled' : 'Disabled'; + toggle.className = 'px-2.5 py-1 text-xs font-semibold rounded-lg border transition ' + + (gov.enabled ? 'bg-emerald-950/70 border-emerald-800 text-emerald-300 hover:bg-emerald-900' + : 'bg-slate-800 border-slate-700 text-slate-400 hover:bg-slate-700'); + + const hist = el('gov-history'); + hist.innerHTML = (gov.history || []).map(h => { + const t = new Date(h.ts * 1000).toLocaleTimeString(); + const up = h.to_level > h.from_level; + return `
| profile | tok/s | load GB/s | ยฐC avg | W avg | SM MHz | n | +
|---|---|---|---|---|---|---|
| ${win ? 'โ ' : ''}${p.profile ?? 'โ'} | +${p.avg_tok_s ?? 'โ'} | ${p.avg_load_gbps ?? 'โ'} | +${p.avg_temp_c ?? 'โ'} | ${p.avg_power_w ?? 'โ'} | +${p.avg_clock_sm ?? 'โ'} | ${p.swaps} |
| offset | tok/s | ยฐC | W | mem MHz | status | +
|---|
Walks the overclock back when the card complains
+Last action: cold start
+ +What is genuinely in RAM, not what we hope is
+Decode throughput per profile, from persisted history
+Sweep a clock offset, measure tok/s, stop at instability
+