Initial commit: HyperSwap GPU Program Swapper with REST API, MCP 2.0, Real-Time Dashboard and Memory Orchestrator
This commit is contained in:
275
README.md
Normal file
275
README.md
Normal file
@@ -0,0 +1,275 @@
|
||||
# HyperSwap // GPU Program Swapper & Memory Orchestrator
|
||||
|
||||
[](https://fastapi.tiangolo.com)
|
||||
[](https://modelcontextprotocol.io)
|
||||
[](https://developer.nvidia.com/cuda-zone)
|
||||
[](https://ubuntu.com)
|
||||
|
||||
**HyperSwap** is an ultra-low-latency VRAM arbitrator, RAM cache pre-warmer, and real-time telemetry dashboard designed specifically for Linux deployment machines that simultaneously host **Ollama LLM workloads** and **ComfyUI Diffusion pipelines** on a single GPU.
|
||||
|
||||
---
|
||||
|
||||
## 1. Architectural Overview & Physics of High-Speed Switching
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
subgraph HostRAM["64 GB DDR5 Host System RAM (Page Cache & Pinned Staging)"]
|
||||
OllamaGGUFs["Ollama GGUF Weights<br/>(Qwen, Gemma, Nemotron)"]
|
||||
ComfySafetensors["ComfyUI Safetensors & VAEs<br/>(53.6 GB Pinned Staging Buffer)"]
|
||||
end
|
||||
|
||||
subgraph GPU["NVIDIA GeForce RTX 4080 SUPER (16 GB VRAM)"]
|
||||
direction LR
|
||||
ActiveLLM["Active LLM<br/>(0–14 GB VRAM)"]
|
||||
ActiveDiffusion["Active Diffusion Pipeline<br/>(0–14 GB VRAM)"]
|
||||
end
|
||||
|
||||
subgraph Orchestrator["HyperSwap Control Plane (:9090)"]
|
||||
REST["REST API & OpenAPI Docs"]
|
||||
MCP["Model Context Protocol (MCP 2.0)"]
|
||||
SSE["1Hz Real-Time SSE Stream"]
|
||||
Arbitrator["VRAM Arbitrator (15ms Soft-Yield)"]
|
||||
Warmer["Page Cache Pre-Warmer"]
|
||||
end
|
||||
|
||||
HostRAM <== "PCIe 4.0 x16 Bus (~31.5 GB/s Hot-Swap)" ==> GPU
|
||||
Orchestrator --> GPU
|
||||
Orchestrator --> HostRAM
|
||||
```
|
||||
|
||||
### The Problem: Disk Bottleneck & VRAM Contention
|
||||
When running both Ollama and ComfyUI on a 16 GB GPU:
|
||||
* An active LLM (e.g. 27B–30B parameter quantized model) uses **11–15 GB VRAM**.
|
||||
* A diffusion model (SDXL, Flux, SD 1.5) requires **4–14 GB VRAM** during generation.
|
||||
* If models are evicted to NVMe storage, reloading weights takes **10–40 seconds** over disk I/O.
|
||||
|
||||
### The Solution: 64 GB RAM Cache + PCIe x16 Hot-Swapping
|
||||
* **Host RAM as Staging**: All active LLMs and diffusion checkpoints remain 100% resident in the 64 GB Linux OS Page Cache and pinned memory buffer.
|
||||
* **PCIe Bus Hot-Swap Speed**: Reloading from host RAM over the PCIe 4.0 x16 bus achieves **~31.5 GB/s** transfer bandwidth, bringing model swap times down to **hundreds of milliseconds**.
|
||||
* **15ms Soft-Yield**: When ComfyUI triggers an image generation, Ollama executes an instant soft-yield (`keep_alive: 0`), dropping VRAM allocation from 14.5 GB to 0 MB in **~15 milliseconds** without discarding model pages from system RAM.
|
||||
|
||||
---
|
||||
|
||||
## 2. REST API Reference
|
||||
|
||||
The HyperSwap server runs on port `9090` by default. Interactive OpenAPI/Swagger docs are accessible at `http://localhost:9090/docs`.
|
||||
|
||||
### Telemetry Endpoints
|
||||
|
||||
#### `GET /api/stats`
|
||||
Returns a unified JSON snapshot of all system sensors, GPU processes, host RAM, Ollama status, ComfyUI queue, and transition logs.
|
||||
|
||||
**Response (200 OK):**
|
||||
```json
|
||||
{
|
||||
"timestamp": 59583.16,
|
||||
"gpu": {
|
||||
"available": true,
|
||||
"device_name": "NVIDIA GeForce RTX 4080 SUPER",
|
||||
"vram_total_gb": 15.99,
|
||||
"vram_used_gb": 1.46,
|
||||
"vram_free_gb": 14.53,
|
||||
"vram_used_pct": 9.1,
|
||||
"gpu_util_pct": 11,
|
||||
"temperature_c": 48,
|
||||
"power_w": 31.4,
|
||||
"fan_pct": 0,
|
||||
"breakdown": {
|
||||
"ollama_gb": 0.0,
|
||||
"comfyui_gb": 0.24,
|
||||
"system_gb": 0.67,
|
||||
"free_gb": 14.53,
|
||||
"processes": [...]
|
||||
}
|
||||
},
|
||||
"ram": {
|
||||
"total_gb": 60.34,
|
||||
"used_gb": 6.72,
|
||||
"cached_gb": 36.21,
|
||||
"free_gb": 17.41,
|
||||
"cache_ratio_pct": 60.0
|
||||
},
|
||||
"ollama": {
|
||||
"online": true,
|
||||
"active_model_name": null,
|
||||
"active_model_vram_gb": 0.0,
|
||||
"installed_models": [...]
|
||||
},
|
||||
"comfyui": {
|
||||
"online": true,
|
||||
"executing": false,
|
||||
"queue_remaining": 0,
|
||||
"vram_free_mb": 14882.4
|
||||
},
|
||||
"history": [...]
|
||||
}
|
||||
```
|
||||
|
||||
#### `GET /api/gpu`
|
||||
Returns hardware sensors (utilization %, temperature, power draw in Watts, fan %, GPU graphics/memory clocks, and active PIDs).
|
||||
|
||||
#### `GET /api/memory`
|
||||
Returns precise `/proc/meminfo` metrics including Total, Used, OS Page Cache, and free memory.
|
||||
|
||||
#### `GET /api/stream`
|
||||
Server-Sent Events (SSE) stream pushing full telemetry updates at 1Hz (`Content-Type: text/event-stream`).
|
||||
|
||||
---
|
||||
|
||||
### Orchestration & Hot-Swap Endpoints
|
||||
|
||||
#### `POST /api/switch-model`
|
||||
Hot-swaps the active Ollama LLM in VRAM and tracks transition timing.
|
||||
|
||||
**Request Body:**
|
||||
```json
|
||||
{
|
||||
"model": "qwen3.8fast:latest",
|
||||
"keep_alive": "30m"
|
||||
}
|
||||
```
|
||||
|
||||
**Response (200 OK):**
|
||||
```json
|
||||
{
|
||||
"success": true,
|
||||
"prev_model": "None",
|
||||
"target_model": "qwen3.8fast:latest",
|
||||
"total_duration_ms": 1420.5,
|
||||
"load_duration_ms": 839.5,
|
||||
"tokens_per_sec": 42.0,
|
||||
"is_ram_hit": true,
|
||||
"response": "Ready."
|
||||
}
|
||||
```
|
||||
|
||||
#### `POST /api/free-vram`
|
||||
Instructs Ollama to soft-yield VRAM down to 0 MB in ~15 milliseconds while keeping model weights in 64GB RAM cache.
|
||||
|
||||
#### `POST /api/comfy-free`
|
||||
Instructs ComfyUI to purge loaded diffusion weights and VRAM cache.
|
||||
|
||||
#### `POST /api/warm-all`
|
||||
Pre-faults and reads all installed Ollama models and ComfyUI Safetensors into the Linux page cache.
|
||||
|
||||
#### `POST /api/warm-model`
|
||||
Pre-warms a specific model or file into RAM.
|
||||
|
||||
**Request Body:**
|
||||
```json
|
||||
{
|
||||
"model_name": "gemma4:26b",
|
||||
"filepath": null
|
||||
}
|
||||
```
|
||||
|
||||
#### `POST /api/benchmark`
|
||||
Runs an automated back-and-forth model swap benchmark and computes average transition latency.
|
||||
|
||||
---
|
||||
|
||||
## 3. Model Context Protocol (MCP) Reference
|
||||
|
||||
HyperSwap includes a native **MCP 2.0 server** ([mcp_server.py](file:///home/drjones/unified-model-manager/mcp_server.py)) that exposes all orchestration and telemetry functions as agentic tools.
|
||||
|
||||
### MCP Tools List
|
||||
|
||||
| Tool Name | Parameters | Description |
|
||||
| :--- | :--- | :--- |
|
||||
| **`get_gpu_status`** | *None* | Live NVIDIA GPU hardware telemetry, VRAM breakdown, temps, power, and PIDs. |
|
||||
| **`get_host_memory_status`** | *None* | 64GB host RAM breakdown, active page cache size, and cache ratio. |
|
||||
| **`switch_ollama_model`** | `model_name` (str), `keep_alive` (str, default "30m") | Hot-swaps active LLM in VRAM, measures latency (ms) and tokens/sec. |
|
||||
| **`soft_yield_ollama_vram`** | `model_name` (optional str) | Yields Ollama VRAM to 0 MB in ~15ms while keeping model weights in RAM cache. |
|
||||
| **`purge_comfyui_vram`** | *None* | Purges loaded diffusion models from ComfyUI pipeline VRAM. |
|
||||
| **`prewarm_all_models_to_ram`** | *None* | Faults all local LLM and diffusion checkpoints into Linux OS page cache. |
|
||||
| **`prewarm_single_model`** | `model_name` (optional str), `filepath` (optional str) | Pre-warms a single GGUF or Safetensors file into RAM. |
|
||||
| **`list_available_models`** | *None* | Lists all installed Ollama models and ComfyUI Safetensors on disk. |
|
||||
| **`get_switch_history`** | `limit` (int, default 20) | Retrieves recent switch events, millisecond latencies, and RAM hit status. |
|
||||
| **`run_model_switch_benchmark`**| `iterations` (int, default 2) | Automated round-trip latency benchmark between installed models. |
|
||||
|
||||
### MCP Resources List
|
||||
|
||||
* `gpu://metrics/live` - Real-time snapshot of GPU sensors and RAM page cache.
|
||||
* `gpu://models/catalog` - Catalog of all discovered GGUF and Safetensors models.
|
||||
* `gpu://history/switches` - Event log of recent model transitions and swap speeds.
|
||||
|
||||
---
|
||||
|
||||
### MCP Client Configurations
|
||||
|
||||
#### Antigravity Configuration (`~/.gemini/config/mcp_config.json`)
|
||||
```json
|
||||
{
|
||||
"mcpServers": {
|
||||
"hyperswap": {
|
||||
"command": "/home/drjones/comfy-mcp-venv/bin/python",
|
||||
"args": ["/home/drjones/unified-model-manager/mcp_server.py", "--stdio"]
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
#### Claude Desktop Configuration (`claude_desktop_config.json`)
|
||||
```json
|
||||
{
|
||||
"mcpServers": {
|
||||
"hyperswap": {
|
||||
"command": "/home/drjones/comfy-mcp-venv/bin/python",
|
||||
"args": ["/home/drjones/unified-model-manager/mcp_server.py", "--stdio"]
|
||||
}
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 4. Web Dashboard & Real-Time Telemetry
|
||||
|
||||
The Web Dashboard is hosted at `http://localhost:9090`.
|
||||
|
||||
* **Hero Memory Gauges**: Visual multi-segment bars representing VRAM allocation across Ollama, ComfyUI, and Desktop, alongside the 64GB host RAM page cache.
|
||||
* **Ollama Control Card**: Real-time active model indicator, hot-swap selector, context size, and 1-click VRAM yield button.
|
||||
* **ComfyUI Pipeline Card**: Live execution state (Idle vs Generating), active prompt queue counter, and VRAM purge controls.
|
||||
* **GPU Hardware Card**: Live gauges for GPU Core Utilization, Temperature (°C), Power Draw (W), Fan Speed (%), and active compute process table.
|
||||
* **Switch Timeline**: Real-time event feed detailing swap durations in milliseconds and RAM cache hit flags.
|
||||
|
||||
---
|
||||
|
||||
## 5. Linux Kernel & Host Tuning
|
||||
|
||||
To ensure that 45–50 GB of model weights remain permanently in RAM without kernel eviction:
|
||||
|
||||
```bash
|
||||
# Prioritize retaining model file cache in RAM (lower pressure = stronger cache retention)
|
||||
sudo sysctl -w vm.vfs_cache_pressure=10
|
||||
|
||||
# Reduce swap aggression for active pages
|
||||
sudo sysctl -w vm.swappiness=10
|
||||
|
||||
# Write changes permanently to /etc/sysctl.d/99-hyperswap.conf
|
||||
echo "vm.vfs_cache_pressure = 10" | sudo tee /etc/sysctl.d/99-hyperswap.conf
|
||||
echo "vm.swappiness = 10" | sudo tee -a /etc/sysctl.d/99-hyperswap.conf
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 6. Systemd Service Management
|
||||
|
||||
The manager runs as a persistent systemd user service:
|
||||
|
||||
```bash
|
||||
# Check status
|
||||
systemctl --user status hyperswap-manager.service
|
||||
|
||||
# Restart service
|
||||
systemctl --user restart hyperswap-manager.service
|
||||
|
||||
# View live logs
|
||||
journalctl --user -u hyperswap-manager.service -f
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 7. License
|
||||
|
||||
MIT License. Developed for Google Antigravity & High-Throughput Linux AI Deployments.
|
||||
Reference in New Issue
Block a user