GPU VRAM (16 GB Dedicated)
Real-time allocation breakdown on RTX 4080 SUPER
Host RAM & Model Page Cache (64 GB)
Models stay resident in RAM for instant PCIe hot-swaps
Ollama LLM Engine
Port :11434 // FlashAttention + Q4 KV Cache
ComfyUI Diffusion Engine
Port :8188 // DynamicVRAM + Pinned Async Offload
GPU Live Telemetry (NVML)
NVIDIA GeForce RTX 4080 SUPER
| PID | Process | VRAM |
|---|---|---|
| Scanning GPU processes... | ||
Model Switch Timeline & Optimizer
Real-time latency logger and memory warmer
GPU Overclock Control
Per-app profiles auto-switch with the VRAM arbitrator
System Health
Every dependency, with impact and how to fix it
VRAM Arbitration
Who holds the GPU, and how handoffs are going
(busy)
later
purges
deferred
Deferred is healthy — an LLM mid-generation cannot unload, so the request queues and applies the moment it finishes. Only stalled (VRAM held while the GPU sits idle) indicates a real problem.
Thermal Governor
Walks the overclock back when the card complains
Last action: cold start
Measured Page-Cache Residency
What is genuinely in RAM, not what we hope is
Is the Overclock Actually Working?
Decode throughput per profile, from persisted history
Overclock Autotune
Sweep a clock offset, measure tok/s, stop at instability