GPU VRAM (16 GB Dedicated)
Real-time allocation breakdown on RTX 4080 SUPER
Host RAM & Model Page Cache (64 GB)
Models stay resident in RAM for instant PCIe hot-swaps
Ollama LLM Engine
Port :11434 // FlashAttention + Q4 KV Cache
ComfyUI Diffusion Engine
Port :8188 // DynamicVRAM + Pinned Async Offload
GPU Live Telemetry (NVML)
NVIDIA GeForce RTX 4080 SUPER
| PID | Process | VRAM |
|---|---|---|
| Scanning GPU processes... | ||
Model Switch Timeline & Optimizer
Real-time latency logger and memory warmer
GPU Overclock Control
Per-app profiles auto-switch with the VRAM arbitrator
VRAM Arbitration
Who holds the GPU, and how handoffs are going
(busy)
later
purges
deferred
Deferred is healthy — an LLM mid-generation cannot unload, so the request queues and applies the moment it finishes. Only stalled (VRAM held while the GPU sits idle) indicates a real problem.
Thermal Governor
Walks the overclock back when the card complains
Last action: cold start
Measured Page-Cache Residency
What is genuinely in RAM, not what we hope is
Is the Overclock Actually Working?
Decode throughput per profile, from persisted history
Overclock Autotune
Sweep a clock offset, measure tok/s, stop at instability