GPU VRAM (16 GB Dedicated)
Real-time allocation breakdown on RTX 4080 SUPER
0.0
/ 16.0 GB
0% USED
Ollama: 0 GB
ComfyUI: 0 GB
System: 0 GB
Ollama:
0 GB
Comfy:
0 GB
System:
0 GB
Free:
0 GB
Host RAM & Model Page Cache (64 GB)
Models stay resident in RAM for instant PCIe hot-swaps
0.0
GB CACHED
60.3 GB Total
Apps Used:
0 GB
Models in RAM:
0 GB
Free RAM:
0 GB
Ollama LLM Engine
Port :11434 // FlashAttention + Q4 KV Cache
Active Model in VRAM
0.0 GB VRAM
None Loaded
0 ctx
LAST SWAP TIME
-- ms
RAM HIT STATUS
--
TOTAL MODELS
0
ComfyUI Diffusion Engine
Port :8188 // DynamicVRAM + Pinned Async Offload
Pipeline Status
IDLE / READY
Dynamic Model Offloader
Queue: 0
Host Pinned Memory:
53.6 GB Staging Buffer
Async PCIe Offloading:
Enabled (2 Streams)
Fast Disk RAM Mmap:
Active
DISCOVERED MODELS
0 Files
VRAM AVAILABLE
15.9 GB
GPU Live Telemetry (NVML)
NVIDIA GeForce RTX 4080 SUPER
GPU Util
0%
Temperature
0°C
Power Draw
0 W
Fan Speed
0%
Active Compute Processes
| PID | Process | VRAM |
|---|---|---|
| Scanning GPU processes... | ||
Model Switch Timeline & Optimizer
Real-time latency logger and memory warmer
Recent Model Swaps
No recent model swaps recorded yet.
64GB RAM Cache holds all models in memory
PCIe x16 (~31.5 GB/s)