HYPERSWAP

v1.0-DEPLOY

NVIDIA RTX 4080 SUPER 16GB // 64GB DDR5 RAM // Ubuntu Linux

Ollama LLM: ONLINE
ComfyUI: ONLINE
SSE 1Hz

GPU VRAM (16 GB Dedicated)

Real-time allocation breakdown on RTX 4080 SUPER

0.0 / 16.0 GB
0% USED
Ollama: 0 GB
ComfyUI: 0 GB
System: 0 GB
Ollama: 0 GB
Comfy: 0 GB
System: 0 GB
Free: 0 GB

Host RAM & Model Page Cache (64 GB)

Models stay resident in RAM for instant PCIe hot-swaps

0.0 GB CACHED
60.3 GB Total
Apps Used: 0 GB
Models in RAM: 0 GB
Free RAM: 0 GB

Ollama LLM Engine

Port :11434 // FlashAttention + Q4 KV Cache

Active Model in VRAM 0.0 GB VRAM
None Loaded
0 ctx
LAST SWAP TIME -- ms
RAM HIT STATUS --
TOTAL MODELS 0

ComfyUI Diffusion Engine

Port :8188 // DynamicVRAM + Pinned Async Offload

Pipeline Status IDLE / READY
Dynamic Model Offloader
Queue: 0
Host Pinned Memory: 53.6 GB Staging Buffer
Async PCIe Offloading: Enabled (2 Streams)
Fast Disk RAM Mmap: Active
DISCOVERED MODELS 0 Files
VRAM AVAILABLE 15.9 GB

GPU Live Telemetry (NVML)

NVIDIA GeForce RTX 4080 SUPER

GPU Util 0%
Temperature 0°C
Power Draw 0 W
Fan Speed
0%
Fan 0: --% | Fan 1: --%
Active Compute Processes
PID Process VRAM
Scanning GPU processes...

Model Switch Timeline & Optimizer

Real-time latency logger and memory warmer

Recent Model Swaps
No recent model swaps recorded yet.
64GB RAM Cache holds all models in memory
PCIe x16 (~31.5 GB/s)

GPU Overclock Control

Per-app profiles auto-switch with the VRAM arbitrator

Active: -- X: --
Power Limit -- W
Core Clock --
Mem Clock --
Temp / Draw --
GPU Fan Cooling Control Dual-fan PWM active speed regulation
AUTO (VBIOS)
Fan 0 (Intake/Core): --%
Fan 1 (Exhaust/VRM): --%
Custom Manual Target: 65%
GPU --
Driver --
VRAM --
Max Core --
Max Mem --
Power --
Live Tuning Graph
Core MHz Mem MHz Temp °C Power W Fan % VRAM GB RAM Cache GB
Fine-tune profile
Lock core to max boost (2900–3105 MHz)
Lock mem to max (11501 MHz)