HYPERSWAP

v1.0-DEPLOY

NVIDIA RTX 4080 SUPER 16GB // 64GB DDR5 RAM // Ubuntu Linux

Ollama LLM: ONLINE
ComfyUI: ONLINE
SSE 1Hz

GPU VRAM (16 GB Dedicated)

Real-time allocation breakdown on RTX 4080 SUPER

0.0 / 16.0 GB
0% USED
Ollama: 0 GB
ComfyUI: 0 GB
System: 0 GB
Ollama: 0 GB
Comfy: 0 GB
System: 0 GB
Free: 0 GB

Host RAM & Model Page Cache (64 GB)

Models stay resident in RAM for instant PCIe hot-swaps

0.0 GB CACHED
60.3 GB Total
Apps Used: 0 GB
Models in RAM: 0 GB
Free RAM: 0 GB

Ollama LLM Engine

Port :11434 // FlashAttention + Q4 KV Cache

Active Model in VRAM 0.0 GB VRAM
None Loaded
0 ctx
LAST SWAP TIME -- ms
RAM HIT STATUS --
TOTAL MODELS 0

ComfyUI Diffusion Engine

Port :8188 // DynamicVRAM + Pinned Async Offload

Pipeline Status IDLE / READY
Dynamic Model Offloader
Queue: 0
Host Pinned Memory: 53.6 GB Staging Buffer
Async PCIe Offloading: Enabled (2 Streams)
Fast Disk RAM Mmap: Active
DISCOVERED MODELS 0 Files
VRAM AVAILABLE 15.9 GB

GPU Live Telemetry (NVML)

NVIDIA GeForce RTX 4080 SUPER

GPU Util 0%
Temperature 0°C
Power Draw 0 W
Fan Speed
0%
Fan 0: --% | Fan 1: --%
Active Compute Processes
PID Process VRAM
Scanning GPU processes...

Model Switch Timeline & Optimizer

Real-time latency logger and memory warmer

Recent Model Swaps
No recent model swaps recorded yet.
64GB RAM Cache holds all models in memory
PCIe x16 (~31.5 GB/s)

GPU Overclock Control

Per-app profiles auto-switch with the VRAM arbitrator

Active: -- X: --
Power Limit -- W
Core Clock --
Mem Clock --
Temp / Draw --
GPU Fan Cooling Control Dual-fan PWM active speed regulation
AUTO (VBIOS)
Fan 0 (Intake/Core): --%
Fan 1 (Exhaust/VRM): --%
Custom Manual Target: 65%
GPU --
Driver --
VRAM --
Max Core --
Max Mem --
Power --
Live Tuning Graph
Core MHz Mem MHz Temp °C Power W Fan % VRAM GB RAM Cache GB
Fine-tune profile
Lock core to max boost (2900–3105 MHz)
Lock mem to max (11501 MHz)

System Health

Every dependency, with impact and how to fix it

checking…

VRAM Arbitration

Who holds the GPU, and how handoffs are going

—
Idle
0
Released
0
Deferred
(busy)
0
Landed
later
0
Stalled
0
Comfy
purges
0
Purges
deferred

Deferred is healthy — an LLM mid-generation cannot unload, so the request queues and applies the moment it finishes. Only stalled (VRAM held while the GPU sits idle) indicates a real problem.

Thermal Governor

Walks the overclock back when the card complains

full level 0 / 3
Escalate
83°C
Recover
72°C
Offset Scale
100%

Last action: cold start

Measured Page-Cache Residency

What is genuinely in RAM, not what we hope is

— GB —

Is the Overclock Actually Working?

Decode throughput per profile, from persisted history

Collecting data…

Overclock Autotune

Sweep a clock offset, measure tok/s, stop at instability

idle
Restores the profile when done, even on error.