HYPERSWAP

v1.0-DEPLOY

NVIDIA RTX 4080 SUPER 16GB // 64GB DDR5 RAM // Ubuntu Linux

Ollama LLM: ONLINE
ComfyUI: ONLINE
SSE 1Hz

GPU VRAM (16 GB Dedicated)

Real-time allocation breakdown on RTX 4080 SUPER

0.0 / 16.0 GB
0% USED
Ollama: 0 GB
ComfyUI: 0 GB
System: 0 GB
Ollama: 0 GB
Comfy: 0 GB
System: 0 GB
Free: 0 GB

Host RAM & Model Page Cache (64 GB)

Models stay resident in RAM for instant PCIe hot-swaps

0.0 GB CACHED
60.3 GB Total
Apps Used: 0 GB
Models in RAM: 0 GB
Free RAM: 0 GB

Ollama LLM Engine

Port :11434 // FlashAttention + Q4 KV Cache

Active Model in VRAM 0.0 GB VRAM
None Loaded
0 ctx
LAST SWAP TIME -- ms
RAM HIT STATUS --
TOTAL MODELS 0

ComfyUI Diffusion Engine

Port :8188 // DynamicVRAM + Pinned Async Offload

Pipeline Status IDLE / READY
Dynamic Model Offloader
Queue: 0
Host Pinned Memory: 53.6 GB Staging Buffer
Async PCIe Offloading: Enabled (2 Streams)
Fast Disk RAM Mmap: Active
DISCOVERED MODELS 0 Files
VRAM AVAILABLE 15.9 GB

GPU Live Telemetry (NVML)

NVIDIA GeForce RTX 4080 SUPER

GPU Util 0%
Temperature 0°C
Power Draw 0 W
Fan Speed
0%
Fan 0: --% | Fan 1: --%
Active Compute Processes
PID Process VRAM
Scanning GPU processes...

Model Switch Timeline & Optimizer

Real-time latency logger and memory warmer

Recent Model Swaps
No recent model swaps recorded yet.
64GB RAM Cache holds all models in memory
PCIe x16 (~31.5 GB/s)

GPU Overclock Control

Per-app profiles auto-switch with the VRAM arbitrator

Active: -- X: --
Power Limit -- W
Core Clock --
Mem Clock --
Temp / Draw --
GPU Fan Cooling Control Dual-fan PWM active speed regulation
AUTO (VBIOS)
Fan 0 (Intake/Core): --%
Fan 1 (Exhaust/VRM): --%
Custom Manual Target: 65%
GPU --
Driver --
VRAM --
Max Core --
Max Mem --
Power --
Live Tuning Graph
Core MHz Mem MHz Temp °C Power W Fan % VRAM GB RAM Cache GB
Fine-tune profile
Lock core to max boost (2900–3105 MHz)
Lock mem to max (11501 MHz)

VRAM Arbitration

Who holds the GPU, and how handoffs are going

—
Idle
0
Released
0
Deferred
(busy)
0
Landed
later
0
Stalled
0
Comfy
purges
0
Purges
deferred

Deferred is healthy — an LLM mid-generation cannot unload, so the request queues and applies the moment it finishes. Only stalled (VRAM held while the GPU sits idle) indicates a real problem.

Thermal Governor

Walks the overclock back when the card complains

full level 0 / 3
Escalate
83°C
Recover
72°C
Offset Scale
100%

Last action: cold start

Measured Page-Cache Residency

What is genuinely in RAM, not what we hope is

— GB —

Is the Overclock Actually Working?

Decode throughput per profile, from persisted history

Collecting data…

Overclock Autotune

Sweep a clock offset, measure tok/s, stop at instability

idle
Restores the profile when done, even on error.