1.6 KiB
1.6 KiB
Portfolio — R630 GPU Passthrough → Always-On Ollama Inference
Problem
A retired Dell R630 PowerEdge (service tag BXK1MR2) needed to become a 24/7 local LLM inference box for a fleet of self-hosted apps — instead of each app spinning up its own Ollama instance or renting cloud GPU time.
Approach
- Hardware: R630 with a Quadro M4000 (8GB). Expanded RAM from 64GB → 117GB, and carved out NVMe into a 32GB SLOG + 174GB L2ARC + 32.5GB swap for ZFS.
- GPU passthrough: Installed Docker +
nvidia-toolkit, attached the M4000 with--gpus all, and stood up Ollama as a systemd service (0.0.0.0:11434, LAN-open,keep_alive 30m,max_loaded_models 1). - The gotcha nobody documents: Ollama 0.34 dropped CUDA for Maxwell (compute 5.2 needs driver 570+, box has 550). CUDA silently fails → Ollama falls back to Vulkan. Diagnosed it, kept Vulkan, and benchmarked it properly instead of assuming it was dead.
Result
ornith-1.5:9b(9B, Q4_K_M, 6.6GB) runs 100% GPU at 13.2 tok/s with a 4K context — comparable to an RTX 3070 at 64K context, on a $100 retired GPU.- The box is now the always-on consolidation lane: every server app (Tarro, PHOTON, Signal Miner, Research Engine, any Flask/Node app with a default LLM) points at one Ollama instead of each running its own.
Tools used
Proxmox, Docker, nvidia-toolkit, Ollama, Vulkan, ZFS (SLOG/L2ARC), systemd, Dell iDRAC8.
Why it's sellable
This is the exact problem people pay to solve: "I have a GPU, why won't Ollama use it?" The Maxwell/CUDA→Vulkan migration is a real, non-obvious fix most people burn hours on.