# Portfolio — R630 GPU Passthrough → Always-On Ollama Inference ## Problem A retired Dell R630 PowerEdge (service tag BXK1MR2) needed to become a 24/7 local LLM inference box for a fleet of self-hosted apps — instead of each app spinning up its own Ollama instance or renting cloud GPU time. ## Approach - **Hardware**: R630 with a Quadro M4000 (8GB). Expanded RAM from 64GB → 117GB, and carved out NVMe into a 32GB SLOG + 174GB L2ARC + 32.5GB swap for ZFS. - **GPU passthrough**: Installed Docker + `nvidia-toolkit`, attached the M4000 with `--gpus all`, and stood up Ollama as a systemd service (`0.0.0.0:11434`, LAN-open, `keep_alive 30m`, `max_loaded_models 1`). - **The gotcha nobody documents**: Ollama 0.34 *dropped CUDA for Maxwell* (compute 5.2 needs driver 570+, box has 550). CUDA silently fails → Ollama falls back to **Vulkan**. Diagnosed it, kept Vulkan, and benchmarked it properly instead of assuming it was dead. ## Result - `ornith-1.5:9b` (9B, Q4_K_M, 6.6GB) runs **100% GPU at 13.2 tok/s** with a 4K context — comparable to an RTX 3070 at 64K context, on a $100 retired GPU. - The box is now the **always-on consolidation lane**: every server app (Tarro, PHOTON, Signal Miner, Research Engine, any Flask/Node app with a default LLM) points at one Ollama instead of each running its own. ## Tools used Proxmox, Docker, nvidia-toolkit, Ollama, Vulkan, ZFS (SLOG/L2ARC), systemd, Dell iDRAC8. ## Why it's sellable This is the exact problem people pay to solve: "I have a GPU, why won't Ollama use it?" The Maxwell/CUDA→Vulkan migration is a real, non-obvious fix most people burn hours on.