Files
tech-skill-monetization/portfolio/01-r630-gpu-ollama.md
2026-10-06 23:43:39 -07:00

36 lines
1.6 KiB
Markdown

# Portfolio — R630 GPU Passthrough → Always-On Ollama Inference
## Problem
A retired Dell R630 PowerEdge (service tag BXK1MR2) needed to become a 24/7 local
LLM inference box for a fleet of self-hosted apps — instead of each app spinning up
its own Ollama instance or renting cloud GPU time.
## Approach
- **Hardware**: R630 with a Quadro M4000 (8GB). Expanded RAM from 64GB → 117GB, and
carved out NVMe into a 32GB SLOG + 174GB L2ARC + 32.5GB swap for ZFS.
- **GPU passthrough**: Installed Docker + `nvidia-toolkit`, attached the M4000 with
`--gpus all`, and stood up Ollama as a systemd service (`0.0.0.0:11434`, LAN-open,
`keep_alive 30m`, `max_loaded_models 1`).
- **The gotcha nobody documents**: Ollama 0.34 *dropped CUDA for Maxwell* (compute 5.2
needs driver 570+, box has 550). CUDA silently fails → Ollama falls back to **Vulkan**.
Diagnosed it, kept Vulkan, and benchmarked it properly instead of assuming it was dead.
## Result
- `ornith-1.5:9b` (9B, Q4_K_M, 6.6GB) runs **100% GPU at 13.2 tok/s** with a 4K context —
comparable to an RTX 3070 at 64K context, on a $100 retired GPU.
- The box is now the **always-on consolidation lane**: every server app (Tarro, PHOTON,
Signal Miner, Research Engine, any Flask/Node app with a default LLM) points at one
Ollama instead of each running its own.
## Tools used
Proxmox, Docker, nvidia-toolkit, Ollama, Vulkan, ZFS (SLOG/L2ARC), systemd, Dell iDRAC8.
## Why it's sellable
This is the exact problem people pay to solve: "I have a GPU, why won't Ollama use it?"
The Maxwell/CUDA→Vulkan migration is a real, non-obvious fix most people burn hours on.