36 lines
1.6 KiB
Markdown
36 lines
1.6 KiB
Markdown
# Portfolio — R630 GPU Passthrough → Always-On Ollama Inference
|
|
|
|
## Problem
|
|
|
|
A retired Dell R630 PowerEdge (service tag BXK1MR2) needed to become a 24/7 local
|
|
LLM inference box for a fleet of self-hosted apps — instead of each app spinning up
|
|
its own Ollama instance or renting cloud GPU time.
|
|
|
|
## Approach
|
|
|
|
- **Hardware**: R630 with a Quadro M4000 (8GB). Expanded RAM from 64GB → 117GB, and
|
|
carved out NVMe into a 32GB SLOG + 174GB L2ARC + 32.5GB swap for ZFS.
|
|
- **GPU passthrough**: Installed Docker + `nvidia-toolkit`, attached the M4000 with
|
|
`--gpus all`, and stood up Ollama as a systemd service (`0.0.0.0:11434`, LAN-open,
|
|
`keep_alive 30m`, `max_loaded_models 1`).
|
|
- **The gotcha nobody documents**: Ollama 0.34 *dropped CUDA for Maxwell* (compute 5.2
|
|
needs driver 570+, box has 550). CUDA silently fails → Ollama falls back to **Vulkan**.
|
|
Diagnosed it, kept Vulkan, and benchmarked it properly instead of assuming it was dead.
|
|
|
|
## Result
|
|
|
|
- `ornith-1.5:9b` (9B, Q4_K_M, 6.6GB) runs **100% GPU at 13.2 tok/s** with a 4K context —
|
|
comparable to an RTX 3070 at 64K context, on a $100 retired GPU.
|
|
- The box is now the **always-on consolidation lane**: every server app (Tarro, PHOTON,
|
|
Signal Miner, Research Engine, any Flask/Node app with a default LLM) points at one
|
|
Ollama instead of each running its own.
|
|
|
|
## Tools used
|
|
|
|
Proxmox, Docker, nvidia-toolkit, Ollama, Vulkan, ZFS (SLOG/L2ARC), systemd, Dell iDRAC8.
|
|
|
|
## Why it's sellable
|
|
|
|
This is the exact problem people pay to solve: "I have a GPU, why won't Ollama use it?"
|
|
The Maxwell/CUDA→Vulkan migration is a real, non-obvious fix most people burn hours on.
|