Files
tech-skill-monetization/portfolio/01-r630-gpu-ollama.md
2026-10-06 23:43:39 -07:00

1.6 KiB

Portfolio — R630 GPU Passthrough → Always-On Ollama Inference

Problem

A retired Dell R630 PowerEdge (service tag BXK1MR2) needed to become a 24/7 local LLM inference box for a fleet of self-hosted apps — instead of each app spinning up its own Ollama instance or renting cloud GPU time.

Approach

  • Hardware: R630 with a Quadro M4000 (8GB). Expanded RAM from 64GB → 117GB, and carved out NVMe into a 32GB SLOG + 174GB L2ARC + 32.5GB swap for ZFS.
  • GPU passthrough: Installed Docker + nvidia-toolkit, attached the M4000 with --gpus all, and stood up Ollama as a systemd service (0.0.0.0:11434, LAN-open, keep_alive 30m, max_loaded_models 1).
  • The gotcha nobody documents: Ollama 0.34 dropped CUDA for Maxwell (compute 5.2 needs driver 570+, box has 550). CUDA silently fails → Ollama falls back to Vulkan. Diagnosed it, kept Vulkan, and benchmarked it properly instead of assuming it was dead.

Result

  • ornith-1.5:9b (9B, Q4_K_M, 6.6GB) runs 100% GPU at 13.2 tok/s with a 4K context — comparable to an RTX 3070 at 64K context, on a $100 retired GPU.
  • The box is now the always-on consolidation lane: every server app (Tarro, PHOTON, Signal Miner, Research Engine, any Flask/Node app with a default LLM) points at one Ollama instead of each running its own.

Tools used

Proxmox, Docker, nvidia-toolkit, Ollama, Vulkan, ZFS (SLOG/L2ARC), systemd, Dell iDRAC8.

Why it's sellable

This is the exact problem people pay to solve: "I have a GPU, why won't Ollama use it?" The Maxwell/CUDA→Vulkan migration is a real, non-obvious fix most people burn hours on.