Files
tech-skill-monetization/portfolio/02-multihost-gpu-routing.md
2026-10-06 23:43:39 -07:00

1.7 KiB

Portfolio — 4-Lane Multi-GPU LLM Routing Architecture

Problem

Three GPU machines (an RTX 4080 SUPER, an RTX 3070, and a Quadro M4000) plus a MacBook were being used ad hoc, with every app guessing where to send its LLM calls. Result: GPU saturation, cold-model timeouts, and latency-critical requests queued behind batch work.

Approach

Designed a 4-lane routing architecture with a single model standard (ornith-1.5:9b) and an explicit per-lane map:

Lane Host GPU Job
premium RTX 4080 SUPER (16GB) CUDA mission-critical speed only
batch RTX 3070 (8GB, WiFi) CUDA stateful, long-context, vision/OCR/embeddings
permanent Quadro M4000 (8GB) Vulkan always-on consolidation for server apps
voice MacBook (unified) Metal/MLX local voice (gemma3:4b)

Result

  • 100K-token context on 16GB VRAM with zero CPU offload — baked num_ctx 102400, num_gpu 99, KV_CACHE_TYPE q4_0, and NUM_BATCH 2048 to push prefill from 308 → 1867 tok/s and decode to ~41 tok/s on the 4080.
  • Single shared GPU for Ollama + ComfyUI (image gen) via HyperSwap (program swapper), so text inference and diffusion don't fight over VRAM.
  • Every consumer (TITAN, Astraea, WorkBrain, PHOTON, Honcho, 90+ cron jobs) mapped to the lane that fits its latency/VRAM class.

Tools used

Ollama (multi-host), CUDA + Vulkan, Flash Attention, KV-cache quantization, HyperSwap, ComfyUI, systemd, LAN routing.

Why it's sellable

"Run multiple LLMs across multiple GPUs without paying a cloud provider" is a real, growing ask. The 100K-context-on-16GB result is a concrete, quantifiable win most consultants can't show.