Snapshot: full project state
This commit is contained in:
40
portfolio/02-multihost-gpu-routing.md
Normal file
40
portfolio/02-multihost-gpu-routing.md
Normal file
@@ -0,0 +1,40 @@
|
||||
# Portfolio — 4-Lane Multi-GPU LLM Routing Architecture
|
||||
|
||||
## Problem
|
||||
|
||||
Three GPU machines (an RTX 4080 SUPER, an RTX 3070, and a Quadro M4000) plus a MacBook
|
||||
were being used ad hoc, with every app guessing where to send its LLM calls. Result:
|
||||
GPU saturation, cold-model timeouts, and latency-critical requests queued behind batch work.
|
||||
|
||||
## Approach
|
||||
|
||||
Designed a **4-lane routing architecture** with a single model standard (`ornith-1.5:9b`)
|
||||
and an explicit per-lane map:
|
||||
|
||||
| Lane | Host | GPU | Job |
|
||||
|---|---|---|---|
|
||||
| **premium** | RTX 4080 SUPER (16GB) | CUDA | mission-critical speed only |
|
||||
| **batch** | RTX 3070 (8GB, WiFi) | CUDA | stateful, long-context, vision/OCR/embeddings |
|
||||
| **permanent** | Quadro M4000 (8GB) | Vulkan | always-on consolidation for server apps |
|
||||
| **voice** | MacBook (unified) | Metal/MLX | local voice (gemma3:4b) |
|
||||
|
||||
## Result
|
||||
|
||||
- **100K-token context on 16GB VRAM with zero CPU offload** — baked `num_ctx 102400`,
|
||||
`num_gpu 99`, `KV_CACHE_TYPE q4_0`, and `NUM_BATCH 2048` to push prefill from 308 →
|
||||
**1867 tok/s** and decode to ~41 tok/s on the 4080.
|
||||
- Single shared GPU for Ollama + ComfyUI (image gen) via HyperSwap (program swapper),
|
||||
so text inference and diffusion don't fight over VRAM.
|
||||
- Every consumer (TITAN, Astraea, WorkBrain, PHOTON, Honcho, 90+ cron jobs) mapped to the
|
||||
lane that fits its latency/VRAM class.
|
||||
|
||||
## Tools used
|
||||
|
||||
Ollama (multi-host), CUDA + Vulkan, Flash Attention, KV-cache quantization, HyperSwap,
|
||||
ComfyUI, systemd, LAN routing.
|
||||
|
||||
## Why it's sellable
|
||||
|
||||
"Run multiple LLMs across multiple GPUs without paying a cloud provider" is a real,
|
||||
growing ask. The 100K-context-on-16GB result is a concrete, quantifiable win most
|
||||
consultants can't show.
|
||||
Reference in New Issue
Block a user