Snapshot: full project state
This commit is contained in:
35
portfolio/01-r630-gpu-ollama.md
Normal file
35
portfolio/01-r630-gpu-ollama.md
Normal file
@@ -0,0 +1,35 @@
|
||||
# Portfolio — R630 GPU Passthrough → Always-On Ollama Inference
|
||||
|
||||
## Problem
|
||||
|
||||
A retired Dell R630 PowerEdge (service tag BXK1MR2) needed to become a 24/7 local
|
||||
LLM inference box for a fleet of self-hosted apps — instead of each app spinning up
|
||||
its own Ollama instance or renting cloud GPU time.
|
||||
|
||||
## Approach
|
||||
|
||||
- **Hardware**: R630 with a Quadro M4000 (8GB). Expanded RAM from 64GB → 117GB, and
|
||||
carved out NVMe into a 32GB SLOG + 174GB L2ARC + 32.5GB swap for ZFS.
|
||||
- **GPU passthrough**: Installed Docker + `nvidia-toolkit`, attached the M4000 with
|
||||
`--gpus all`, and stood up Ollama as a systemd service (`0.0.0.0:11434`, LAN-open,
|
||||
`keep_alive 30m`, `max_loaded_models 1`).
|
||||
- **The gotcha nobody documents**: Ollama 0.34 *dropped CUDA for Maxwell* (compute 5.2
|
||||
needs driver 570+, box has 550). CUDA silently fails → Ollama falls back to **Vulkan**.
|
||||
Diagnosed it, kept Vulkan, and benchmarked it properly instead of assuming it was dead.
|
||||
|
||||
## Result
|
||||
|
||||
- `ornith-1.5:9b` (9B, Q4_K_M, 6.6GB) runs **100% GPU at 13.2 tok/s** with a 4K context —
|
||||
comparable to an RTX 3070 at 64K context, on a $100 retired GPU.
|
||||
- The box is now the **always-on consolidation lane**: every server app (Tarro, PHOTON,
|
||||
Signal Miner, Research Engine, any Flask/Node app with a default LLM) points at one
|
||||
Ollama instead of each running its own.
|
||||
|
||||
## Tools used
|
||||
|
||||
Proxmox, Docker, nvidia-toolkit, Ollama, Vulkan, ZFS (SLOG/L2ARC), systemd, Dell iDRAC8.
|
||||
|
||||
## Why it's sellable
|
||||
|
||||
This is the exact problem people pay to solve: "I have a GPU, why won't Ollama use it?"
|
||||
The Maxwell/CUDA→Vulkan migration is a real, non-obvious fix most people burn hours on.
|
||||
40
portfolio/02-multihost-gpu-routing.md
Normal file
40
portfolio/02-multihost-gpu-routing.md
Normal file
@@ -0,0 +1,40 @@
|
||||
# Portfolio — 4-Lane Multi-GPU LLM Routing Architecture
|
||||
|
||||
## Problem
|
||||
|
||||
Three GPU machines (an RTX 4080 SUPER, an RTX 3070, and a Quadro M4000) plus a MacBook
|
||||
were being used ad hoc, with every app guessing where to send its LLM calls. Result:
|
||||
GPU saturation, cold-model timeouts, and latency-critical requests queued behind batch work.
|
||||
|
||||
## Approach
|
||||
|
||||
Designed a **4-lane routing architecture** with a single model standard (`ornith-1.5:9b`)
|
||||
and an explicit per-lane map:
|
||||
|
||||
| Lane | Host | GPU | Job |
|
||||
|---|---|---|---|
|
||||
| **premium** | RTX 4080 SUPER (16GB) | CUDA | mission-critical speed only |
|
||||
| **batch** | RTX 3070 (8GB, WiFi) | CUDA | stateful, long-context, vision/OCR/embeddings |
|
||||
| **permanent** | Quadro M4000 (8GB) | Vulkan | always-on consolidation for server apps |
|
||||
| **voice** | MacBook (unified) | Metal/MLX | local voice (gemma3:4b) |
|
||||
|
||||
## Result
|
||||
|
||||
- **100K-token context on 16GB VRAM with zero CPU offload** — baked `num_ctx 102400`,
|
||||
`num_gpu 99`, `KV_CACHE_TYPE q4_0`, and `NUM_BATCH 2048` to push prefill from 308 →
|
||||
**1867 tok/s** and decode to ~41 tok/s on the 4080.
|
||||
- Single shared GPU for Ollama + ComfyUI (image gen) via HyperSwap (program swapper),
|
||||
so text inference and diffusion don't fight over VRAM.
|
||||
- Every consumer (TITAN, Astraea, WorkBrain, PHOTON, Honcho, 90+ cron jobs) mapped to the
|
||||
lane that fits its latency/VRAM class.
|
||||
|
||||
## Tools used
|
||||
|
||||
Ollama (multi-host), CUDA + Vulkan, Flash Attention, KV-cache quantization, HyperSwap,
|
||||
ComfyUI, systemd, LAN routing.
|
||||
|
||||
## Why it's sellable
|
||||
|
||||
"Run multiple LLMs across multiple GPUs without paying a cloud provider" is a real,
|
||||
growing ask. The 100K-context-on-16GB result is a concrete, quantifiable win most
|
||||
consultants can't show.
|
||||
33
portfolio/03-sk4pro-android-tv.md
Normal file
33
portfolio/03-sk4pro-android-tv.md
Normal file
@@ -0,0 +1,33 @@
|
||||
# Portfolio — Android TV Box Rooting + Emulation (SK4 Pro)
|
||||
|
||||
## Problem
|
||||
|
||||
Turn a UGOOS SK4 Pro Android TV box (Amlogic, Android 14) into a retro-emulation
|
||||
console — including PS2 emulation via NetherSX2 — which required root, BIOS/ROM
|
||||
staging, and core configuration.
|
||||
|
||||
## Approach
|
||||
|
||||
- **Device**: SK4 Pro, Amlogic SoC, Android 14 with A/B partitions.
|
||||
- **Root**: Pushed Magisk 30.7, generated a `magisk_patched` boot image, and staged the
|
||||
patched `init_boot` partition for the A/B slot.
|
||||
- **Emulation**: Configured RetroArch PlayStation cores (needs BIOS + ROMs), staged
|
||||
`scph5501.bin`-class PS1 BIOS and `.bin/.cue`/`.chd` ROMs, and prepped NetherSX2 (PS2)
|
||||
with its required PS2 BIOS set (SCPH-39001/70012/77001/90001).
|
||||
- **Access**: ADB over network (port 5555) + SMB Samba share into `/sdcard` for ROM
|
||||
management when the ADB device was still in "unauthorized" state.
|
||||
|
||||
## Result
|
||||
|
||||
A rooted, emulation-ready TV box with N64/PS1/PS2 paths staged, Jellyfin side-loaded for
|
||||
media, and a clean ROM/BIOS layout — all managed remotely over the LAN.
|
||||
|
||||
## Tools used
|
||||
|
||||
Amlogic A/B flashing, Magisk, ADB, RetroArch, NetherSX2, SMB/Samba, Jellyfin.
|
||||
|
||||
## Why it's sellable
|
||||
|
||||
This is the "hacker" proof — real firmware/boot-image work on Android hardware, not just
|
||||
app installs. Same skillset applies to ESP32-S3 embedded firmware (I also built a voice
|
||||
assistant on ESP32-S3). Demonstrates I can go below the OS layer when the job needs it.
|
||||
Reference in New Issue
Block a user