self-hosted AI · homelab · hardware

I get local AI running on
hardware you own.

GPU passthrough, multi-host Ollama routing, and self-hosted LLMs that never leave your network. If your GPU won't cooperate with Ollama, I can fix it.

4
GPU lanes routed
100K
ctx on 16GB VRAM
13.2
tok/s on a $100 GPU
3
hosts, one model std
01 · Selected work

Proof, not promises.

Real builds from real hardware. Each one solves the exact problem people pay to have solved.

R630 → Always-On Ollama Inference (GPU Passthrough)

Dell R630 · Quadro M4000 · Vulkan fallback · 13.2 tok/s
Problem
A retired R630 needed to become a 24/7 local LLM box for a fleet of apps — instead of each app spinning up its own Ollama or renting cloud GPU time.
Approach
Docker + nvidia-toolkit, Ollama as a systemd service (LAN-open, keep_alive, model pinning). Diagnosed the undocumented gotcha: Ollama 0.34 dropped CUDA for Maxwell, so it silently fell back to Vulkan — benchmarked it properly instead of declaring it dead.
Result
ornith-1.5:9b at 100% GPU, 13.2 tok/s on a $100 retired card — the always-on consolidation lane every server app now points at.
ProxmoxDockerOllamaVulkanZFSsystemdiDRAC

4-Lane Multi-GPU LLM Routing Architecture

RTX 4080 SUPER · RTX 3070 · M4000 · MacBook · one model standard
Problem
Three GPUs + a MacBook were being used ad hoc — every app guessed where to send its LLM call, causing GPU saturation, cold-model timeouts, and latency-critical jobs queued behind batch work.
Approach
Four explicit lanes (premium / batch / always-on / voice) with a single model standard and per-lane latency/VRAM routing. Shared one GPU between Ollama and ComfyUI via HyperSwap.
Result
100K-token context on 16GB VRAM with zero CPU offload; prefill pushed 308 → 1867 tok/s. Every consumer mapped to the lane that fits its class.
Ollama multi-hostCUDAVulkanFlash AttentionKV-cache quantHyperSwapComfyUI

Android TV Box Rooting + Emulation (SK4 Pro)

UGOOS SK4 Pro · Amlogic · Android 14 · Magisk · NetherSX2
Problem
Turn an Android TV box into a retro-emulation console (N64/PS1/PS2) — which required root, BIOS/ROM staging, and core config, all below the app layer.
Approach
Pushed Magisk 30.7, patched the A/B boot image, staged PS1/PS2 BIOS + ROMs, configured RetroArch cores, and managed it over ADB + SMB.
Result
A rooted, emulation-ready TV box with clean ROM/BIOS layout, side-loaded Jellyfin, managed fully remote over the LAN.
Amlogic A/BMagiskADBRetroArchNetherSX2SambaESP32-S3
02 · What I do

Services & rates.

Self-hosted AI / GPU passthrough

$75/hr

Local LLM on your own hardware — GPU passthrough, Ollama setup, multi-host routing, context/VRAM optimization.

Homelab & sysadmin

$65/hr

Proxmox, Docker, ZFS, backups (PBS), networking. Turn a pile of servers into one clean infrastructure.

Hardware / embedded

$65/hr

Android TV rooting, ESP32 firmware, ADB automation, boot-image patching — the below-the-OS-layer stuff.

Defensive security review

$95/hr

Authorized vulnerability assessment and app-security review. Scoped, documented, lawful.

03 · Contact

Have a GPU that won't cooperate?

Tell me your hardware and what's stuck. I'll tell you if it's a 2-command fix or a real job.

Email me