Snapshot: full project state

This commit is contained in:
2026-10-06 23:43:39 -07:00
commit 0ec1dedf31
12 changed files with 1250 additions and 0 deletions

93
outreach.md Normal file
View File

@@ -0,0 +1,93 @@
# Outreach — Live Targets + Drafted Messages + Tracking Log
## Method
Find real people stuck on the exact problem I've solved (Ollama/GPU passthrough / local
LLM infra), and offer a specific fix — never a template. Reddit help threads first (zero
cost, genuine need), then Upwork job posts.
---
## Live targets (found, verified)
### Reddit threads — "Ollama won't use GPU" (direct fit)
1. **r/ollama — "No GPU utilization"**
`https://www.reddit.com/r/ollama/comments/1t4srec/no_gpu_utilization/`
→ The user's GPU isn't being used. First step: `ollama ps` shows PROCESSOR=CPU.
2. **r/LocalLLaMA — "Ollama not using GPU, need help"**
`https://www.reddit.com/r/LocalLLaMA/comments/1jw5m8k/ollama_not_using_gpu_need_help/`
3. **r/LocalLLaMA — "Proxmox and LXC Passthrough for Ollama Best Practices?"**
`https://www.reddit.com/r/LocalLLaMA/comments/1irmk5n/proxmox_and_lxc_passthrough_for_ollama_best/`
→ Direct passthrough-in-LXC question. My exact lane.
4. **r/LocalLLaMA — "Need help setting up a local LLM server with RTX 3060"**
`https://www.reddit.com/r/LocalLLaMA/comments/1lvm3kv/need_help_setting_up_a_local_llm_server_with_rtx/`
### Upwork — Proxmox/AI infra (hire intent)
5. Upwork "Proxmox VE Specialists for Hire" landing page → browse open GPU passthrough /
homelab / self-hosted AI jobs, apply only to the specific-fit ones.
`https://www.upwork.com/hire/proxmox-ve-freelancers/`
---
## Drafted messages (specific, non-templated — reference THEIR problem)
### Target 1 — r/ollama "No GPU utilization"
> Before you reinstall anything — run `ollama ps` while a model is loaded and check the
> PROCESSOR column. If it says CPU, that's the real signal. Then `nvidia-smi` in a second
> terminal while it's inferring: if Ollama doesn't appear as a GPU process, the driver
> isn't the issue — the model is too big for VRAM and it's spilling to CPU, or you're on
> a card where CUDA silently fell back to Vulkan (I hit exactly this on a Quadro M4000
> after Ollama 0.34 dropped Maxwell CUDA). What GPU + what model are you running? I can
> tell you which it is from those two commands.
*(Value: gives the diagnostic, demonstrates the exact non-obvious Maxwell/Vulkan fix.)*
### Target 3 — r/LocalLLaMA "Proxmox and LXC Passthrough for Ollama"
> The LXC vs full VM passthrough question comes down to whether you need it exposed to
> ONE container or shared. For a single always-on Ollama container, device passthrough
> (`/dev/nvidia*`) + `nvidia-container-toolkit` in the LXC beats full VM passthrough —
> you keep snapshot/backup and don't lose the card to a VM. I run a 4-node setup (4080
> CUDA lane + 3070 batch lane + M4000 Vulkan always-on lane + MacBook voice) — happy to
> share the exact systemd/env config if you say which card you're passing through.
### Target 4 — r/LocalLLaMA "RTX 3060 local LLM server"
> 3060 = 12GB, that's a solid self-host box. The trap is default Ollama will spill a
> 13B+ model to CPU and you'll think it's broken. Pin a 7–9B Q4 model with `num_gpu 99`
> and `OLLAMA_KV_CACHE_TYPE q4_0` and you'll stay 100% GPU. I did 100K context on 16GB
> with zero offload using Flash Attention + KV quant — the 3060 will do ~24–32K ctx
> comfortably. Want the exact override.conf I use?
---
## Tracking log
| # | Platform | Target | Status | Follow-up date |
|---|----------|--------|--------|----------------|
| 1 | Reddit | r/ollama "No GPU utilization" | ⬜ draft ready | — |
| 2 | Reddit | r/LocalLLaMA "not using GPU" | ⬜ draft ready | — |
| 3 | Reddit | r/LocalLLaMA "LXC passthrough" | ⬜ draft ready | — |
| 4 | Reddit | r/LocalLLaMA "RTX 3060 server" | ⬜ draft ready | — |
| 5 | Upwork | Proxmox/AI infra jobs | ⬜ browse + apply | — |
**Rules:**
- Quality over volume — 5 targeted > 50 generic.
- No unpaid "exposure" work. A quick free *diagnostic* (2 commands) is fine — it earns
trust; a free *build* is not.
- Send from YOUR real Reddit account, human-paced (max 1–2/day at first). Mass DMs from a
fresh account = shadowban.
- Log every response in this table (status: ⬜ sent / 🟡 replied / ✅ gig / ❌ dead).
---
## Content flywheel (optional, after first win)
Turn the R630 GPU→Ollama build into a YouTube/blog tutorial, cross-linking the Gumroad
guide + Upwork profile. The Maxwell/Vulkan gotcha is genuinely novel enough to rank.