# Vision Model Benchmark — Android Control via MCP **Date:** Aug 13, 2026 · **Bench host:** shadow-death (10.30.20.128, RTX 4080 SUPER 16GB, Ollama 0.32.9) **Target:** tiny vision model for GamingPC (10.30.20.186, RTX 3070 8GB) to drive Android via ADB/MCP **Device:** Android Cloud VM1301 (Android-x86 9.0, 1024x768, `Android.local:5555`) ## Method 8 test situations captured from a real Android device with **uiautomator XML ground truth**: launcher dialog, Chrome welcome, home screen, Settings, Chrome/Wikipedia article, Calculator, Settings search, app drawer. Each model ran 4 tasks per situation (identify, describe, **ground coordinates**, agent JSON) = 320 scored calls. Grounding score = pixel-distance accuracy against the real element center from the XML dump. No vibes — every score is measured. ## Results | Model | Size | identify | describe | **ground** | agent JSON | tok/s | VRAM (bench) | |---|---|---|---|---|---|---|---| | moondream:1.8b | 1.7G | 0.12 | 0.21 | 0.00 | 0.00 | 331 | 1.2G | | granite3.2-vision:2b | 1.7G | 0.62 | 0.04 | 0.00 | 0.38 | 223 | 3.1G | | **qwen2.5vl:3b** | **2.6G** | **1.00** | **0.55** | **0.57 (0.91 visible)** | **0.62** | **219** | 8.2G* | | llava-phi3:3.8b | 2.9G | 0.62 | 0.14 | 0.27 | 0.00 | 196 | 4.0G | | gemma3:4b | 3.3G | 1.00 | 0.40 | 0.30 | 0.38 | 161 | 3.0G | | llava:7b | 4.7G | 0.62 | 0.10 | 0.03 | 0.62 | 145 | 8.7G | | qwen2.5vl:7b | 5.7G | 0.88 | 0.52 | 0.57 (0.91 visible) | 0.38 | 129 | 12.7G | | minicpm-v4.5:8b | 6.1G | 1.00 | 0.64 | 0.54 (0.81 visible) | 0.12 | 118 | 11.1G | *bench VRAM includes default 4096 context. On GamingPC with num_ctx=2048: **3.2GB total, fully on GPU.** **Grounding on screens where the target was actually visible** (dialog/chrome/settings/article/search): | Model | avg | per-screen | |---|---|---| | qwen2.5vl:3b | **0.91** | .90 / .93 / .85 / .94 / .93 | | qwen2.5vl:7b | **0.91** | .90 / .94 / .84 / .94 / .93 | | minicpm-v4.5 | 0.81 | .50 / .75 / .93 / .90 / .95 | | llava-phi3 | 0.43 | inconsistent (0/.32/0/.86/.98) | | gemma3:4b | 0.27 | partial only | | llava:7b / moondream / granite3.2-vision | ~0 | unusable for tapping | ## Winner: qwen2.5vl:3b - **Grounding: 0.91** — identical to its 7b big brother at half the size and 40% more speed - 1.00 app identification, 219 tok/s, warm answers ~0.5–1.5s - On GamingPC 8GB: **3.2GB VRAM** (num_ctx 2048), fits alongside granite4.1:3b keep-warm - Live tap test on calculator: 7→(59,310) vs true (73,357) — 20–50px accuracy, hits the buttons ## Control-loop demo (GamingPC → VM1301) Vision-only loop for "type 7+8 and press =": model tapped **7 → 8 → + → =** in the correct sequence with ~20px precision. Multi-step autonomy flaked afterward (pressed back, looped on 7/8) because a stateless single-screenshot loop can't self-verify. **Fix that works:** keep step history in the prompt + let the agent verify with uiautomator (exact bounds) after every tap. Production architecture = Hermes agent decides → qwen2.5vl:3b sees/grounds → android-adb MCP executes → uiautomator dump verifies. ## Model verdicts (the interesting ones) - **qwen2.5vl:7b** — no grounding advantage over the 3b. Skip on 8GB cards. - **minicpm-v4.5:8b** — best descriptions (0.64) + good grounding (0.81), but 6.1GB, 118 tok/s, and its JSON output wanders (multi-action plans, its own schema). Good fallback. - **minicpm-v4.6** (1.6GB, on GamingPC) — **thinking-only variant: content is always empty.** Unusable for direct control. - **granite3.2-vision:2b** — returns *tool calls* (`getpixel`) instead of answers. Wrong fit for prompt-based control, interesting for tool-use harnesses. - **llava:7b / moondream / llava-phi3** — grounding garbage. Don't bother. - **llama3.2-vision:11b** — **fails to load on Ollama 0.32.9** (`unknown model architecture: 'mllama'`). Not available on this stack. ## Pitfalls hit (all fixed) 1. **Android screen sleep → black screencaps** (4.9KB). Fix: `svc power stayon true`. 2. **VM1301 IP drifts** — kernel `ip=.26` is overridden by Android's EthernetService DHCP (router gave .79). Fix: `adb connect Android.local:5555` (mDNS) — DHCP-proof. android-adb MCP now auto-discovers: env → online device → Android.local → legacy .84. 3. **Stale adb daemon** → "No route to host" despite TCP working. Fix: `adb kill-server`. 4. **Markdown fences around model JSON** — strict `json.loads` scores 0; fence-tolerant parsing recovers 0.0 → 0.62 on agent task. 5. **qwen2.5vl outputs** either `{"x": int}` or `{"coordinate": [x,y]}` or bbox lists — harness must normalize all three. 6. **urllib routed LAN Ollama through a system proxy** → 404s. Fix: `ProxyHandler({})`. 7. Realme phone (10.30.20.104) and UGOOS TV (.84) both offline — VM1301 is the live target. ## Deployment (done) - `qwen2.5vl:3b` + `qwen2.5vl:7b` installed on GamingPC (10.30.20.186:11434) - ADB MCP patched for device auto-discovery; verified working against VM1301 - Recommended call: `num_ctx=2048`, `temperature=0`, ask for `{"coordinate":[x,y]}` JSON - Demo harness: `~/vision_bench/android_control_demo.py ` - Raw data: `~/vision_bench/bench_results.jsonl` · analysis: `analyze.py`, `final_scores.py` ## Recommended follow-ups - Pin qwen2.5vl:3b in `keep_ollama_hot.py` on .186 (vision goto, next to ornith) - Wire Hermes auxiliary vision to qwen2.5vl:3b@.186 (config.yaml auxiliary.vision) - Add step-history + uiautomator verification to the control harness for multi-step tasks - Replace dead UGOOS TV .84 default references; keep mDNS path as primary