5.5 KiB
Vision Model Benchmark — Android Control via MCP
Date: Aug 13, 2026 · Bench host: shadow-death (10.30.20.128, RTX 4080 SUPER 16GB, Ollama 0.32.9)
Target: tiny vision model for GamingPC (10.30.20.186, RTX 3070 8GB) to drive Android via ADB/MCP
Device: Android Cloud VM1301 (Android-x86 9.0, 1024x768, Android.local:5555)
Method
8 test situations captured from a real Android device with uiautomator XML ground truth: launcher dialog, Chrome welcome, home screen, Settings, Chrome/Wikipedia article, Calculator, Settings search, app drawer. Each model ran 4 tasks per situation (identify, describe, ground coordinates, agent JSON) = 320 scored calls. Grounding score = pixel-distance accuracy against the real element center from the XML dump. No vibes — every score is measured.
Results
| Model | Size | identify | describe | ground | agent JSON | tok/s | VRAM (bench) |
|---|---|---|---|---|---|---|---|
| moondream:1.8b | 1.7G | 0.12 | 0.21 | 0.00 | 0.00 | 331 | 1.2G |
| granite3.2-vision:2b | 1.7G | 0.62 | 0.04 | 0.00 | 0.38 | 223 | 3.1G |
| qwen2.5vl:3b | 2.6G | 1.00 | 0.55 | 0.57 (0.91 visible) | 0.62 | 219 | 8.2G* |
| llava-phi3:3.8b | 2.9G | 0.62 | 0.14 | 0.27 | 0.00 | 196 | 4.0G |
| gemma3:4b | 3.3G | 1.00 | 0.40 | 0.30 | 0.38 | 161 | 3.0G |
| llava:7b | 4.7G | 0.62 | 0.10 | 0.03 | 0.62 | 145 | 8.7G |
| qwen2.5vl:7b | 5.7G | 0.88 | 0.52 | 0.57 (0.91 visible) | 0.38 | 129 | 12.7G |
| minicpm-v4.5:8b | 6.1G | 1.00 | 0.64 | 0.54 (0.81 visible) | 0.12 | 118 | 11.1G |
*bench VRAM includes default 4096 context. On GamingPC with num_ctx=2048: 3.2GB total, fully on GPU.
Grounding on screens where the target was actually visible (dialog/chrome/settings/article/search):
| Model | avg | per-screen |
|---|---|---|
| qwen2.5vl:3b | 0.91 | .90 / .93 / .85 / .94 / .93 |
| qwen2.5vl:7b | 0.91 | .90 / .94 / .84 / .94 / .93 |
| minicpm-v4.5 | 0.81 | .50 / .75 / .93 / .90 / .95 |
| llava-phi3 | 0.43 | inconsistent (0/.32/0/.86/.98) |
| gemma3:4b | 0.27 | partial only |
| llava:7b / moondream / granite3.2-vision | ~0 | unusable for tapping |
Winner: qwen2.5vl:3b
- Grounding: 0.91 — identical to its 7b big brother at half the size and 40% more speed
- 1.00 app identification, 219 tok/s, warm answers ~0.5–1.5s
- On GamingPC 8GB: 3.2GB VRAM (num_ctx 2048), fits alongside granite4.1:3b keep-warm
- Live tap test on calculator: 7→(59,310) vs true (73,357) — 20–50px accuracy, hits the buttons
Control-loop demo (GamingPC → VM1301)
Vision-only loop for "type 7+8 and press =": model tapped 7 → 8 → + → = in the correct sequence with ~20px precision. Multi-step autonomy flaked afterward (pressed back, looped on 7/8) because a stateless single-screenshot loop can't self-verify. Fix that works: keep step history in the prompt + let the agent verify with uiautomator (exact bounds) after every tap. Production architecture = Hermes agent decides → qwen2.5vl:3b sees/grounds → android-adb MCP executes → uiautomator dump verifies.
Model verdicts (the interesting ones)
- qwen2.5vl:7b — no grounding advantage over the 3b. Skip on 8GB cards.
- minicpm-v4.5:8b — best descriptions (0.64) + good grounding (0.81), but 6.1GB, 118 tok/s, and its JSON output wanders (multi-action plans, its own schema). Good fallback.
- minicpm-v4.6 (1.6GB, on GamingPC) — thinking-only variant: content is always empty. Unusable for direct control.
- granite3.2-vision:2b — returns tool calls (
getpixel) instead of answers. Wrong fit for prompt-based control, interesting for tool-use harnesses. - llava:7b / moondream / llava-phi3 — grounding garbage. Don't bother.
- llama3.2-vision:11b — fails to load on Ollama 0.32.9 (
unknown model architecture: 'mllama'). Not available on this stack.
Pitfalls hit (all fixed)
- Android screen sleep → black screencaps (4.9KB). Fix:
svc power stayon true. - VM1301 IP drifts — kernel
ip=.26is overridden by Android's EthernetService DHCP (router gave .79). Fix:adb connect Android.local:5555(mDNS) — DHCP-proof. android-adb MCP now auto-discovers: env → online device → Android.local → legacy .84. - Stale adb daemon → "No route to host" despite TCP working. Fix:
adb kill-server. - Markdown fences around model JSON — strict
json.loadsscores 0; fence-tolerant parsing recovers 0.0 → 0.62 on agent task. - qwen2.5vl outputs either
{"x": int}or{"coordinate": [x,y]}or bbox lists — harness must normalize all three. - urllib routed LAN Ollama through a system proxy → 404s. Fix:
ProxyHandler({}). - Realme phone (10.30.20.104) and UGOOS TV (.84) both offline — VM1301 is the live target.
Deployment (done)
qwen2.5vl:3b+qwen2.5vl:7binstalled on GamingPC (10.30.20.186:11434)- ADB MCP patched for device auto-discovery; verified working against VM1301
- Recommended call:
num_ctx=2048,temperature=0, ask for{"coordinate":[x,y]}JSON - Demo harness:
~/vision_bench/android_control_demo.py <model> <goal> - Raw data:
~/vision_bench/bench_results.jsonl· analysis:analyze.py,final_scores.py
Recommended follow-ups
- Pin qwen2.5vl:3b in
keep_ollama_hot.pyon .186 (vision goto, next to ornith) - Wire Hermes auxiliary vision to qwen2.5vl:3b@.186 (config.yaml auxiliary.vision)
- Add step-history + uiautomator verification to the control harness for multi-step tasks
- Replace dead UGOOS TV .84 default references; keep mDNS path as primary