Files
vision_bench/REPORT.md
2026-10-06 23:43:27 -07:00

5.5 KiB
Raw Blame History

Vision Model Benchmark — Android Control via MCP

Date: Aug 13, 2026 · Bench host: shadow-death (10.30.20.128, RTX 4080 SUPER 16GB, Ollama 0.32.9) Target: tiny vision model for GamingPC (10.30.20.186, RTX 3070 8GB) to drive Android via ADB/MCP Device: Android Cloud VM1301 (Android-x86 9.0, 1024x768, Android.local:5555)

Method

8 test situations captured from a real Android device with uiautomator XML ground truth: launcher dialog, Chrome welcome, home screen, Settings, Chrome/Wikipedia article, Calculator, Settings search, app drawer. Each model ran 4 tasks per situation (identify, describe, ground coordinates, agent JSON) = 320 scored calls. Grounding score = pixel-distance accuracy against the real element center from the XML dump. No vibes — every score is measured.

Results

Model Size identify describe ground agent JSON tok/s VRAM (bench)
moondream:1.8b 1.7G 0.12 0.21 0.00 0.00 331 1.2G
granite3.2-vision:2b 1.7G 0.62 0.04 0.00 0.38 223 3.1G
qwen2.5vl:3b 2.6G 1.00 0.55 0.57 (0.91 visible) 0.62 219 8.2G*
llava-phi3:3.8b 2.9G 0.62 0.14 0.27 0.00 196 4.0G
gemma3:4b 3.3G 1.00 0.40 0.30 0.38 161 3.0G
llava:7b 4.7G 0.62 0.10 0.03 0.62 145 8.7G
qwen2.5vl:7b 5.7G 0.88 0.52 0.57 (0.91 visible) 0.38 129 12.7G
minicpm-v4.5:8b 6.1G 1.00 0.64 0.54 (0.81 visible) 0.12 118 11.1G

*bench VRAM includes default 4096 context. On GamingPC with num_ctx=2048: 3.2GB total, fully on GPU.

Grounding on screens where the target was actually visible (dialog/chrome/settings/article/search):

Model avg per-screen
qwen2.5vl:3b 0.91 .90 / .93 / .85 / .94 / .93
qwen2.5vl:7b 0.91 .90 / .94 / .84 / .94 / .93
minicpm-v4.5 0.81 .50 / .75 / .93 / .90 / .95
llava-phi3 0.43 inconsistent (0/.32/0/.86/.98)
gemma3:4b 0.27 partial only
llava:7b / moondream / granite3.2-vision ~0 unusable for tapping

Winner: qwen2.5vl:3b

  • Grounding: 0.91 — identical to its 7b big brother at half the size and 40% more speed
  • 1.00 app identification, 219 tok/s, warm answers ~0.5–1.5s
  • On GamingPC 8GB: 3.2GB VRAM (num_ctx 2048), fits alongside granite4.1:3b keep-warm
  • Live tap test on calculator: 7→(59,310) vs true (73,357) — 20–50px accuracy, hits the buttons

Control-loop demo (GamingPC → VM1301)

Vision-only loop for "type 7+8 and press =": model tapped 7 → 8 → + → = in the correct sequence with ~20px precision. Multi-step autonomy flaked afterward (pressed back, looped on 7/8) because a stateless single-screenshot loop can't self-verify. Fix that works: keep step history in the prompt + let the agent verify with uiautomator (exact bounds) after every tap. Production architecture = Hermes agent decides → qwen2.5vl:3b sees/grounds → android-adb MCP executes → uiautomator dump verifies.

Model verdicts (the interesting ones)

  • qwen2.5vl:7b — no grounding advantage over the 3b. Skip on 8GB cards.
  • minicpm-v4.5:8b — best descriptions (0.64) + good grounding (0.81), but 6.1GB, 118 tok/s, and its JSON output wanders (multi-action plans, its own schema). Good fallback.
  • minicpm-v4.6 (1.6GB, on GamingPC) — thinking-only variant: content is always empty. Unusable for direct control.
  • granite3.2-vision:2b — returns tool calls (getpixel) instead of answers. Wrong fit for prompt-based control, interesting for tool-use harnesses.
  • llava:7b / moondream / llava-phi3 — grounding garbage. Don't bother.
  • llama3.2-vision:11b — fails to load on Ollama 0.32.9 (unknown model architecture: 'mllama'). Not available on this stack.

Pitfalls hit (all fixed)

  1. Android screen sleep → black screencaps (4.9KB). Fix: svc power stayon true.
  2. VM1301 IP drifts — kernel ip=.26 is overridden by Android's EthernetService DHCP (router gave .79). Fix: adb connect Android.local:5555 (mDNS) — DHCP-proof. android-adb MCP now auto-discovers: env → online device → Android.local → legacy .84.
  3. Stale adb daemon → "No route to host" despite TCP working. Fix: adb kill-server.
  4. Markdown fences around model JSON — strict json.loads scores 0; fence-tolerant parsing recovers 0.0 → 0.62 on agent task.
  5. qwen2.5vl outputs either {"x": int} or {"coordinate": [x,y]} or bbox lists — harness must normalize all three.
  6. urllib routed LAN Ollama through a system proxy → 404s. Fix: ProxyHandler({}).
  7. Realme phone (10.30.20.104) and UGOOS TV (.84) both offline — VM1301 is the live target.

Deployment (done)

  • qwen2.5vl:3b + qwen2.5vl:7b installed on GamingPC (10.30.20.186:11434)
  • ADB MCP patched for device auto-discovery; verified working against VM1301
  • Recommended call: num_ctx=2048, temperature=0, ask for {"coordinate":[x,y]} JSON
  • Demo harness: ~/vision_bench/android_control_demo.py <model> <goal>
  • Raw data: ~/vision_bench/bench_results.jsonl · analysis: analyze.py, final_scores.py
  • Pin qwen2.5vl:3b in keep_ollama_hot.py on .186 (vision goto, next to ornith)
  • Wire Hermes auxiliary vision to qwen2.5vl:3b@.186 (config.yaml auxiliary.vision)
  • Add step-history + uiautomator verification to the control harness for multi-step tasks
  • Replace dead UGOOS TV .84 default references; keep mDNS path as primary