d042f6685feacf279d4bd35600e5475249c5e922
Vision Bench — Android-Control Vision Model Benchmark
Measured (not vibes) comparison of tiny Ollama vision models for driving Android via ADB/MCP. Ground truth = uiautomator XML dumps; 8 real device situations × 4 tasks × 4 models = 320 scored calls.
Scores (identify / ground / agent JSON)
- qwen2.5vl:3b — 1.00 / 0.57 (0.91 visible) / 0.62 ← winner
- granite3.2-vision:2b — 0.62 / 0.00 / 0.38
- llava-phi3:3.8b — 0.62 / 0.27 / 0.00
- moondream:1.8b — 0.12 / 0.00 / 0.00
Files
bench.py harness · capture_situations.sh ground-truth capture ·
final_scores.py / analyze.py scoring · bench_results.jsonl raw data ·
REPORT.md full write-up · android_control_demo.py working agent demo
Description
Languages
Python
90.7%
Shell
9.3%