2026-10-06 23:43:27 -07:00
2026-10-06 23:43:27 -07:00
2026-10-06 23:43:27 -07:00
2026-10-06 23:43:27 -07:00
2026-10-06 23:43:27 -07:00
2026-10-06 23:43:27 -07:00
2026-10-06 23:43:27 -07:00
2026-10-06 23:43:27 -07:00
2026-10-06 23:43:27 -07:00
2026-10-06 23:43:27 -07:00
2026-10-06 23:43:27 -07:00
2026-10-06 23:43:27 -07:00
2026-10-06 23:43:27 -07:00
2026-10-06 23:43:27 -07:00
2026-10-06 23:43:27 -07:00
2026-10-06 23:43:27 -07:00
2026-10-06 23:43:27 -07:00

Vision Bench — Android-Control Vision Model Benchmark

Measured (not vibes) comparison of tiny Ollama vision models for driving Android via ADB/MCP. Ground truth = uiautomator XML dumps; 8 real device situations × 4 tasks × 4 models = 320 scored calls.

Scores (identify / ground / agent JSON)

  • qwen2.5vl:3b — 1.00 / 0.57 (0.91 visible) / 0.62 ← winner
  • granite3.2-vision:2b — 0.62 / 0.00 / 0.38
  • llava-phi3:3.8b — 0.62 / 0.27 / 0.00
  • moondream:1.8b — 0.12 / 0.00 / 0.00

Files

bench.py harness · capture_situations.sh ground-truth capture · final_scores.py / analyze.py scoring · bench_results.jsonl raw data · REPORT.md full write-up · android_control_demo.py working agent demo

Description
Vision Bench — Android-Control Vision Model Benchmark
Readme 105 KiB
Languages
Python 90.7%
Shell 9.3%