18 lines
731 B
Markdown
18 lines
731 B
Markdown
# Vision Bench — Android-Control Vision Model Benchmark
|
||
|
||
Measured (not vibes) comparison of tiny Ollama vision models for driving
|
||
Android via ADB/MCP. Ground truth = uiautomator XML dumps; 8 real device
|
||
situations × 4 tasks × 4 models = 320 scored calls.
|
||
|
||
## Scores (identify / ground / agent JSON)
|
||
- **qwen2.5vl:3b — 1.00 / 0.57 (0.91 visible) / 0.62** ← winner
|
||
- granite3.2-vision:2b — 0.62 / 0.00 / 0.38
|
||
- llava-phi3:3.8b — 0.62 / 0.27 / 0.00
|
||
- moondream:1.8b — 0.12 / 0.00 / 0.00
|
||
|
||
## Files
|
||
`bench.py` harness · `capture_situations.sh` ground-truth capture ·
|
||
`final_scores.py` / `analyze.py` scoring · `bench_results.jsonl` raw data ·
|
||
`REPORT.md` full write-up · `android_control_demo.py` working agent demo
|
||
|