Snapshot: full project state
This commit is contained in:
17
README.md
Normal file
17
README.md
Normal file
@@ -0,0 +1,17 @@
|
||||
# Vision Bench — Android-Control Vision Model Benchmark
|
||||
|
||||
Measured (not vibes) comparison of tiny Ollama vision models for driving
|
||||
Android via ADB/MCP. Ground truth = uiautomator XML dumps; 8 real device
|
||||
situations × 4 tasks × 4 models = 320 scored calls.
|
||||
|
||||
## Scores (identify / ground / agent JSON)
|
||||
- **qwen2.5vl:3b — 1.00 / 0.57 (0.91 visible) / 0.62** ← winner
|
||||
- granite3.2-vision:2b — 0.62 / 0.00 / 0.38
|
||||
- llava-phi3:3.8b — 0.62 / 0.27 / 0.00
|
||||
- moondream:1.8b — 0.12 / 0.00 / 0.00
|
||||
|
||||
## Files
|
||||
`bench.py` harness · `capture_situations.sh` ground-truth capture ·
|
||||
`final_scores.py` / `analyze.py` scoring · `bench_results.jsonl` raw data ·
|
||||
`REPORT.md` full write-up · `android_control_demo.py` working agent demo
|
||||
|
||||
Reference in New Issue
Block a user