Files
vision_bench/REPORT.md
2026-10-06 23:43:27 -07:00

90 lines
5.5 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Vision Model Benchmark — Android Control via MCP
**Date:** Aug 13, 2026 · **Bench host:** shadow-death (10.30.20.128, RTX 4080 SUPER 16GB, Ollama 0.32.9)
**Target:** tiny vision model for GamingPC (10.30.20.186, RTX 3070 8GB) to drive Android via ADB/MCP
**Device:** Android Cloud VM1301 (Android-x86 9.0, 1024x768, `Android.local:5555`)
## Method
8 test situations captured from a real Android device with **uiautomator XML ground truth**:
launcher dialog, Chrome welcome, home screen, Settings, Chrome/Wikipedia article, Calculator,
Settings search, app drawer. Each model ran 4 tasks per situation (identify, describe, **ground
coordinates**, agent JSON) = 320 scored calls. Grounding score = pixel-distance accuracy against
the real element center from the XML dump. No vibes — every score is measured.
## Results
| Model | Size | identify | describe | **ground** | agent JSON | tok/s | VRAM (bench) |
|---|---|---|---|---|---|---|---|
| moondream:1.8b | 1.7G | 0.12 | 0.21 | 0.00 | 0.00 | 331 | 1.2G |
| granite3.2-vision:2b | 1.7G | 0.62 | 0.04 | 0.00 | 0.38 | 223 | 3.1G |
| **qwen2.5vl:3b** | **2.6G** | **1.00** | **0.55** | **0.57 (0.91 visible)** | **0.62** | **219** | 8.2G* |
| llava-phi3:3.8b | 2.9G | 0.62 | 0.14 | 0.27 | 0.00 | 196 | 4.0G |
| gemma3:4b | 3.3G | 1.00 | 0.40 | 0.30 | 0.38 | 161 | 3.0G |
| llava:7b | 4.7G | 0.62 | 0.10 | 0.03 | 0.62 | 145 | 8.7G |
| qwen2.5vl:7b | 5.7G | 0.88 | 0.52 | 0.57 (0.91 visible) | 0.38 | 129 | 12.7G |
| minicpm-v4.5:8b | 6.1G | 1.00 | 0.64 | 0.54 (0.81 visible) | 0.12 | 118 | 11.1G |
*bench VRAM includes default 4096 context. On GamingPC with num_ctx=2048: **3.2GB total, fully on GPU.**
**Grounding on screens where the target was actually visible** (dialog/chrome/settings/article/search):
| Model | avg | per-screen |
|---|---|---|
| qwen2.5vl:3b | **0.91** | .90 / .93 / .85 / .94 / .93 |
| qwen2.5vl:7b | **0.91** | .90 / .94 / .84 / .94 / .93 |
| minicpm-v4.5 | 0.81 | .50 / .75 / .93 / .90 / .95 |
| llava-phi3 | 0.43 | inconsistent (0/.32/0/.86/.98) |
| gemma3:4b | 0.27 | partial only |
| llava:7b / moondream / granite3.2-vision | ~0 | unusable for tapping |
## Winner: qwen2.5vl:3b
- **Grounding: 0.91** — identical to its 7b big brother at half the size and 40% more speed
- 1.00 app identification, 219 tok/s, warm answers ~0.5–1.5s
- On GamingPC 8GB: **3.2GB VRAM** (num_ctx 2048), fits alongside granite4.1:3b keep-warm
- Live tap test on calculator: 7→(59,310) vs true (73,357) — 20–50px accuracy, hits the buttons
## Control-loop demo (GamingPC → VM1301)
Vision-only loop for "type 7+8 and press =": model tapped **7 → 8 → + → =** in the correct
sequence with ~20px precision. Multi-step autonomy flaked afterward (pressed back, looped on
7/8) because a stateless single-screenshot loop can't self-verify. **Fix that works:** keep step
history in the prompt + let the agent verify with uiautomator (exact bounds) after every tap.
Production architecture = Hermes agent decides → qwen2.5vl:3b sees/grounds → android-adb MCP
executes → uiautomator dump verifies.
## Model verdicts (the interesting ones)
- **qwen2.5vl:7b** — no grounding advantage over the 3b. Skip on 8GB cards.
- **minicpm-v4.5:8b** — best descriptions (0.64) + good grounding (0.81), but 6.1GB, 118 tok/s,
and its JSON output wanders (multi-action plans, its own schema). Good fallback.
- **minicpm-v4.6** (1.6GB, on GamingPC) — **thinking-only variant: content is always empty.**
Unusable for direct control.
- **granite3.2-vision:2b** — returns *tool calls* (`getpixel`) instead of answers. Wrong fit for
prompt-based control, interesting for tool-use harnesses.
- **llava:7b / moondream / llava-phi3** — grounding garbage. Don't bother.
- **llama3.2-vision:11b** — **fails to load on Ollama 0.32.9** (`unknown model architecture:
'mllama'`). Not available on this stack.
## Pitfalls hit (all fixed)
1. **Android screen sleep → black screencaps** (4.9KB). Fix: `svc power stayon true`.
2. **VM1301 IP drifts** — kernel `ip=.26` is overridden by Android's EthernetService DHCP
(router gave .79). Fix: `adb connect Android.local:5555` (mDNS) — DHCP-proof. android-adb
MCP now auto-discovers: env → online device → Android.local → legacy .84.
3. **Stale adb daemon** → "No route to host" despite TCP working. Fix: `adb kill-server`.
4. **Markdown fences around model JSON** — strict `json.loads` scores 0; fence-tolerant parsing
recovers 0.0 → 0.62 on agent task.
5. **qwen2.5vl outputs** either `{"x": int}` or `{"coordinate": [x,y]}` or bbox lists — harness
must normalize all three.
6. **urllib routed LAN Ollama through a system proxy** → 404s. Fix: `ProxyHandler({})`.
7. Realme phone (10.30.20.104) and UGOOS TV (.84) both offline — VM1301 is the live target.
## Deployment (done)
- `qwen2.5vl:3b` + `qwen2.5vl:7b` installed on GamingPC (10.30.20.186:11434)
- ADB MCP patched for device auto-discovery; verified working against VM1301
- Recommended call: `num_ctx=2048`, `temperature=0`, ask for `{"coordinate":[x,y]}` JSON
- Demo harness: `~/vision_bench/android_control_demo.py <model> <goal>`
- Raw data: `~/vision_bench/bench_results.jsonl` · analysis: `analyze.py`, `final_scores.py`
## Recommended follow-ups
- Pin qwen2.5vl:3b in `keep_ollama_hot.py` on .186 (vision goto, next to ornith)
- Wire Hermes auxiliary vision to qwen2.5vl:3b@.186 (config.yaml auxiliary.vision)
- Add step-history + uiautomator verification to the control harness for multi-step tasks
- Replace dead UGOOS TV .84 default references; keep mDNS path as primary