Mobile Agent Evaluation on AndroidControl Low (test)
93.7Task Success RateQwen2.5-VL-72B
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Qwen2.5-VL-72B2025.08 | 93.7 | — | — | |
| Qwen2.5-VL-32B2025.08 | 93.3 | — | — | |
| CRAFT-GUI-32BTraining Stage=stage32025.08 | 92.7 | 93.6 | 98.9 | |
| CRAFT-GUI-32BTraining Stage=stage22025.08 | 92.2 | 93.1 | 98.8 | |
| UI-TARS-72B2025.08 | 91.3 | 89.9 | 98.1 | |
| CRAFT-GUI-32BTraining Stage=stage12025.08 | 91.3 | 92.4 | 98 | |
| OS-Atlas-7B2025.08 | 85.2 | 88 | 93.6 | |
| Aguvis-72B2025.08 | 84.4 | — | — | |
| Qwen2-VL-7B2025.08 | 82.6 | 86.5 | 91.9 | |
| InternVL-2-4B2025.08 | 80.1 | 84.1 | 90.9 | |
| SeeClick2025.08 | 75 | 73.4 | 93 | |
| Aria-UI2025.08 | 67.3 | 87.7 | — | |
| ClaudeVersion=Claude-computer-use2025.08 | 19.4 | 0 | 74.3 | |
| GPT-4o2025.08 | 19.4 | 0 | 74.3 | |
| Aria-UITHEvaluation Protocol=W. Training Set2024.12 | 0.6733 | 0.8769 | — | |
| Aria-UIIHEvaluation Protocol=W. Training Set2024.12 | 0.6726 | 0.872 | — | |
| Aria-UIEvaluation Protocol=W. Training Set2024.12 | 0.663 | 0.8571 | — | |
| Aria-UIEvaluation Protocol=Zero-shot2024.12 | 0.5439 | 0.797 | — | |
| UGroundEvaluation Protocol=W. Training Set2024.12 | 0.4685 | 0.7428 | — | |
| Qwen2-VLEvaluation Protocol=Zero-shot2024.12 | 0.3253 | 0.6424 | — | |
| SeeClickEvaluation Protocol=Zero-shot2024.12 | 0.1772 | 0.4555 | — | |
| GPT-4oEvaluation Protocol=Zero-shot2024.12 | 0.0512 | 0.1636 | — |