GUI world modeling on MobileWorldBench
2.78AccuracyMobileWorldModel-8B
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| MobileWorldModel-8Bvisual hint=true, VLM-as-a-judge=Gemini-3.1-Pro2026.05 | 2.78 | 2.89 | 2.7 | 8.37 | |
| MobileWorldModel-8Bvisual hint=false, VLM-as-a-judge=Gemini-3.1-Pro2026.05 | 2.66 | 2.76 | 2.64 | 8.07 | |
| GPT-5.2VLM-as-a-judge=Gemini-3.1-Pro2026.05 | 2.62 | 2.48 | 3.02 | 8.11 | |
| Gemini-3.1-ProVLM-as-a-judge=Gemini-3.1-Pro2026.05 | 2.57 | 2.42 | 2.97 | 7.96 | |
| MobileWorldModel-4Bvisual hint=true, VLM-as-a-judge=Gemini-3.1-Pro2026.05 | 2.49 | 2.6 | 2.5 | 7.6 | |
| MobileWorldModel-4Bvisual hint=false, VLM-as-a-judge=Gemini-3.1-Pro2026.05 | 2.47 | 2.59 | 2.43 | 7.48 | |
| Claude-4.6-SonnetVLM-as-a-judge=Gemini-3.1-Pro2026.05 | 2.19 | 2.08 | 2.4 | 6.67 | |
| Qwen3-VL-4BVLM-as-a-judge=Gemini-3.1-Pro2026.05 | 1.98 | 1.83 | 2.1 | 5.83 | |
| Qwen3-VL-8BVLM-as-a-judge=Gemini-3.1-Pro2026.05 | 1.98 | 2.02 | 1.97 | 5.97 | |
| Qwen3-VL-32BVLM-as-a-judge=Gemini-3.1-Pro2026.05 | 1.93 | 1.92 | 2.11 | 5.95 |