GUI Navigation on Multimodal-Mind2Web Cross-Website
62.2Step Success RateHTML-T5-XL
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| HTML-T5-XLModel Size=3B, Input Modality=Text, Select From Top=true2025.03 | 62.2 | 68.4 | 71 | |
| Explorer-7BEvaluation Protocol=Supervised Fine-Tuning, Train Data=Syn. + M2W2025.02 | 56.7 | 60.5 | 90.7 | |
| AutoWebGLMModel Size=6B, Input Modality=Text, Select From Top=true2025.03 | 56.4 | — | — | |
| CogAgentModel Size=18B, Input Modality=Image, Select From Top=true2025.03 | 54 | — | — | |
| Explorer-7BEvaluation Protocol=Supervised Fine-Tuning, Train Data=M2W2025.02 | 52 | 56.3 | 89.7 | |
| AgentTrek-7BEvaluation Protocol=Supervised Fine-Tuning, Train Data=Syn. + M2W2025.02 | 51.4 | 57.6 | 88.1 | |
| Explorer-4BEvaluation Protocol=Supervised Fine-Tuning, Train Data=Syn. + M2W2025.02 | 51.4 | 55.6 | 89.5 | |
| SpiritSight-26BModel Size=26B, Input Modality=Image, Select From Top=false2025.03 | 48.1 | 57 | 85.7 | |
| LLaMA2-7BModel Size=7B, Input Modality=Text, Select From Top=true2025.03 | 47.1 | — | — | |
| LPOTraining Protocol=After Preference Optimization2025.06 | 46.4 | 64.4 | 74.4 | |
| Explorer-4BEvaluation Protocol=Supervised Fine-Tuning, Train Data=M2W2025.02 | 45 | 49.1 | 87.2 | |
| Explorer-7BEvaluation Protocol=Supervised Fine-Tuning, Train Data=Syn. traj.2025.02 | 44.5 | 48.7 | 87.7 | |
| SpiritSight-8BModel Size=8B, Input Modality=Image, Select From Top=false2025.03 | 44 | 52.2 | 84.7 | |
| RGUI-R1Training Protocol=After Preference Optimization2025.06 | 43.5 | 61.2 | 67.6 | |
| Explorer-4BEvaluation Protocol=Supervised Fine-Tuning, Train Data=Syn. traj.2025.02 | 39.3 | 44.1 | 87.7 | |
| Base ModelTraining Protocol=After Supervised Fine-Tuning2025.06 | 38.4 | 60.7 | 56.9 | |
| SpiritSight-2BModel Size=2B, Input Modality=Image, Select From Top=false2025.03 | 37.8 | 44 | 83.6 | |
| OmniParserInput Modality=Image, Select From Top=false2025.03 | 36.5 | 41 | 84.8 | |
| RInfiGUI-R1Training Protocol=After Preference Optimization2025.06 | 34.4 | 62.2 | 49.5 | |
| ScribeAgent-32BEvaluation Protocol=Supervised Fine-Tuning, Train Data=Syn. traj.2025.02 | 32.5 | 34.1 | 52.7 | |
| SeeActInput Modality=Text+Image, Select From Top=false2025.03 | 32.4 | 38 | 67.8 | |
| SeeActEvaluation Protocol=In-Context Learning2025.02 | 32.4 | 38 | 69.8 | |
| ReadAgent-PModel Size=340B, Input Modality=Text, Select From Top=false2025.03 | 31.1 | 37.4 | 75.1 | |
| MiniCPM-GUILanguage Model Size=Larger2024.06 | 29.7 | 34.7 | — | |
| GPT-4Evaluation Protocol=In-Context Learning2025.02 | 27 | 30.2 | 61 | |
| RUI-R1Training Protocol=After Preference Optimization2025.06 | 22.1 | 56.5 | 31.5 | |
| EDGE-9.6BEvaluation Protocol=Supervised Fine-Tuning, Train Data=Syn. + M2W2025.02 | 21.1 | — | — | |
| SeeClick-9.6BEvaluation Protocol=Supervised Fine-Tuning, Train Data=Syn. + M2W2025.02 | 18.8 | 21.9 | 82.9 | |
| MiniCPM-GUIModel Size=3B, Input Modality=Image, Select From Top=false2025.03 | 17.3 | 20.3 | 81.7 | |
| MiniCPM-3.1BEvaluation Protocol=Supervised Fine-Tuning, Train Data=Syn. + M2W2025.02 | 17.3 | 20.3 | 81.7 | |
| MiniCPM-GUILanguage Model Size=Standard2024.06 | 17.3 | 20.3 | — | |
| SeeClickModel Size=9.6B, Input Modality=Image, Select From Top=false2025.03 | 16.4 | 21.4 | 80.6 | |
| SeeClick2024.06 | 16.4 | 21.4 | — | |
| Qwen-GUI2024.06 | 15.6 | 19.3 | — | |
| GPT-3.5Evaluation Protocol=In-Context Learning2025.02 | 14.1 | 14.9 | 56.5 | |
| Fuyu-GUIModel Size=8B, Input Modality=Image, Select From Top=false2025.03 | 12.2 | 13.9 | 80.7 | |
| Fuyu-GUI2024.06 | 12.2 | 13.9 | — |