GUI Navigation on Multimodal-Mind2Web Cross-Task
71.5Step Success RateHTML-T5-XL
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| HTML-T5-XLModel Size=3B, Input Modality=Text, Select From Top=true2025.03 | 71.5 | 76.4 | 78.8 | |
| AutoWebGLMModel Size=6B, Input Modality=Text, Select From Top=true2025.03 | 66.4 | — | — | |
| CogAgentModel Size=18B, Input Modality=Image, Select From Top=true2025.03 | 62.3 | — | — | |
| AgentTrek-7BEvaluation Protocol=Supervised Fine-Tuning, Train Data=Syn. + M2W2025.02 | 55.7 | 60.8 | 88.9 | |
| SpiritSight-26BModel Size=26B, Input Modality=Image, Select From Top=false2025.03 | 54.7 | 60.5 | 89.7 | |
| Explorer-7BEvaluation Protocol=Supervised Fine-Tuning, Train Data=Syn. + M2W2025.02 | 53.2 | 56.5 | 90.3 | |
| LLaMA2-7BModel Size=7B, Input Modality=Text, Select From Top=true2025.03 | 52.7 | — | — | |
| SpiritSight-8BModel Size=8B, Input Modality=Image, Select From Top=false2025.03 | 52.7 | 59.2 | 88.9 | |
| Explorer-4BEvaluation Protocol=Supervised Fine-Tuning, Train Data=Syn. + M2W2025.02 | 50.7 | 53.4 | 88.1 | |
| LPOTraining Protocol=After Preference Optimization2025.06 | 49.5 | 64.3 | 76.7 | |
| Explorer-7BEvaluation Protocol=Supervised Fine-Tuning, Train Data=M2W2025.02 | 48.3 | 51.8 | 88 | |
| RGUI-R1Training Protocol=After Preference Optimization2025.06 | 46.6 | 62.5 | 71.6 | |
| SpiritSight-2BModel Size=2B, Input Modality=Image, Select From Top=false2025.03 | 44.9 | 51.7 | 87.2 | |
| Explorer-4BEvaluation Protocol=Supervised Fine-Tuning, Train Data=M2W2025.02 | 44.8 | 48.1 | 88 | |
| SeeActInput Modality=Text+Image, Select From Top=false2025.03 | 40.2 | 46.4 | 73.4 | |
| SeeActEvaluation Protocol=In-Context Learning2025.02 | 40.2 | 46.4 | 73.4 | |
| Explorer-7BEvaluation Protocol=Supervised Fine-Tuning, Train Data=Syn. traj.2025.02 | 39.6 | 43.6 | 86.6 | |
| OmniParserInput Modality=Image, Select From Top=false2025.03 | 39.4 | 42.4 | 87.6 | |
| Base ModelTraining Protocol=After Supervised Fine-Tuning2025.06 | 38.2 | 60.3 | 57.4 | |
| RInfiGUI-R1Training Protocol=After Preference Optimization2025.06 | 35.8 | 62.6 | 51.3 | |
| ScribeAgent-32BEvaluation Protocol=Supervised Fine-Tuning, Train Data=Syn. traj.2025.02 | 35.6 | 38 | 52.9 | |
| Explorer-4BEvaluation Protocol=Supervised Fine-Tuning, Train Data=Syn. traj.2025.02 | 33.2 | 36.5 | 82.9 | |
| GPT-4Evaluation Protocol=In-Context Learning2025.02 | 32.3 | 40.8 | 63.1 | |
| EDGE-9.6BEvaluation Protocol=Supervised Fine-Tuning, Train Data=Syn. + M2W2025.02 | 30 | — | — | |
| ReadAgent-PModel Size=340B, Input Modality=Text, Select From Top=false2025.03 | 29.2 | 33.7 | 72.5 | |
| SeeClickModel Size=9.6B, Input Modality=Image, Select From Top=false2025.03 | 25.5 | 28.3 | 87 | |
| RUI-R1Training Protocol=After Preference Optimization2025.06 | 24.9 | 59.5 | 34.5 | |
| SeeClick-9.6BEvaluation Protocol=Supervised Fine-Tuning, Train Data=Syn. + M2W2025.02 | 23.7 | 26.3 | 86.2 | |
| MiniCPM-GUIModel Size=3B, Input Modality=Image, Select From Top=false2025.03 | 20.8 | 23.8 | 86.8 | |
| MiniCPM-3.1BEvaluation Protocol=Supervised Fine-Tuning, Train Data=Syn. + M2W2025.02 | 20.8 | 23.8 | 86.8 | |
| GPT-3.5Evaluation Protocol=In-Context Learning2025.02 | 16.8 | 19.4 | 59.2 | |
| Fuyu-GUIModel Size=8B, Input Modality=Image, Select From Top=false2025.03 | 15.6 | 19.1 | 86.1 |