GUI Navigation on Multimodal-Mind2Web Cross-Domain
67.1Step Success RateHTML-T5-XL
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| HTML-T5-XLModel Size=3B, Input Modality=Text, Select From Top=true2025.03 | 67.1 | 73 | 75.6 | |
| CogAgentModel Size=18B, Input Modality=Image, Select From Top=true2025.03 | 59.4 | — | — | |
| AutoWebGLMModel Size=6B, Input Modality=Text, Select From Top=true2025.03 | 55.8 | — | — | |
| Explorer-7BEvaluation Protocol=Supervised Fine-Tuning, Train Data=Syn. + M2W2025.02 | 53 | 55.7 | 90.4 | |
| AgentTrek-7BEvaluation Protocol=Supervised Fine-Tuning, Train Data=Syn. + M2W2025.02 | 52.6 | 56 | 87.5 | |
| LLaMA2-7BModel Size=7B, Input Modality=Text, Select From Top=true2025.03 | 50.3 | — | — | |
| LPOTraining Protocol=After Preference Optimization2025.06 | 49.6 | 65.2 | 74.8 | |
| SpiritSight-26BModel Size=26B, Input Modality=Image, Select From Top=false2025.03 | 49.2 | 54.1 | 87.2 | |
| Explorer-7BEvaluation Protocol=Supervised Fine-Tuning, Train Data=M2W2025.02 | 48.1 | 50.9 | 88.9 | |
| RGUI-R1Training Protocol=After Preference Optimization2025.06 | 47.9 | 65 | 71.1 | |
| Explorer-4BEvaluation Protocol=Supervised Fine-Tuning, Train Data=Syn. + M2W2025.02 | 47.2 | 49.8 | 88.8 | |
| Explorer-7BEvaluation Protocol=Supervised Fine-Tuning, Train Data=Syn. traj.2025.02 | 44.7 | 47.6 | 87.2 | |
| Explorer-4BEvaluation Protocol=Supervised Fine-Tuning, Train Data=M2W2025.02 | 44.6 | 46.9 | 87.7 | |
| SpiritSight-8BModel Size=8B, Input Modality=Image, Select From Top=false2025.03 | 44.4 | 50.1 | 86 | |
| OmniParserInput Modality=Image, Select From Top=false2025.03 | 42 | 45.5 | 85.7 | |
| Base ModelTraining Protocol=After Supervised Fine-Tuning2025.06 | 40.7 | 63.8 | 58.5 | |
| RInfiGUI-R1Training Protocol=After Preference Optimization2025.06 | 40 | 65.1 | 53.1 | |
| Explorer-4BEvaluation Protocol=Supervised Fine-Tuning, Train Data=Syn. traj.2025.02 | 39.8 | 42.5 | 86.3 | |
| ScribeAgent-32BEvaluation Protocol=Supervised Fine-Tuning, Train Data=Syn. traj.2025.02 | 37.3 | 39.4 | 54.7 | |
| SpiritSight-2BModel Size=2B, Input Modality=Image, Select From Top=false2025.03 | 36.9 | 42.4 | 83.5 | |
| SeeActInput Modality=Text+Image, Select From Top=false2025.03 | 36.8 | 42.4 | 69.3 | |
| SeeActEvaluation Protocol=In-Context Learning2025.02 | 36.8 | 42.4 | 69.3 | |
| ReadAgent-PModel Size=340B, Input Modality=Text, Select From Top=false2025.03 | 33.4 | 37.2 | 76.3 | |
| GPT-4Evaluation Protocol=In-Context Learning2025.02 | 29.7 | 35.4 | 61.9 | |
| RUI-R1Training Protocol=After Preference Optimization2025.06 | 27.1 | 61.6 | 37.2 | |
| GPT-3.5Evaluation Protocol=In-Context Learning2025.02 | 24.1 | 25.2 | 57.9 | |
| EDGE-9.6BEvaluation Protocol=Supervised Fine-Tuning, Train Data=Syn. + M2W2025.02 | 22.4 | — | — | |
| SeeClickModel Size=9.6B, Input Modality=Image, Select From Top=false2025.03 | 20.8 | 23.2 | 84.8 | |
| SeeClick-9.6BEvaluation Protocol=Supervised Fine-Tuning, Train Data=Syn. + M2W2025.02 | 20.2 | 22.1 | 84.1 | |
| MiniCPM-GUIModel Size=3B, Input Modality=Image, Select From Top=false2025.03 | 14.6 | 17.9 | 74.5 | |
| MiniCPM-3.1BEvaluation Protocol=Supervised Fine-Tuning, Train Data=Syn. + M2W2025.02 | 14.6 | 17.9 | 74.5 | |
| Fuyu-GUIModel Size=8B, Input Modality=Image, Select From Top=false2025.03 | 11.7 | 14.2 | 83.1 |