Element Grounding on Multimodal-Mind2Web (out-of-distribution)
57.6Cross-Task GeneralizationAria-UITH
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Aria-UITHInput=Image, Planner=GPT-4o, Evaluation Protocol=Zero-shot, History=Textual action history2024.12 | 57.6 | 58 | 61.2 | 58.9 | |
| Aria-UIIHInput=Image, Planner=GPT-4o, Evaluation Protocol=Zero-shot, History=Text-image interleaved history2024.12 | 57.6 | 57.7 | 61.4 | 58.9 | |
| Aria-UIInput=Image, Planner=GPT-4o, Evaluation Protocol=Zero-shot2024.12 | 56.1 | 57 | 59.5 | 57.5 | |
| UGroundInput=Image, Planner=GPT-4o, Evaluation Protocol=Zero-shot2024.12 | 47.7 | 46 | 46.6 | 46.8 | |
| ChoiceInput=Image + HTML Tree, Planner=GPT-4, Evaluation Protocol=Zero-shot2024.12 | 46.4 | 38 | 42.4 | 42.3 | |
| UGroundInput=Image, Planner=GPT-4, Evaluation Protocol=Zero-shot2024.12 | 45.1 | 44.7 | 44.6 | 44.8 | |
| OmniParserInput=Image, Planner=GPT-4, Evaluation Protocol=Zero-shot2024.12 | 42.4 | 41 | 45.4 | 42.9 | |
| SeeClickInput=Image, Planner=GPT-4o, Evaluation Protocol=Zero-shot2024.12 | 32.1 | 33.1 | 33.5 | 32.9 | |
| SOMInput=Image + HTML Tree, Planner=GPT-4, Evaluation Protocol=Zero-shot2024.12 | 29.6 | 20.1 | 27 | 25.6 | |
| SeeClickInput=Image, Planner=GPT-4, Evaluation Protocol=Zero-shot2024.12 | 29.6 | 28.5 | 30.7 | 29.6 |