GUI Grounding on ScreenSpot-Pro (test)
74.1Element AccuracyVISTA
Evaluation Results
| Method | Links | |
|---|---|---|
| VISTAModel Scale=≥30B, Inference Method=Multi-view (+ MVP)2026.06 | 74.1 | |
| VISTAModel Scale=≈8B, Inference Method=Multi-view (+ MVP)2026.06 | 72 | |
| VISTAModel Scale=≈4B, Inference Method=Multi-view (+ MVP)2026.06 | 71.6 | |
| VISTAModel Scale=≥30B, Inference Method=Single-view2026.06 | 67 | |
| Holo2Model Scale=≥30B, Inference Method=Single-view2026.06 | 66.1 | |
| MAI-UIModel Scale=≈8B, Inference Method=Single-view2026.06 | 65.8 | |
| VISTAModel Scale=≈8B, Inference Method=Single-view2026.06 | 65.8 | |
| GTA1Model Scale=≥30B, Inference Method=Single-view2026.06 | 63.6 | |
| VISTAModel Scale=≈4B, Inference Method=Single-view2026.06 | 63.4 | |
| Step-GUIModel Scale=≈8B, Inference Method=Single-view2026.06 | 62.6 | |
| UI-VenusModel Scale=72B, Inference Method=Single-view2026.06 | 61.9 | |
| OpenCUAModel Scale=72B, Inference Method=Single-view2026.06 | 60.8 | |
| Step-GUIModel Scale=≈4B, Inference Method=Single-view2026.06 | 60 | |
| Holo2Model Scale=≈8B, Inference Method=Single-view2026.06 | 58.9 | |
| Holo1.5-7B + SafeGroundRisk Level=0.342026.02 | 58.66 | |
| GUI-OwlModel Scale=≥30B, Inference Method=Single-view2026.06 | 58 | |
| Holo1.5-7B + SafeGroundRisk Level=0.382026.02 | 57.87 | |
| Holo1.5-7B + SafeGroundRisk Level=0.422026.02 | 55.73 | |
| Qwen3-VLModel Scale=≈4B, Inference Method=Single-view2026.06 | 55.5 | |
| OpenCUAModel Scale=≥30B, Inference Method=Single-view2026.06 | 55.3 | |
| GUI-Actor-2.5VL-7B + SafeGroundRisk Level=0.342026.02 | 55.18 | |
| GUI-Actor-2VL-7B + SafeGroundRisk Level=0.342026.02 | 55.18 | |
| GUI-OwlModel Scale=≈8B, Inference Method=Single-view2026.06 | 54.9 | |
| GUI-Actor-2.5VL-7B + SafeGroundRisk Level=0.382026.02 | 54.86 | |
| UI-TARS-1.5-7B + SafeGroundRisk Level=0.382026.02 | 54.7 | |
| GUI-Actor-2VL-7B + SafeGroundRisk Level=0.422026.02 | 53.99 | |
| Qwen3-VLModel Scale=≥30B, Inference Method=Single-view2026.06 | 53.7 | |
| UI-TARS-1.5-7B + SafeGroundRisk Level=0.342026.02 | 53.68 | |
| GUI-Actor-2.5VL-7B + SafeGroundRisk Level=0.422026.02 | 53.6 | |
| Holo1.5-3B + SafeGroundRisk Level=0.342026.02 | 53.44 | |
| Gemini-onlyRisk Level=0.342026.02 | 53.28 | |
| GUI-Actor-2VL-7B + SafeGroundRisk Level=0.382026.02 | 53.28 | |
| Holo1.5-7B + SafeGroundRisk Level=0.462026.02 | 53.2 | |
| GTA1-7B + SafeGroundRisk Level=0.462026.02 | 53.12 | |
| UI-TARS-1.5-7B + SafeGroundRisk Level=0.422026.02 | 53.04 | |
| GUI-Actor-2VL-7B + SafeGroundRisk Level=0.462026.02 | 52.96 | |
| Holo1.5-3B + SafeGroundRisk Level=0.382026.02 | 52.73 | |
| Qwen3-VLModel Scale=≈8B, Inference Method=Single-view2026.06 | 52.7 | |
| Holo1.5-7BCascading=None2026.02 | 52.41 | |
| Holo1.5-7B + SafeGroundRisk Level=0.502026.02 | 52.41 | |
| Holo1.5-3B + SafeGroundRisk Level=0.422026.02 | 52.02 | |
| GUI-Actor-2.5VL-7B + SafeGroundRisk Level=0.462026.02 | 51.38 | |
| Trifuse+GUI-AIMAModel Size=7B, Training Protocol=Training-based (TB)2026.02 | 51.3 | |
| UI-VenusModel Scale=≈8B, Inference Method=Single-view2026.06 | 50.8 | |
| GUI-Actor-2VL-7B + SafeGroundRisk Level=0.502026.02 | 50.67 | |
| UI-TARS-1.5-7B + SafeGroundRisk Level=0.462026.02 | 50.43 | |
| GTA1Model Scale=≈8B, Inference Method=Single-view2026.06 | 50.1 | |
| OpenCUAModel Scale=≈8B, Inference Method=Single-view2026.06 | 50 | |
| GTA1-7B + SafeGroundRisk Level=0.502026.02 | 49.96 | |
| GUI-AIMAModel Size=3B, Training Protocol=Training-based (TB)2026.02 | 49.8 | |
| Holo1.5-3B + SafeGroundRisk Level=0.462026.02 | 49.25 | |
| GUI-Actor-2.5VL-7B + SafeGroundRisk Level=0.502026.02 | 49.17 | |
| UI-TARS-1.5-7B + SafeGroundRisk Level=0.502026.02 | 47.91 | |
| Holo1.5-3B + SafeGroundRisk Level=0.502026.02 | 47.35 | |
| GTA1-7BCascading=None2026.02 | 46.88 | |
| Trifuse+GUI-ActorModel Size=7B, Training Protocol=Training-based (TB)2026.02 | 45.8 | |
| GUI-Actor-2.5VL-7BCascading=None2026.02 | 45.69 | |
| Holo1.5-3BCascading=None2026.02 | 45.45 | |
| GUI-ActorModel Size=7B, Training Protocol=Training-based (TB)2026.02 | 44.6 | |
| Trifuse+GUI-ActorModel Size=3B, Training Protocol=Training-based (TB)2026.02 | 42.7 | |
| GUI-ActorModel Size=3B, Training Protocol=Training-based (TB)2026.02 | 42.2 | |
| UI-TARS-1.5-7BCascading=None2026.02 | 41.58 | |
| GUI-Actor-2VL-7BCascading=None2026.02 | 40.79 | |
| UI-TARS-1.5Model Scale=≈8B, Inference Method=Single-view2026.06 | 35.7 | |
| TrifuseModel Size=7B, Training Protocol=Training-free (TF)2026.02 | 29.7 | |
| TrifuseModel Size=3B, Training Protocol=Training-free (TF)2026.02 | 18.9 | |
| TAGModel Size=8.5B, Training Protocol=Training-free (TF)2026.02 | 3 |