GUI Grounding on ScreenSpot
92Avg AccGUI-G2
Evaluation Results
| Method | Links | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GUI-G2Size=7B2025.10 | 92 | 96.7 | 90.8 | 95.9 | 88.6 | 90.9 | 86.9 | — | — | — | — | — | |
| HyperClickSize=7B2025.10 | 91.5 | 95.6 | 91.7 | 93.8 | 82.9 | 92.2 | 88.4 | — | — | — | — | — | |
| GAIR2025.12 | 91 | 96 | 88.2 | 95.9 | 84.3 | 91.8 | 86.4 | — | — | — | — | — | |
| UI-TARS-7BParameters=7B2025.12 | 89.5 | 94.5 | 85.2 | 95.9 | 85.7 | 90 | 83.5 | — | — | — | — | — | |
| UI-TARSSize=7B2025.10 | 89.5 | 94.5 | 85.2 | 95.9 | 85.7 | 90 | 83.5 | — | — | — | — | — | |
| UGround-V1-72BParameters=72B2025.12 | 89.4 | 94.1 | 83.4 | 94.9 | 85.7 | 90.4 | 87.9 | — | — | — | — | — | |
| Aguvis-72B2025.02 | 89.2 | — | — | — | — | — | — | — | — | — | — | — | |
| UI-R1-ESize=3B2025.10 | 89.2 | 97.1 | 83 | 95.4 | 77.9 | 91.7 | 85 | — | — | — | — | — | |
| HyperClickSize=3B2025.10 | 88.5 | 96.7 | 83.9 | 92.8 | 80.7 | 88.7 | 83.5 | — | — | — | — | — | |
| AGUVIS-72BParameters=72B2025.12 | 88.4 | 94.5 | 85.2 | 95.4 | 77.9 | 91.3 | 85.9 | — | — | — | — | — | |
| GUI-ActorSize=7B2025.10 | 88.3 | 94.9 | 82.1 | 91.8 | 80 | 91.3 | 85.4 | — | — | — | — | — | |
| SE-GUISize=7B2025.10 | 88.2 | — | — | — | — | — | — | — | — | — | — | — | |
| InfiGUI-R1-3BParameters=3B2025.12 | 87.5 | 97.1 | 81.2 | 94.3 | 77.1 | 91.7 | 77.6 | — | — | — | — | — | |
| Qwen2.5-VL-72B2025.02 | 87.1 | — | — | — | — | — | — | — | — | — | — | — | |
| Full ModelBackbone=Qwen3-VL-8B-Instruct, Pruning Ratio=0%2026.02 | 87.03 | — | — | — | — | — | — | — | — | — | — | — | |
| POPBackbone=Qwen3-VL-8B-Instruct, Pruning Ratio=33.3%2026.02 | 86.4 | — | — | — | — | — | — | — | — | — | — | — | |
| UGround-V1-7BParameters=7B2025.12 | 86.3 | 93 | 79.9 | 93.8 | 76.4 | 90.9 | 84 | — | — | — | — | — | |
| TongUISize=7B2025.10 | 86 | 91.9 | 79.5 | 93.8 | 80 | 89.1 | 81.6 | — | — | — | — | — | |
| GUI-C²-3BTraining Samples=4.6K2026.05 | 85.8 | — | — | — | — | — | — | 85.1 | 85.5 | 87.1 | — | — | |
| WandaBackbone=Qwen3-VL-8B-Instruct, Pruning Ratio=30%2026.02 | 85.22 | — | — | — | — | — | — | — | — | — | — | — | |
| OS-Atlas-Base-7BPlanner=GPT-4o2024.10 | 85.14 | 93.77 | 79.91 | 90.21 | 66.43 | 92.61 | 79.13 | — | — | — | — | — | |
| OS-Atlas-7BParameters=7B2025.12 | 85.1 | 93.8 | 79.9 | 90.2 | 66.4 | 92.6 | 79.1 | — | — | — | — | — | |
| Qwen2.5-VLSize=7B2025.10 | 84.7 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL-7BTraining Samples=-2026.05 | 84.7 | — | — | — | — | — | — | — | — | — | — | — | |
| Aguvis-7BTraining Samples=1M2026.05 | 84.4 | — | — | — | — | — | — | 84.7 | 86.9 | 82.4 | — | — | |
| Gemini 2.02025.02 | 84 | — | — | — | — | — | — | — | — | — | — | — | |
| Two-Stage Entropy-Guided GRPOBackbone=Qwen2.5-VL-3B, Noise Level=0%2025.11 | 83.6 | — | — | — | — | — | — | — | — | — | — | — | |
| UI-R1Size=3B2025.10 | 83.3 | 95.6 | 84.7 | 90.2 | 59.3 | 85.2 | 73.3 | — | — | — | — | — | |
| Claude2025.02 | 83 | — | — | — | — | — | — | — | — | — | — | — | |
| SSL-7BProtocol=Reinforcement Learning, Paradigm=SSL, Training Data=Mix-3K2026.01 | 83 | — | — | 93.3 | 70 | 93 | 75.7 | — | — | — | — | — | |
| AGUVIS-7BParameters=7B2025.12 | 83 | 95.6 | 77.7 | 93.8 | 67.1 | 88.3 | 75.2 | — | — | — | — | — | |
| RL-Continuous-7BProtocol=Reinforcement Learning, Paradigm=RLVR, Training Data=Mix-3K2026.01 | 82.5 | — | — | 92.2 | 70 | 91.3 | 76.6 | — | — | — | — | — | |
| OS-AtlasSize=7B2025.10 | 82.5 | 93 | 72.9 | 91.8 | 62.9 | 90.89 | 74.3 | — | — | — | — | — | |
| OSAtlas-7BTraining Samples=13M2026.05 | 82.5 | — | — | — | — | — | — | 84.5 | 85 | 78.8 | — | — | |
| OS-Atlas-Base-7BPlanner=None2024.10 | 82.47 | 93.04 | 72.93 | 91.75 | 62.86 | 90.87 | 74.27 | — | — | — | — | — | |
| Two-Stage Entropy-Guided GRPOBackbone=Qwen2.5-VL-3B, Noise Level=20%2025.11 | 82.4 | — | — | — | — | — | — | — | — | — | — | — | |
| UI-TARS-2BTraining Samples=2M2026.05 | 82.3 | — | — | — | — | — | — | 79.8 | 85 | 81.4 | — | — | |
| GRPOBackbone=Qwen2.5-VL-3B, Noise Level=0%2025.11 | 82.2 | — | — | — | — | — | — | — | — | — | — | — | |
| RL-Binary-7BProtocol=Reinforcement Learning, Paradigm=RLVR, Training Data=Mix-3K2026.01 | 82.1 | — | — | 92.7 | 68.6 | 91.3 | 75.7 | — | — | — | — | — | |
| GRPOBackbone=Qwen2.5-VL-3B, Noise Level=20%2025.11 | 81.8 | — | — | — | — | — | — | — | — | — | — | — | |
| UGround-7BPlanner=GPT-4o2024.10 | 81.4 | 93.4 | 76.9 | 92.8 | 67.9 | 88.7 | 68.9 | — | — | — | — | — | |
| Two-Stage Entropy-Guided GRPOBackbone=Qwen2.5-VL-3B, Noise Level=50%2025.11 | 80.2 | — | — | — | — | — | — | — | — | — | — | — | |
| Os-Atlas-7BProtocol=Post-Training2026.01 | 79.9 | — | — | 91.7 | 62.8 | 90.8 | 74.2 | — | — | — | — | — | |
| SSL-3BProtocol=Reinforcement Learning, Paradigm=SSL, Training Data=Mix-3K2026.01 | 78.3 | — | — | 91.3 | 65 | 88.3 | 68.4 | — | — | — | — | — | |
| Two-Stage Entropy-Guided GRPOBackbone=Qwen2-VL-7B, Noise Level=0%2025.11 | 78 | — | — | — | — | — | — | — | — | — | — | — | |
| QwenVL2.5-7B*Protocol=Post-Training, SFT=Mix-3K2026.01 | 77.3 | — | — | 90.3 | 62.8 | 87.8 | 68.2 | — | — | — | — | — | |
| UI-R1-3BProtocol=Post-Training2026.01 | 77 | — | — | 90.2 | 59.3 | 85.2 | 73.3 | — | — | — | — | — | |
| OS-Atlas-Base-4BPlanner=GPT-4o2024.10 | 76.81 | 94.14 | 73.8 | 77.84 | 47.14 | 86.52 | 65.53 | — | — | — | — | — | |
| RL-Binary-3BProtocol=Reinforcement Learning, Paradigm=RLVR, Training Data=Mix-3K2026.01 | 76.8 | — | — | 90.3 | 61.4 | 87.4 | 67.9 | — | — | — | — | — | |
| RL-Continuous-3BProtocol=Reinforcement Learning, Paradigm=RLVR, Training Data=Mix-3K2026.01 | 76.6 | — | — | 93.3 | 62.8 | 84.3 | 66 | — | — | — | — | — | |
| Two-Stage Entropy-Guided GRPOBackbone=Qwen2-VL-7B, Noise Level=20%2025.11 | 76.6 | — | — | — | — | — | — | — | — | — | — | — | |
| GRPOBackbone=Qwen2.5-VL-3B, Noise Level=50%2025.11 | 76.2 | — | — | — | — | — | — | — | — | — | — | — | |
| Two-Stage Entropy-Guided GRPOBackbone=Qwen2.5-VL-3B, Noise Level=100%2025.11 | 75.8 | — | — | — | — | — | — | — | — | — | — | — | |
| QwenVL2.5-7BProtocol=Zero Shot2026.01 | 75.4 | — | — | 89.7 | 60 | 86.9 | 65.1 | — | — | — | — | — | |
| GRPOBackbone=Qwen2-VL-7B, Noise Level=0%2025.11 | 75.4 | — | — | — | — | — | — | — | — | — | — | — | |
| ShowUISize=2B, #Train=256K, Zero-shot=true2024.11 | 75.1 | — | — | — | — | — | — | — | — | — | — | — | |
| ShowUI-2BTraining Samples=256K2026.05 | 75.1 | — | — | — | — | — | — | 76.2 | 84.8 | 70.8 | — | — | |
| ShowUI-GSize=2B, #Train=119K, Zero-shot=true, Training Mode=Grounding data only2024.11 | 74.9 | — | — | — | — | — | — | — | — | — | — | — | |
| UGround-7BProtocol=Post-Training2026.01 | 74.2 | — | — | 82.5 | 63.6 | 80.4 | 70.4 | — | — | — | — | — | |
| FOCUS-2BProtocol=Post-Training2026.01 | 74 | — | — | 80.9 | 65 | 81.7 | 68.5 | — | — | — | — | — | |
| GRPOBackbone=Qwen2-VL-7B, Noise Level=20%2025.11 | 74 | — | — | — | — | — | — | — | — | — | — | — | |
| UGround-7BPlanner=None2024.10 | 73.3 | 82.8 | 60.3 | 82.5 | 63.6 | 80.4 | 70.4 | — | — | — | — | — | |
| UGroundSize=7B, #Train=1.3M, Zero-shot=true2024.11 | 73.3 | — | — | — | — | — | — | — | — | — | — | — | |
| UGroundSize=7B2025.10 | 73.3 | 82.8 | 60.3 | 82.5 | 63.6 | 80.4 | 70.4 | — | — | — | — | — | |
| UGround-7BTraining Samples=10M2026.05 | 73.3 | — | — | — | — | — | — | 78.3 | 75.9 | 75.8 | — | — | |
| OmniParserLocal Semantics (LS)=true, Grounding Model=Finetuned Interactable Region Detection (ID)2024.08 | 73 | 93.9 | 57 | 91.3 | 63.6 | 81.3 | 51 | — | — | — | — | — | |
| OmniParserBackbone=GPT-4V, Zero-shot=true2024.11 | 73 | — | — | — | — | — | — | — | — | — | — | — | |
| GRPOBackbone=Qwen2.5-VL-3B, Noise Level=100%2025.11 | 71 | — | — | — | — | — | — | — | — | — | — | — | |
| ShowUI-2BProtocol=Post-Training2026.01 | 70.7 | — | — | 76.3 | 61.1 | 81.7 | 63.6 | — | — | — | — | — | |
| BaseBackbone=Qwen2.5-VL-3B, Noise Level=-2025.11 | 70.6 | — | — | — | — | — | — | — | — | — | — | — | |
| OS-Atlas-Base-4BPlanner=None2024.10 | 70.13 | 85.71 | 58.52 | 72.16 | 45.71 | 82.61 | 63.11 | — | — | — | — | — | |
| Two-Stage Entropy-Guided GRPOBackbone=Qwen2-VL-7B, Noise Level=50%2025.11 | 69.8 | — | — | — | — | — | — | — | — | — | — | — | |
| Two-Stage Entropy-Guided GRPOBackbone=InternVL-3.5-2B, Noise Level=20%2025.11 | 69.8 | — | — | — | — | — | — | — | — | — | — | — | |
| Two-Stage Entropy-Guided GRPOBackbone=InternVL-3.5-2B, Noise Level=0%2025.11 | 69.8 | — | — | — | — | — | — | — | — | — | — | — | |
| GRPO w. Two.Backbone=InternVL-3.5-2B, Annotation noise level=20%2025.11 | 69.8 | — | — | — | — | — | — | — | — | — | — | — | |
| GRPO w. Two.Backbone=InternVL-3.5-2B, Annotation noise level=0%2025.11 | 69.8 | — | — | — | — | — | — | — | — | — | — | — | |
| GRPO w. Min.Backbone=InternVL-3.5-2B, Annotation noise level=0%2025.11 | 69.2 | — | — | — | — | — | — | — | — | — | — | — | |
| OmniParserLocal Semantics (LS)=true, Grounding Model=Grounding DINO (GD)2024.08 | 68.7 | 94.8 | 53.7 | 89.3 | 44.9 | 83 | 45.1 | — | — | — | — | — | |
| OSAtlas-4BTraining Samples=13M2026.05 | 68.5 | — | — | — | — | — | — | 69.9 | 56.2 | 74.9 | — | — | |
| GRPOBackbone=InternVL-3.5-2B, Noise Level=0%2025.11 | 66.8 | — | — | — | — | — | — | — | — | — | — | — | |
| GRPOBackbone=InternVL-3.5-2B, Annotation noise level=0%2025.11 | 66.8 | — | — | — | — | — | — | — | — | — | — | — | |
| R-VLMBackbone=Qwen-VL 9.6B, Training method=LoRA2025.07 | 66.3 | 85 | 61.1 | 81.4 | 52.8 | 66.5 | 51.4 | — | — | — | — | — | |
| GRPO w. Max.Backbone=InternVL-3.5-2B, Annotation noise level=20%2025.11 | 66 | — | — | — | — | — | — | — | — | — | — | — | |
| Os-Atlas-4BProtocol=Post-Training2026.01 | 65.9 | — | — | 72.1 | 45.7 | 82.6 | 63.1 | — | — | — | — | — | |
| GRPO w. Min.Backbone=InternVL-3.5-2B, Annotation noise level=20%2025.11 | 65.6 | — | — | — | — | — | — | — | — | — | — | — | |
| GRPO w. Max.Backbone=InternVL-3.5-2B, Annotation noise level=0%2025.11 | 65.2 | — | — | — | — | — | — | — | — | — | — | — | |
| QwenVL2.5-3B*Protocol=Post-Training, SFT=Mix-3K2026.01 | 63.4 | — | — | 85.7 | 46.2 | 73 | 48.5 | — | — | — | — | — | |
| GRPOBackbone=InternVL-3.5-2B, Noise Level=20%2025.11 | 63.2 | — | — | — | — | — | — | — | — | — | — | — | |
| GRPOBackbone=InternVL-3.5-2B, Annotation noise level=20%2025.11 | 63.2 | — | — | — | — | — | — | — | — | — | — | — | |
| SeeClickPlanner=GPT-4o2024.10 | 62.89 | 83.52 | 59.39 | 82.47 | 35 | 66.96 | 35.44 | — | — | — | — | — | |
| GRPOBackbone=Qwen2-VL-7B, Noise Level=50%2025.11 | 61.2 | — | — | — | — | — | — | — | — | — | — | — | |
| GRPO w. Min.Backbone=InternVL-3.5-2B, Annotation noise level=50%2025.11 | 59.8 | — | — | — | — | — | — | — | — | — | — | — | |
| Two-Stage Entropy-Guided GRPOBackbone=InternVL-3.5-2B, Noise Level=50%2025.11 | 59.2 | — | — | — | — | — | — | — | — | — | — | — | |
| GRPO w. Two.Backbone=InternVL-3.5-2B, Annotation noise level=50%2025.11 | 59.2 | — | — | — | — | — | — | — | — | — | — | — | |
| OmniParserLocal Semantics (LS)=false, Grounding Model=Grounding DINO (GD)2024.08 | 58.38 | 92.7 | 49.4 | 64.9 | 26.3 | 77.3 | 39.7 | — | — | — | — | — | |
| GRPO w. Max.Backbone=InternVL-3.5-2B, Annotation noise level=50%2025.11 | 57.6 | — | — | — | — | — | — | — | — | — | — | — | |
| GRPOBackbone=InternVL-3.5-2B, Noise Level=50%2025.11 | 56.8 | — | — | — | — | — | — | — | — | — | — | — | |
| GRPOBackbone=InternVL-3.5-2B, Annotation noise level=50%2025.11 | 56.8 | — | — | — | — | — | — | — | — | — | — | — | |
| Two-Stage Entropy-Guided GRPOBackbone=Qwen2-VL-2B, Noise Level=0%2025.11 | 55.6 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL-3BTraining Samples=-2026.05 | 55.5 | — | — | — | — | — | — | — | — | — | — | — |