Mobile UI Control on AndroidWorld
71.6Overall Task Success RateGUI-Owl-7B w/ Android Coach
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| GUI-Owl-7B w/ Android Coach#Params=7B2026.04 | 71.6 | 90.7 | 58.3 | 35.1 | |
| GUI-Owl-7B w/ GRPO#Params=7B2026.04 | 68.1 | 86.3 | 55.6 | 33.3 | |
| GUI-Owl-7B w/ PPO#Params=7B2026.04 | 68.1 | 85.2 | 57.4 | 33.3 | |
| GUI-Owl-7BOnline Data=Y, Training Steps=–, Training Trajectories=–2026.04 | 66.4 | — | — | — | |
| GUI-Owl-7B Base Model#Params=7B2026.04 | 64.7 | 84.2 | 50.9 | 28.1 | |
| GLM-4.5V (106B)Online Data=–, Training Steps=–, Training Trajectories=–2026.04 | 57 | — | — | — | |
| UI-Venus-Navi-7BOnline Data=N, Training Steps=350K, Training Trajectories=–2026.04 | 49.1 | — | — | — | |
| UI-TARS-72B-SFTOnline Data=Y, Training Steps=–, Training Trajectories=145K2026.04 | 46.6 | — | — | — | |
| Qwen2.5-7B-Instruct-ASLType=Fine-Tuned, Agent=T3A*, Model=Qwen2.5-7B-Instruct-ASL (Ours), Input Modality=Text2025.06 | 46.3 | 72.2 | 32.1 | 12.5 | |
| T3A* (SFT)Type=Fine-Tuned, Agent=T3A*, Model=Qwen2.5-7B-Instruct-SFT, Input Modality=Text2025.06 | 42.5 | 72.2 | 26.8 | 3.13 | |
| T3AType=Prompt-Driven, Agent=T3A, Model=GPT-4o, Input Modality=Image2025.06 | 41.9 | 64.9 | 26.2 | 14.6 | |
| GLM-4.1V-ThinkingOnline Data=–, Training Steps=–, Training Trajectories=–2026.04 | 41.7 | — | — | — | |
| UI-TARS-1.5-7B w/ Android Coach#Params=7B2026.04 | 41.1 | 56.3 | 27.8 | 17.5 | |
| Claude-Sonnet-4 (SoM)#Params=-2026.04 | 41 | — | — | — | |
| UI-TARS-72B-DPO#Params=72B2026.04 | 40.5 | 57.4 | 27.8 | 10.5 | |
| UI-TARS-1.5-7B w/ GRPO#Params=7B2026.04 | 38.2 | 51.9 | 28.7 | 12.3 | |
| UI-TARS-1.5-7B w/ PPO#Params=7B2026.04 | 37.4 | 50.8 | 26.9 | 14 | |
| M3AType=Prompt-Driven, Agent=M3A, Model=GPT-4o, Input Modality=Image + Text2025.06 | 36.6 | 60.5 | 20.2 | 8.3 | |
| GPT-4o (SoM)#Params=-2026.04 | 34.5 | — | — | — | |
| SOLAR-RLOnline Data=N, Training Steps=94K, Training Trajectories=15K2026.04 | 33.7 | — | — | — | |
| UI-TARS-7B-SFTOnline Data=Y, Training Steps=–, Training Trajectories=145K2026.04 | 33.3 | — | — | — | |
| UI-TARS-1.5-7B Base Model#Params=7B2026.04 | 32.8 | 43.7 | 25.9 | 10.5 | |
| Qwen2.5-VL Agent + VisCriticbase_agent=Qwen2.5-VL Agent, critic_module=VisCritic2026.06 | 29.8 | — | — | — | |
| Qwen2.5-VL Agent + GUI-Critic-R1base_agent=Qwen2.5-VL Agent, critic_module=GUI-Critic-R12026.06 | 28.2 | — | — | — | |
| Qwen2.5-VL Agent + GuidNavbase_agent=Qwen2.5-VL Agent, critic_module=GuidNav2026.06 | 27.8 | — | — | — | |
| Qwen2.5-VL Agent + GUI-PRAbase_agent=Qwen2.5-VL Agent, critic_module=GUI-PRA2026.06 | 27 | — | — | — | |
| Aguvis-72BOnline Data=N, Training Steps=–, Training Trajectories=35K2026.04 | 26.1 | — | — | — | |
| Qwen2.5VL-32B-Instruct#Params=32B2026.04 | 25.9 | 37.2 | 14.8 | 10.5 | |
| ShowUI + VisCriticbase_agent=ShowUI, critic_module=VisCritic2026.06 | 24.8 | — | — | — | |
| Qwen2.5-VL Agentbase_agent=Qwen2.5-VL Agent2026.06 | 24.6 | — | — | — | |
| Gemini-Pro-1.5 (SoM)#Params=-2026.04 | 22.8 | — | — | — | |
| ShowUI + GUI-PRAbase_agent=ShowUI, critic_module=GUI-PRA2026.06 | 22.1 | — | — | — | |
| SeeActType=Prompt-Driven, Agent=SeeAct, Model=GPT-4o, Input Modality=Image + Text2025.06 | 22 | 34.2 | 15.5 | 4.2 | |
| CogAgent + VisCriticbase_agent=CogAgent, critic_module=VisCritic2026.06 | 20.8 | — | — | — | |
| ShowUIbase_agent=ShowUI2026.06 | 19.8 | — | — | — | |
| SeeClick + VisCriticbase_agent=SeeClick, critic_module=VisCritic2026.06 | 19.1 | — | — | — | |
| SeeClick + GUI-Critic-R1base_agent=SeeClick, critic_module=GUI-Critic-R12026.06 | 18.3 | — | — | — | |
| SeeClick + BacktrackAgentbase_agent=SeeClick, critic_module=BacktrackAgent2026.06 | 17.9 | — | — | — | |
| OS-Genesis-7B-AW#Params=7B2026.04 | 17.8 | 26.8 | 11.1 | 1.8 | |
| AgentCPM-GUI-8B#Params=8B2026.04 | 17.5 | 29 | 5.6 | 3.5 | |
| CogAgentbase_agent=CogAgent2026.06 | 17.5 | — | — | — | |
| SeeClick + GUI-PRAbase_agent=SeeClick, critic_module=GUI-PRA2026.06 | 17.1 | — | — | — | |
| Qwen2.5-VL-7B-Instruct#Params=7B2026.04 | 14.9 | 23.5 | 6.5 | 3.5 | |
| SeeClickbase_agent=SeeClick2026.06 | 14.7 | — | — | — | |
| T3A*Type=Prompt-Driven, Agent=T3A*, Model=Qwen2.5-7B-Instruct, Input Modality=Text2025.06 | 2.5 | 5.6 | 0 | 0 |