Single-step action prediction on Cloud Console Benchmark 400-sample
95.25AccuracyGemini 3 Pro Preview
Evaluation Results
| Method | Links | |
|---|---|---|
| Gemini 3 Pro PreviewModel Category=Frontier2026.06 | 95.25 | |
| GPT-5.5Model Category=Frontier2026.06 | 94.75 | |
| AliyunConsoleAgent-32BModel Category=AliyunConsoleAgent (Ours), Training Protocol=SFT, Base Model=Qwen3-VL-32B2026.06 | 92.75 | |
| Kimi K2.6Model Category=Frontier2026.06 | 92.33 | |
| Qwen3.6-PlusModel Category=Frontier2026.06 | 90.75 | |
| AliyunConsoleAgent (Qwen3-VL-8B)Model Category=AliyunConsoleAgent (Ours), Training Protocol=SFT, Base Model=Qwen3-VL-8B2026.06 | 88.84 | |
| Qwen3-VL-32B-InstructModel Category=Open-Source Base2026.06 | 83 | |
| Qwen3-VL-8B-InstructModel Category=Open-Source Base2026.06 | 71.17 |