Agentic Tool-Use on ACEBench
37.5ACE-E ScoreGEAR
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GEARBackbone=Qwen3-8B, Evaluation Context=Standard2026.05 | 37.5 | 45.3 | — | |
| GEARBackbone=Qwen3-4B, Evaluation Context=Standard2026.05 | 36.3 | 38.9 | — | |
| MT-GRPOBackbone=Qwen3-8B, Evaluation Context=Standard2026.05 | 33.7 | 37.8 | — | |
| MT-GRPOBackbone=Qwen3-4B, Evaluation Context=Standard2026.05 | 32.8 | 35.9 | — | |
| ARPOBackbone=Qwen3-8B, Evaluation Context=Standard2026.05 | 32.5 | 36.3 | — | |
| GEARBackbone=Qwen3-4B, Evaluation Context=Multi-Domain Generalization2026.05 | 32.5 | 36.4 | — | |
| BaseBackbone=Qwen3-8B, Evaluation Context=Standard2026.05 | 31.5 | 35.7 | — | |
| GRPOBackbone=Qwen3-8B, Evaluation Context=Standard2026.05 | 31.3 | 35.8 | — | |
| GRPOBackbone=Qwen3-4B, Evaluation Context=Standard2026.05 | 31.2 | 35.6 | — | |
| GRPOBackbone=Qwen3-4B, Evaluation Context=Multi-Domain Generalization2026.05 | 30.7 | 33.4 | — | |
| OPSD+RLBackbone=Qwen3-8B, Evaluation Context=Standard2026.05 | 30.5 | 32.9 | — | |
| BaseBackbone=Qwen3-4B, Evaluation Context=Standard2026.05 | 26.5 | 30.2 | — | |
| OPSDBackbone=Qwen3-8B, Evaluation Context=Standard2026.05 | 26.5 | 28.8 | — | |
| OPSD+RLBackbone=Qwen3-4B, Evaluation Context=Standard2026.05 | 24.3 | 28.1 | — | |
| ARPOBackbone=Qwen3-4B, Evaluation Context=Standard2026.05 | 23.6 | 26.2 | — | |
| OPSDBackbone=Qwen3-4B, Evaluation Context=Standard2026.05 | 22.5 | 26.5 | — | |
| DeepSeek-V3.1Model Identifier=DeepSeek-V3.12026.05 | — | — | 0.408 | |
| Gemini-2.5Model Identifier=gemini-2.5-pro-preview-05-062026.05 | — | — | 0.634 | |
| GPT-5Model Identifier=gpt-5-2025-08-072026.05 | — | — | 0.325 | |
| Kimi-K2Model Identifier=Kimi-K22026.05 | — | — | 0.65 | |
| MAVENBase Model=GPT-OSS-120b2026.05 | — | — | 0.75 | |
| o3Model Identifier=o3-2025-04-162026.05 | — | — | 0.633 | |
| o4-miniModel Identifier=o4-mini-2025-04-162026.05 | — | — | 0.6 | |
| Qwen3-Th-235BModel Identifier=Qwen3-235B2026.05 | — | — | 0.391 |