Agentic Workflow Success on τ2-bench
76.5Airline Success RateLongCat-Flash-Thinking-2601
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| LongCat-Flash-Thinking-2601User Simulator=GPT-4.1-2025-04-142026.02 | 76.5 | 88.6 | 99.3 | — | — | 88.2 | |
| Gemini-3-proUser Simulator=Gemini-3-pro2026.02 | 73 | 85.3 | 98 | — | — | 85.4 | |
| Claude-Sonnet-4.52026.02 | 70 | 86.2 | 98 | — | — | 84.7 | |
| Qwen3-Max-ThinkingUser Simulator=GPT-4.1-2025-04-142026.02 | 69 | 79.4 | 98.2 | — | — | 82.2 | |
| LongCat-Flash-ThinkingUser Simulator=GPT-4.1-2025-04-142026.02 | 67.5 | 71.5 | 83.1 | — | — | 74 | |
| LOGIGEN-32B(RL)User Simulator=DeepSeek-V3.22026.02 | 64.7 | 85.1 | 88.6 | — | — | 79.5 | |
| LOGIGEN-32B(SFT)User Simulator=DeepSeek-V3.22026.02 | 64 | 82 | 86.6 | — | — | 77.5 | |
| DeepSeek-V3.2-ThinkingUser Simulator=DeepSeek-V3.22026.02 | 63.8 | 81.1 | 96.2 | — | — | 80.4 | |
| GPT-5User Simulator=GPT-4.1-2025-04-142026.02 | 62.5 | 81.6 | 95.8 | — | — | 80 | |
| LOGIGEN-8B(SFT)User Simulator=DeepSeek-V3.22026.02 | 61.5 | 74.1 | 80.7 | — | — | 72.1 | |
| InfTool2025.12 | 60 | 67 | 67.5 | — | — | — | |
| AgentScaler-30B-A3B2026.02 | 60 | 70.2 | 55.3 | — | — | 61.8 | |
| Qwen3-MaxUser Simulator=GPT-4.1-2025-04-142026.02 | 59.5 | 72.2 | 84.2 | — | — | 72 | |
| Qwen3-235B-A22B-Thinking-25072026.02 | 58 | 71.9 | 45.6 | — | — | 58.7 | |
| Kimi-K2-Thinking2026.02 | 56.5 | 70.6 | 65.8 | — | — | 64.3 | |
| AgentScaler-4B2026.02 | 56 | 62.3 | 48.2 | — | — | 55.5 | |
| LOGIGEN-8B(RL)User Simulator=DeepSeek-V3.22026.02 | 54.7 | 79.5 | 81.3 | — | — | 71.8 | |
| Nemotron-3-Nano-30B-A3B2026.02 | 48 | 56.9 | 42.2 | — | — | 49 | |
| MUA-RL-32BUser Simulator=GPT-4.1-2025-04-142026.02 | 45.4 | 67.3 | 28.3 | — | — | 47 | |
| DeepSeek-V3.1-Terminus-ThinkingUser Simulator=DeepSeek-V3.22026.02 | 44 | 65.4 | 23.7 | — | — | 44.4 | |
| AgentScaler-8B2026.02 | 44 | 58.8 | 45.4 | — | — | 49.4 | |
| Qwen3-32BUser Simulator=DeepSeek-V3.2, Note=Internal evaluation2026.02 | 42 | 53.3 | 26.9 | — | — | 40.7 | |
| AWMModel Size=8B2026.02 | 38.5 | 41.23 | 23.47 | 33.45 | 55.4 | — | |
| GPT-OSS-120B-A5B2026.02 | 38 | 57 | 45.6 | — | — | 46.9 | |
| EnvScaler-8BUser Simulator=GPT-4.1-2025-04-142026.02 | 36 | 53.6 | — | — | — | — | |
| LongCat-GEM-32BUser Simulator=GPT-4.1-2025-04-142026.02 | 35.5 | 55.5 | — | — | — | — | |
| SimulatorModel Size=8B2026.02 | 34 | 32.24 | 29.17 | 31.3 | 54.32 | — | |
| EnvScaler-4BUser Simulator=GPT-4.1-2025-04-142026.02 | 34 | 48.1 | — | — | — | — | |
| EnvScalerModel Size=4B2026.02 | 31.5 | 44.3 | 12.5 | 28.96 | 51.8 | — | |
| EnvScalerModel Size=8B2026.02 | 31.5 | 49.56 | 32.68 | 39.39 | 63.31 | — | |
| AWMModel Size=14B2026.02 | 31.5 | 63.6 | 17.76 | 39.03 | 57.19 | — | |
| Qwen3-8BUser Simulator=DeepSeek-V3.2, Note=Internal evaluation2026.02 | 30.5 | 38.6 | 23.3 | — | — | 30.8 | |
| BaseModel Size=14B2026.02 | 27 | 65.35 | 12.28 | 36.69 | 55.4 | — | |
| BaseModel Size=8B2026.02 | 26.5 | 34.43 | 18.42 | 26.44 | 50.72 | — | |
| Instruct baseline2025.12 | 26 | 40 | 21 | — | — | — | |
| TOUCAN-32BUser Simulator=GPT-4o2026.02 | 22 | 52.6 | 20.2 | — | — | 31.6 | |
| LongCat-GEM-8BUser Simulator=GPT-4.1-2025-04-142026.02 | 22 | 44.5 | — | — | — | — | |
| SimulatorModel Size=14B2026.02 | 21.5 | 48.9 | 18.2 | 31.39 | 55.4 | — | |
| BaseModel Size=4B2026.02 | 21 | 19.96 | 9.43 | 15.83 | 34.89 | — | |
| TOUCAN-7BUser Simulator=GPT-4o2026.02 | 20 | 22.8 | 10.5 | — | — | 17.7 | |
| AWMModel Size=4B2026.02 | 19 | 30.26 | 16.45 | 22.57 | 43.89 | — | |
| MUA-RL-8BUser Simulator=GPT-4.1-2025-04-142026.02 | 19 | 49.8 | 21.8 | — | — | 30.2 | |
| SimulatorModel Size=4B2026.02 | 15.5 | 18.64 | 7.9 | 13.67 | 35.25 | — | |
| GLM-4.7-Thinking2026.02 | — | — | — | — | — | 87.4 | |
| MiniMax-M22026.02 | — | — | 87 | — | — | 77.2 |