Multi-turn Tool-use on BFCL multi-turn V3
68Average Success RateGLM-4.6 355B
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| GLM-4.6 355BModel Type=Open-source2026.05 | 68 | 74.5 | 68 | 63 | 66.5 | |
| Claude-Sonnet-4.5Model Type=Closed-source2026.05 | 61.38 | 69 | 65 | 52.5 | 59 | |
| Claude-Sonnet-4-5-20250929Model Category=Closed-source Model2026.06 | 61.38 | 69 | 65 | 52.5 | 59 | |
| Gemini-3-Pro-PreviewModel Type=Closed-source2026.05 | 60.75 | 64.5 | 60 | 54.5 | 64 | |
| Gemini-3-Pro-PreviewModel Category=Closed-source Model2026.06 | 60.75 | 64.5 | 60 | 54.5 | 64 | |
| Grok-4-1-fast-reasoningModel Category=Closed-source Model2026.06 | 58.88 | 70.5 | 59.5 | 43 | 62.5 | |
| xLAM-2-3b-fc-rModel Type=Open-source2026.05 | 58.38 | 71.5 | 59 | 57.5 | 45.5 | |
| xLAM-2-3b-fc-rSize=3B, Model Category=Open-source Model2026.06 | 58.38 | 71.5 | 59 | 57.5 | 45.5 | |
| Qwen3-4B-Instruct + RODS (ours)Size=4B, Strategy=RODS2026.06 | 56 | 68 | 59 | 44 | 53 | |
| Claude-Haiku-4-5-20251001Model Category=Closed-source Model2026.06 | 53.63 | 63.5 | 42.5 | 52.5 | 56 | |
| Nanbeige4-3B-Thinking-2511Size=3B, Model Category=Open-source Model2026.06 | 51.12 | 58.5 | 54 | 45 | 47 | |
| Kimi-K2-InstructSize=1043B, Model Category=Open-source Model2026.06 | 50.63 | 62 | 41 | 44.5 | 55 | |
| Qwen3-4B-Instruct + EnvTuningSize=4B, Strategy=EnvTuning2026.06 | 50.5 | 64 | 52 | 35 | 51 | |
| Qwen3-4B-Instruct + Static datasetSize=4B, Strategy=Static dataset2026.06 | 50 | 62 | 51 | 35 | 52 | |
| Qwen3-32BModel Type=Open-source2026.05 | 47.88 | 56 | 52.5 | 40 | 43 | |
| Qwen3-32BSize=32B, Model Category=Open-source Model2026.06 | 47.88 | 56 | 52.5 | 40 | 43 | |
| ToolACE-2-Llama-3.1-8B + SEALBackbone=ToolACE-2-Llama-3.1-8B, Method Variant=SEAL, Training Budget=400-sample2026.05 | 46.75 | 58 | 46 | 44 | 39 | |
| Grok-4-1-fast-non-reasoningModel Category=Closed-source Model2026.06 | 46.75 | 58 | 39.5 | 37.5 | 52 | |
| DeepSeek-V3.2-ExpSize=671B, Model Category=Open-source Model2026.06 | 44.88 | 55 | 49 | 27 | 48.5 | |
| Qwen3-235B-A22B-InstructSize=235B, Model Category=Open-source Model2026.06 | 44.63 | 54 | 42.5 | 31.5 | 50.5 | |
| GPT-4o-2024-11-20Model Type=Closed-source2026.05 | 42.5 | 55.5 | 34.5 | 29 | 51 | |
| GPT-4o-2024-11-20Model Category=Closed-source Model2026.06 | 42.5 | 55.5 | 34.5 | 29 | 51 | |
| Qwen2.5-7B-Instruct + SEALBackbone=Qwen2.5-7B-Instruct, Method Variant=SEAL, Training Budget=400-sample2026.05 | 40.25 | 58 | 36 | 34 | 33 | |
| ToolACE-MTSize=8B, Model Category=Open-source Model2026.06 | 40.25 | 57.5 | 31.5 | 34 | 38 | |
| Qwen2.5-7B-Instruct + RODS (ours)Size=7B, Strategy=RODS2026.06 | 40.25 | 54 | 43.5 | 33.5 | 30 | |
| ToolACE-2-Llama-3.1-8B + Vanilla RLBackbone=ToolACE-2-Llama-3.1-8B, Method Variant=Vanilla RL, Training Budget=400-sample2026.05 | 38.5 | 52 | 30 | 40 | 32 | |
| ToolACE-2-8BSize=8B, Model Category=Open-source Model2026.06 | 38.38 | 49 | 28 | 30.5 | 46 | |
| Qwen2.5-7B-Instruct + EnvTuningSize=7B, Strategy=EnvTuning2026.06 | 37.75 | 51.5 | 41 | 30.5 | 28 | |
| Qwen2.5-7B-Instruct + Static datasetSize=7B, Strategy=Static dataset2026.06 | 36.92 | 50.33 | 40.33 | 29.33 | 27.67 | |
| Gemini-2.5-FlashModel Type=Closed-source2026.05 | 36.25 | 41.5 | 36 | 32 | 35.5 | |
| ToolACE-2-Llama-3.1-8B (Original)Backbone=ToolACE-2-Llama-3.1-8B, Method Variant=Original2026.05 | 32 | 45 | 26 | 35 | 22 | |
| Llama-3.1-8B-Instruct + RODS (ours)Size=8B, Strategy=RODS2026.06 | 30.88 | 32 | 28 | 24.5 | 39 | |
| Qwen2.5-7B-Instruct + Vanilla RLBackbone=Qwen2.5-7B-Instruct, Method Variant=Vanilla RL, Training Budget=400-sample2026.05 | 30.75 | 46 | 27 | 27 | 23 | |
| Qwen3-30B-A3B-ThinkingSize=30B, Model Category=Open-source Model2026.06 | 30 | 43.5 | 10.5 | 25 | 41 | |
| Gemini-2.5-Pro-PreviewModel Category=Closed-source Model2026.06 | 28.75 | 32 | 29 | 22 | 32 | |
| Llama-3.1-8B-Instruct + EnvTuningSize=8B, Strategy=EnvTuning2026.06 | 28.38 | 28 | 25.5 | 23 | 37 | |
| Llama-3.1-8B-Instruct + Static datasetSize=8B, Strategy=Static dataset2026.06 | 28.25 | 28.2 | 25.85 | 22.15 | 36.8 | |
| GPT-5.2-2025-12-11Model Category=Closed-source Model2026.06 | 28.13 | 36.5 | 18 | 27.5 | 30.5 | |
| Qwen2.5-14B-InstructModel Type=Open-source2026.05 | 25.25 | 33 | 26 | 22 | 20 | |
| Qwen3-4B-InstructSize=4B2026.06 | 22.13 | 26.5 | 21 | 15.5 | 25.5 | |
| Llama-4-MaverickSize=400B, Model Category=Open-source Model2026.06 | 20.25 | 27 | 22 | 14 | 18 | |
| Gemini-2.5-FlashModel Category=Closed-source Model2026.06 | 16.75 | 14.5 | 16.5 | 17.5 | 18.5 | |
| Qwen2.5-3B-Instruct + SEALBackbone=Qwen2.5-3B-Instruct, Method Variant=SEAL, Training Budget=400-sample2026.05 | 14 | 19 | 15 | 12 | 10 | |
| Qwen2.5-7B-Instruct (Original)Backbone=Qwen2.5-7B-Instruct, Method Variant=Original2026.05 | 14 | 22 | 14 | 10 | 10 | |
| Qwen2.5-3B-Instruct + Vanilla RLBackbone=Qwen2.5-3B-Instruct, Method Variant=Vanilla RL, Training Budget=400-sample2026.05 | 9.25 | 16 | 9 | 6 | 6 | |
| Qwen2.5-7B-InstructSize=7B2026.06 | 7 | 9.33 | 9.33 | 6.33 | 3 | |
| Qwen2.5-3B-Instruct (Original)Backbone=Qwen2.5-3B-Instruct, Method Variant=Original2026.05 | 5.75 | 11 | 6 | 3 | 3 | |
| Llama-3.1-8B-InstructSize=8B2026.06 | 5.48 | 6.15 | 6.8 | 3.2 | 5.75 |