Multi-Step Tool Orchestration on ComplexFuncBench 1.0 (test)
36.4Turn AccuracyQwen3-8B-1e-6
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3-8B-1e-6Model-LR=Qwen3-8B-1e-6, Training Strategy=GRPO, Samples per workflow template=32026.03 | 36.4 | |
| Qwen3-8B-1e-6Model-LR=Qwen3-8B-1e-6, Training Strategy=Zero-shot, Samples per workflow template=02026.03 | 34.4 | |
| Qwen3-8B-1e-6Model-LR=Qwen3-8B-1e-6, Training Strategy=SFT, Samples per workflow template=32026.03 | 32.6 | |
| Qwen3-8B-1e-6Model-LR=Qwen3-8B-1e-6, Training Strategy=SFT, Samples per workflow template=102026.03 | 31.7 | |
| Qwen3-8B-1e-6Model-LR=Qwen3-8B-1e-6, Training Strategy=SFT+GRPO, Samples per workflow template=32026.03 | 31.1 | |
| Qwen3-8B-1e-6Model-LR=Qwen3-8B-1e-6, Training Strategy=GRPO, Samples per workflow template=102026.03 | 30.2 | |
| Qwen3-8B-1e-6Model-LR=Qwen3-8B-1e-6, Training Strategy=SFT+GRPO, Samples per workflow template=102026.03 | 29.9 | |
| Qwen3-8B-5e-6Model-LR=Qwen3-8B-5e-6, Training Strategy=GRPO, Samples per workflow template=102026.03 | 29.6 | |
| Qwen3-8B-5e-6Model-LR=Qwen3-8B-5e-6, Training Strategy=GRPO, Samples per workflow template=32026.03 | 28.7 | |
| Qwen3-8B-5e-6Model-LR=Qwen3-8B-5e-6, Training Strategy=SFT+GRPO, Samples per workflow template=102026.03 | 28.4 | |
| Qwen2.5-7B-1e-6Model-LR=Qwen2.5-7B-1e-6, Training Strategy=SFT+GRPO, Samples per workflow template=32026.03 | 27.1 | |
| Qwen2.5-7B-1e-6Model-LR=Qwen2.5-7B-1e-6, Training Strategy=GRPO, Samples per workflow template=32026.03 | 25.8 | |
| Qwen2.5-7B-5e-6Model-LR=Qwen2.5-7B-5e-6, Training Strategy=GRPO, Samples per workflow template=32026.03 | 24.1 | |
| Qwen2.5-7B-1e-6Model-LR=Qwen2.5-7B-1e-6, Training Strategy=GRPO, Samples per workflow template=102026.03 | 23.9 | |
| Qwen3-8B-5e-6Model-LR=Qwen3-8B-5e-6, Training Strategy=SFT+GRPO, Samples per workflow template=32026.03 | 22.9 | |
| Qwen2.5-7B-5e-6Model-LR=Qwen2.5-7B-5e-6, Training Strategy=GRPO, Samples per workflow template=102026.03 | 21.3 | |
| Qwen2.5-7B-5e-6Model-LR=Qwen2.5-7B-5e-6, Training Strategy=SFT+GRPO, Samples per workflow template=32026.03 | 16.5 | |
| Qwen2.5-7B-1e-6Model-LR=Qwen2.5-7B-1e-6, Training Strategy=SFT+GRPO, Samples per workflow template=102026.03 | 16 | |
| Qwen2.5-7B-1e-6Model-LR=Qwen2.5-7B-1e-6, Training Strategy=SFT, Samples per workflow template=102026.03 | 13.5 | |
| Qwen2.5-7B-1e-6Model-LR=Qwen2.5-7B-1e-6, Training Strategy=Zero-shot, Samples per workflow template=02026.03 | 13.4 | |
| Qwen2.5-7B-5e-6Model-LR=Qwen2.5-7B-5e-6, Training Strategy=SFT+GRPO, Samples per workflow template=102026.03 | 12.1 | |
| Qwen2.5-7B-1e-6Model-LR=Qwen2.5-7B-1e-6, Training Strategy=SFT, Samples per workflow template=32026.03 | 11.5 |