Tool-Calling and Answer Generation on APIGen-MT (test)
90.18Action RecallQwen3-1.7B + RL
Evaluation Results
| Method | Links | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Qwen3-1.7B + RLModel Size=1.7B, Training=Reinforcement Learning (Reasoning-Action Synergy), Thinking Stage=Reasoning included2025.12 | 90.18 | 91.41 | 89.93 | 90.66 | 85.07 | 72.55 | 89.03 | 90.63 | 89.82 | 97.51 | |
| Qwen3-1.7B + SFT (no think)Model Size=1.7B, Training=Supervised Fine-Tuning (no thinking), Thinking Stage=None2025.12 | 89.12 | 94.63 | 85.83 | 90.02 | 85.12 | 72.34 | 83.3 | 93.55 | 88.13 | 97.31 | |
| Qwen3-1.7B + SFT (Cold Start - think)Model Size=1.7B, Training=Cold-start Supervised Fine-Tuning, Thinking Stage=Reasoning included2025.12 | 87.45 | 89.06 | 87.26 | 88.15 | 80.7 | 70.98 | 86.12 | 88.07 | 87.08 | 96.5 | |
| Qwen3-1.7BModel Size=1.7B, Training=Base2025.12 | 58.9 | 61.77 | 60.84 | 61.3 | 37.71 | 57.85 | 58.47 | 59.42 | 58.94 | 75.65 |