Tool Use on BFCL V3
62.1AccuracyQwen2.5-Instruct-14B with TOUCAN-SFT + COVERT-RL
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen2.5-Instruct-14B with TOUCAN-SFT + COVERT-RLModel Scale=14B, Training Strategy=TOUCAN-SFT + COVERT-RL2026.04 | 62.1 | |
| Qwen3 4BThinking Mode=true, FC Format=true, temperature=1, top-p=0.95, top-k=20, presence penalty=1.52025.12 | 61.7 | |
| Qwen2.5-Instruct-14B with COVERT-RLModel Scale=14B, Training Strategy=COVERT-RL2026.04 | 59.9 | |
| Qwen2.5-Instruct-14B with TOUCAN-SFTModel Scale=14B, Training Strategy=TOUCAN-SFT2026.04 | 59.5 | |
| Qwen2.5-Instruct-7B with TOUCAN-SFT + COVERT-RLModel Scale=7B, Training Strategy=TOUCAN-SFT + COVERT-RL2026.04 | 59.1 | |
| Youtu-LLM 2BThinking Mode=true, FC Format=true, temperature=1, top-p=0.95, top-k=20, presence penalty=1.52025.12 | 58 | |
| Qwen2.5-Instruct-7B with COVERT-RLModel Scale=7B, Training Strategy=COVERT-RL2026.04 | 57.2 | |
| Qwen2.5-Instruct-7B with TOUCAN-SFTModel Scale=7B, Training Strategy=TOUCAN-SFT2026.04 | 57 | |
| Qwen2.5-Instruct-14BModel Scale=14B, Training Strategy=Base2026.04 | 56.5 | |
| Qwen3 1.7BThinking Mode=true, FC Format=true, temperature=1, top-p=0.95, top-k=20, presence penalty=1.52025.12 | 55.5 | |
| Qwen2.5-Instruct-7BModel Scale=7B, Training Strategy=Base2026.04 | 54.1 | |
| SmolLM3 3BThinking Mode=true, FC Format=true, temperature=1, top-p=0.95, top-k=20, presence penalty=1.52025.12 | 31.5 |