Tool-use on ACEBench
61.8AccuracyQwen2.5-Instruct-14B with TOUCAN-SFT + COVERT-RL
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen2.5-Instruct-14B with TOUCAN-SFT + COVERT-RLModel Scale=14B, Training Strategy=TOUCAN-SFT + COVERT-RL2026.04 | 61.8 | |
| Qwen2.5-Instruct-14B with COVERT-RLModel Scale=14B, Training Strategy=COVERT-RL2026.04 | 59.3 | |
| Qwen2.5-Instruct-7B with TOUCAN-SFT + COVERT-RLModel Scale=7B, Training Strategy=TOUCAN-SFT + COVERT-RL2026.04 | 53.9 | |
| Qwen2.5-Instruct-14BModel Scale=14B, Training Strategy=Base2026.04 | 53 | |
| Qwen2.5-Instruct-7B with COVERT-RLModel Scale=7B, Training Strategy=COVERT-RL2026.04 | 51.2 | |
| Qwen2.5-Instruct-14B with TOUCAN-SFTModel Scale=14B, Training Strategy=TOUCAN-SFT2026.04 | 48.8 | |
| Qwen2.5-Instruct-7BModel Scale=7B, Training Strategy=Base2026.04 | 42.6 | |
| Qwen2.5-Instruct-7B with TOUCAN-SFTModel Scale=7B, Training Strategy=TOUCAN-SFT2026.04 | 41.8 |