Tool-use Reasoning on ShopTrajQA 64k context
60.1AccuracyQwen3-4B-SFT-RLVR
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3-4B-SFT-RLVRTraining Stage=SFT-RLVR2026.06 | 60.1 | |
| gpt-oss-120BInference Mode=pure text2026.06 | 57.2 | |
| Qwen3-1.7B-SFT-RLVRTraining Stage=SFT-RLVR2026.06 | 51 | |
| Qwen3-4BInference Mode=pure text2026.06 | 47.6 | |
| Qwen3-4B-SFTTraining Stage=SFT2026.06 | 41 | |
| gpt-oss-120BInference Mode=tool-calling2026.06 | 35.5 | |
| Qwen3-1.7B-SFTTraining Stage=SFT2026.06 | 27.7 | |
| Qwen3-Coder-30BInference Mode=pure text2026.06 | 24.7 | |
| Qwen3-Coder-30BInference Mode=tool-calling2026.06 | 22.4 | |
| gpt-oss-20BInference Mode=pure text2026.06 | 21.4 | |
| gpt-oss-20BInference Mode=tool-calling2026.06 | 18.8 | |
| Qwen3-1.7BInference Mode=tool-calling2026.06 | 14.7 | |
| Qwen3-4BInference Mode=tool-calling2026.06 | 9.3 |