Tool-use Reasoning on ShopTrajQA 32k context
62.6Accuracygpt-oss-120B
Evaluation Results
| Method | Links | |
|---|---|---|
| gpt-oss-120BInference Mode=pure text2026.06 | 62.6 | |
| Qwen3-4B-SFT-RLVRTraining Stage=SFT-RLVR2026.06 | 62.5 | |
| Qwen3-1.7B-SFT-RLVRTraining Stage=SFT-RLVR2026.06 | 59.2 | |
| Qwen3-4BInference Mode=pure text2026.06 | 48 | |
| Qwen3-Coder-30BInference Mode=pure text2026.06 | 38.4 | |
| Qwen3-4B-SFTTraining Stage=SFT2026.06 | 36.4 | |
| gpt-oss-120BInference Mode=tool-calling2026.06 | 28.7 | |
| Qwen3-1.7B-SFTTraining Stage=SFT2026.06 | 28.1 | |
| Qwen3-Coder-30BInference Mode=tool-calling2026.06 | 25 | |
| Qwen3-1.7BInference Mode=pure text2026.06 | 23 | |
| gpt-oss-20BInference Mode=pure text2026.06 | 22.3 | |
| Qwen3-1.7BInference Mode=tool-calling2026.06 | 17.3 | |
| gpt-oss-20BInference Mode=tool-calling2026.06 | 17.1 | |
| Qwen3-4BInference Mode=tool-calling2026.06 | 13.8 |