Planning on T-Eval official subset (test)
92.2PrecisionQwen2.5-72B-Instruct
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Qwen2.5-72B-Instruct2026.07 | 92.2 | 88.1 | 89.2 | |
| DeepSeek-V32026.07 | 91.1 | 87.4 | 88.5 | |
| GPT-4o2026.07 | 90.4 | 86.4 | 87.5 | |
| DeepSeek-R12026.07 | 90.2 | 87.2 | 87.8 | |
| o1-preview2026.07 | 90 | 86.5 | 87.4 | |
| Qwen3-8BTraining=CARL2026.07 | 89.5 | 88.5 | 88.1 | |
| Qwen3-8BTraining=RFT2026.07 | 88.6 | 88.2 | 87.3 | |
| DeepSeek-R1-Distill-Llama-8BTraining=CARL2026.07 | 88.4 | 88.6 | 87.5 | |
| QwQ-32B2026.07 | 88.3 | 84.6 | 85.6 | |
| DeepSeek-R1-Distill-Llama-8BTraining=RFT2026.07 | 88 | 86 | 86.4 | |
| Qwen3-8B2026.07 | 86.6 | 83.5 | 84.2 | |
| Llama-3.1-70B-Instruct2026.07 | 85.4 | 81.9 | 83 | |
| DeepSeek-R1-Distill-Llama-8B2026.07 | 81.8 | 79.3 | 79.4 | |
| Llama-3.1-8B-Instruct2026.07 | 81.5 | 76.5 | 78.9 |