Competition Mathematics on MATH 5,000 problems (test)
97.6L1 AccuracyTRI (SFT + DPO)
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| TRI (SFT + DPO)Fine-tuning Protocol=SFT + DPO2026.06 | 97.6 | 95.2 | 88.9 | 73.4 | 52.7 | 1,267 | |
| TRI (Full)Fine-tuning Protocol=SFT + DPO + Repair2026.06 | 97.6 | 95.4 | 89.3 | 74.7 | 53.7 | 1,268 | |
| Qwen2.5-72B + CoT-SCBackbone=Qwen2.5-72B, Decoding Strategy=CoT-SC, Note=reproduced from official checkpoints2026.06 | 96.8 | 93.4 | 85.9 | 68.5 | 47.3 | 29,472 | |
| TRI (SFT only)Fine-tuning Protocol=SFT only2026.06 | 96.4 | 93.1 | 85.6 | 68.1 | 49.8 | 1,534 | |
| InternLM-StepProver + CoT-SCBackbone=InternLM-StepProver, Decoding Strategy=CoT-SC(k=8), Note=results from original papers2026.06 | 96.2 | 92.6 | 84.7 | 67.2 | 46.8 | 15,896 | |
| Qwen2.5-72B + CoTBackbone=Qwen2.5-72B, Decoding Strategy=CoT, Note=reproduced from official checkpoints2026.06 | 95.3 | 91.7 | 82.4 | 63.8 | 42.1 | 1,842 | |
| Llama-3.1-70B + ToTBackbone=Llama-3.1-70B, Decoding Strategy=ToT(b=5), Note=reproduced from official checkpoints2026.06 | 94.7 | 90.8 | 81.3 | 63.1 | 40.8 | 9,561 | |
| InternLM-StepProver + CoTBackbone=InternLM-StepProver, Decoding Strategy=CoT, Note=results from original papers2026.06 | 94.1 | 90.3 | 80.9 | 62.7 | 41.3 | 1,987 | |
| Llama-3.1-70B + CoTBackbone=Llama-3.1-70B, Decoding Strategy=CoT, Note=reproduced from official checkpoints2026.06 | 93.6 | 88.9 | 78.1 | 58.3 | 36.4 | 1,913 |