Mathematical Reasoning on Olympiad (Pass@1 accuracy)
68.3Pass@1 AccuracyPieceHint-Nemotron-1.5B
Evaluation Results
| Method | Links | |
|---|---|---|
| PieceHint-Nemotron-1.5BModel Name=Nemotron-1.5B, Model Size=1.5B, Method Type=PieceHint2026.04 | 68.3 | |
| SFPOModel=DS-distilled-Qwen-7B2025.10 | 65.73 | |
| DeepSearch-1.5Bnumber of samples=32, model scale=1.5B, evaluation cluster=128×H100 96G2025.09 | 65.72 | |
| Nemotron-Research-Reasoning-Qwen-1.5B v2number of samples=32, model scale=1.5B, evaluation cluster=128×H100 96G2025.09 | 64.69 | |
| Nemotron-Research-Reasoning-Qwen-1.5B v1number of samples=32, model scale=1.5B, evaluation cluster=128×H100 96G2025.09 | 64.56 | |
| PieceHint-Qwen3-1.7BModel Name=Qwen3-1.7B, Model Size=1.7B, Method Type=PieceHint2026.04 | 62.4 | |
| GRPOModel=DS-distilled-Qwen-7B2025.10 | 61.24 | |
| Nemotron-1.5BModel Name=Nemotron-1.5B, Model Size=1.5B, Method Type=Baseline2026.04 | 61.2 | |
| DeepSeek-R1-Distill-32BModel Name=DeepSeek-R1-Distill-32B, Model Size=32B, Method Type=Baseline2026.04 | 61.2 | |
| DeepScaleR-1.5Bnumber of samples=32, model scale=1.5B, evaluation cluster=128×H100 96G2025.09 | 58.95 | |
| Qwen3-4BModel Name=Qwen3-4B, Model Size=4B, Method Type=Baseline2026.04 | 58.5 | |
| DeepSeek-R1-Distill-7BModel Name=DeepSeek-R1-Distill-7B, Model Size=7B, Method Type=Baseline2026.04 | 55.1 | |
| STILL-3-1.5Bnumber of samples=32, model scale=1.5B, evaluation cluster=128×H100 96G2025.09 | 53.84 | |
| Qwen3-1.7BModel Name=Qwen3-1.7B, Model Size=1.7B, Method Type=Baseline2026.04 | 53.7 | |
| Open-RS1-1.5Bnumber of samples=32, model scale=1.5B, evaluation cluster=128×H100 96G2025.09 | 52.82 | |
| Open-RS2-1.5Bnumber of samples=32, model scale=1.5B, evaluation cluster=128×H100 96G2025.09 | 52.63 | |
| Open-RS3-1.5Bnumber of samples=32, model scale=1.5B, evaluation cluster=128×H100 96G2025.09 | 52.25 | |
| Qwen2.5-32B + SimpleRLEvaluation Approach=LLM-as-a-judge (Ours), Model Size=32B, Training Strategy=SimpleRL2026.04 | 52 | |
| DeepSeek-R1-Distill-Qwen-1.5Bnumber of samples=32, model scale=1.5B, evaluation cluster=128×H100 96G2025.09 | 51.55 | |
| SFPOModel=DS-distilled-Qwen-1.5B2025.10 | 50.67 | |
| PieceHint-DeepSeek-1.5BModel Name=DeepSeek-R1-Distill-1.5B, Model Size=1.5B, Method Type=PieceHint2026.04 | 50.1 | |
| GRPOModel=DS-distilled-Qwen-1.5B2025.10 | 49.85 | |
| BaseModel=DS-distilled-Qwen-7B2025.10 | 48.89 | |
| SFPOModel=Qwen3-4B-Base2025.10 | 48.67 | |
| GRPOModel=Qwen3-4B-Base2025.10 | 48.59 | |
| Qwen2.5-14B + SimpleRLEvaluation Approach=LLM-as-a-judge (Ours), Model Size=14B, Training Strategy=SimpleRL2026.04 | 48.1 | |
| Qwen2.5-7B + SimpleRL (Lighteval)Evaluation Approach=LLM-as-a-judge (Ours), Model Size=7B, Training Strategy=SimpleRL, Evaluation Framework=Lighteval2026.04 | 46.1 | |
| Qwen2.5-32B + SimpleRLEvaluation Approach=Symbolic (Baseline), Model Size=32B, Training Strategy=SimpleRL2026.04 | 46 | |
| SFPOModel=Qwen2.5-Math-7B2025.10 | 45.59 | |
| GRPOModel=Qwen2.5-Math-7B2025.10 | 44.8 | |
| Qwen2.5-14B + SimpleRLEvaluation Approach=Symbolic (Baseline), Model Size=14B, Training Strategy=SimpleRL2026.04 | 43.5 | |
| DeepSeek-R1-Distill-1.5BModel Name=DeepSeek-R1-Distill-1.5B, Model Size=1.5B, Method Type=Baseline2026.04 | 43.4 | |
| Qwen2.5-7B + SimpleRLEvaluation Approach=LLM-as-a-judge (Ours), Model Size=7B, Training Strategy=SimpleRL2026.04 | 42.7 | |
| Qwen2.5-7B + SimpleRLEvaluation Approach=Symbolic (Baseline), Model Size=7B, Training Strategy=SimpleRL2026.04 | 41.2 | |
| DCRLBackbone=Qwen3-8B-Base2026.03 | 40.9 | |
| Qwen2.5-Math-1.5B-Instructnumber of samples=32, model scale=1.5B, evaluation cluster=128×H100 96G2025.09 | 40 | |
| SFPOModel=Qwen2.5-Math-1.5B2025.10 | 39.72 | |
| GRPOModel=Qwen2.5-Math-1.5B2025.10 | 39.42 | |
| BaseModel=Qwen2.5-Math-7B2025.10 | 38.54 | |
| Qwen2.5-Math-1.5B-Oat-Zeronumber of samples=32, model scale=1.5B, evaluation cluster=128×H100 96G2025.09 | 37.78 | |
| DCRLBackbone=Qwen3-4B-Base2026.03 | 37.2 | |
| MTTeacher Model=Qwen2.5-Math-7B-Instruct2026.04 | 37 | |
| BaseModel=DS-distilled-Qwen-1.5B2025.10 | 36.94 | |
| VCRDTeacher Model=Qwen2.5-Math-7B-Instruct, Student Model=Qwen2.5-Math-1.5B2026.04 | 35.4 | |
| distillm22026.04 | 34.1 | |
| DistillLM-2Teacher Model=Qwen2.5-Math-7B-Instruct, Student Model=Qwen2.5-Math-1.5B2026.04 | 34.1 | |
| Qwen2.5-7B (Lighteval)Evaluation Approach=LLM-as-a-judge (Ours), Model Size=7B, Evaluation Framework=Lighteval2026.04 | 33.9 | |
| VCRD-Probapproach=PRM-free2026.04 | 33.6 | |
| DistilLLMTeacher Model=Qwen2.5-Math-7B-Instruct, Student Model=Qwen2.5-Math-1.5B2026.04 | 33.5 | |
| GKDTeacher Model=Qwen2.5-Math-7B-Instruct, Student Model=Qwen2.5-Math-1.5B2026.04 | 33.3 | |
| ABKDTeacher Model=Qwen2.5-Math-7B-Instruct, Student Model=Qwen2.5-Math-1.5B2026.04 | 33.2 | |
| KDTeacher Model=Qwen2.5-Math-7B-Instruct, Student Model=Qwen2.5-Math-1.5B2026.04 | 32.4 | |
| Qwen2.5-14BEvaluation Approach=LLM-as-a-judge (Ours), Model Size=14B2026.04 | 31.4 | |
| Qwen2.5-Math-1.5Bnumber of samples=32, model scale=1.5B, evaluation cluster=128×H100 96G2025.09 | 30.74 | |
| Qwen2.5-7BEvaluation Approach=LLM-as-a-judge (Ours), Model Size=7B2026.04 | 30.6 | |
| Qwen2.5-32BEvaluation Approach=LLM-as-a-judge (Ours), Model Size=32B2026.04 | 30.6 | |
| Qwen3-1.7B-BaseTraining Strategy=Full-shot2026.04 | 29.48 | |
| Qwen2.5-14BEvaluation Approach=Symbolic (Baseline), Model Size=14B2026.04 | 28.6 | |
| BaseModel=Qwen2.5-Math-1.5B2025.10 | 28.45 | |
| Qwen2.5-7BEvaluation Approach=Symbolic (Baseline), Model Size=7B2026.04 | 28.1 | |
| Qwen2.5-32BEvaluation Approach=Symbolic (Baseline), Model Size=32B2026.04 | 27.9 | |
| Qwen2.5-7B (Lighteval)Evaluation Approach=Symbolic (Baseline), Model Size=7B, Evaluation Framework=Lighteval2026.04 | 27.8 | |
| Qwen2.5-7B + SimpleRL (Lighteval)Evaluation Approach=Symbolic (Baseline), Model Size=7B, Training Strategy=SimpleRL, Evaluation Framework=Lighteval2026.04 | 27.1 | |
| Qwen3-1.7B-BaseTraining Strategy=Only-General2026.04 | 26.07 | |
| HEALTraining Strategy=HEAL2026.04 | 25.78 | |
| Qwen3-1.7B-BaseTraining Strategy=Hybrid2026.04 | 24.15 | |
| Qwen3-1.7B-BaseTraining Strategy=Few-shot2026.04 | 23.7 | |
| Qwen3-1.7B-BaseTraining Strategy=Base2026.04 | 21.48 | |
| BaseModel=Qwen3-4B-Base2025.10 | 20.66 | |
| MSStudent Model=Qwen2.5-Math-1.5B2026.04 | 16.9 | |
| DCRLBackbone=Llama3.2-3B-Instruct2026.03 | 15.4 | |
| Llama3.1-8BEvaluation Approach=LLM-as-a-judge (Ours), Model Size=8B2026.04 | 3.5 | |
| Llama3.1-8BEvaluation Approach=Symbolic (Baseline), Model Size=8B2026.04 | 2.7 |