Mathematical Reasoning on AMC 23 (Accuracy, Avg, Speedup)
95AccuracyGRPO
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GRPOBackbone=Qwen-3-8B-Base, Method=GRPO2026.03 | 95 | 56.66 | — | |
| GRPO + ARROLBackbone=Qwen-3-8B-Base, Method=GRPO + ARROL2026.03 | 95 | 59.53 | 1.62 | |
| GRPO + ARROLBackbone=Qwen-3-4B-Base, Method=GRPO + ARROL2026.03 | 92.5 | 53.54 | 1.63 | |
| Skywork-OR1-Math-7BDistillation Domain=Single-Domain, Teacher Model=Skywork-OR1-Math-7B2026.05 | 92.34 | — | — | |
| Skywork-OR1-7BDistillation Domain=Multi-Domain, Teacher Model=Skywork-OR1-7B2026.05 | 91.79 | — | — | |
| GRPOBackbone=Qwen-3-4B-Base, Method=GRPO2026.03 | 87.5 | 51.01 | — | |
| GRPO + ARROLBackbone=Qwen-3-1.7B-Base, Method=GRPO + ARROL2026.03 | 82.5 | 37.09 | 1.61 | |
| TrOPDDistillation Domain=Single-Domain, Teacher Model=Skywork-OR1-Math-7B, Student Model=DeepSeek-R1-Distill-Qwen-1.5B, Training Steps=200, Learning Rate=5 x 10^-6, Sample Rollouts=4, Max Generation Length=80962026.05 | 77.03 | — | — | |
| REOPOLDDistillation Domain=Single-Domain, Teacher Model=Skywork-OR1-Math-7B, Student Model=DeepSeek-R1-Distill-Qwen-1.5B, Training Steps=200, Learning Rate=5 x 10^-6, Sample Rollouts=4, Max Generation Length=80962026.05 | 75.78 | — | — | |
| OPDDistillation Domain=Single-Domain, Teacher Model=Skywork-OR1-Math-7B, Student Model=DeepSeek-R1-Distill-Qwen-1.5B, Training Steps=200, Learning Rate=5 x 10^-6, Sample Rollouts=4, Max Generation Length=80962026.05 | 75.39 | — | — | |
| EOPDDistillation Domain=Single-Domain, Teacher Model=Skywork-OR1-Math-7B, Student Model=DeepSeek-R1-Distill-Qwen-1.5B, Training Steps=200, Learning Rate=5 x 10^-6, Sample Rollouts=4, Max Generation Length=80962026.05 | 75.23 | — | — | |
| GRPOBackbone=Qwen-3-1.7B-Base, Method=GRPO2026.03 | 75 | 34.79 | — | |
| Entropy OPD 20%Distillation Domain=Single-Domain, Teacher Model=Skywork-OR1-Math-7B, Student Model=DeepSeek-R1-Distill-Qwen-1.5B, Training Steps=200, Learning Rate=5 x 10^-6, Sample Rollouts=4, Max Generation Length=80962026.05 | 73.82 | — | — | |
| REOPOLD 2StageDistillation Domain=Single-Domain, Teacher Model=Skywork-OR1-Math-7B, Student Model=DeepSeek-R1-Distill-Qwen-1.5B, Training Steps=200, Learning Rate=5 x 10^-6, Sample Rollouts=4, Max Generation Length=80962026.05 | 73.35 | — | — | |
| DeepSeek-Qwen2.5-1.5BStudent Model=DeepSeek-Qwen2.5-1.5B2026.05 | 71.01 | — | — | |
| TrOPDDistillation Domain=Multi-Domain, Teacher Model=Skywork-OR1-7B, Student Model=DeepSeek-R1-Distill-Qwen-1.5B, Training Steps=200, Learning Rate=5 x 10^-6, Sample Rollouts=4, Max Generation Length=80962026.05 | 70.93 | — | — | |
| few-shotModel=OPEN-REASONER-7B2026.05 | 65 | — | — | |
| REOPOLDDistillation Domain=Multi-Domain, Teacher Model=Skywork-OR1-7B, Student Model=DeepSeek-R1-Distill-Qwen-1.5B, Training Steps=200, Learning Rate=5 x 10^-6, Sample Rollouts=4, Max Generation Length=80962026.05 | 63.9 | — | — | |
| OPDDistillation Domain=Multi-Domain, Teacher Model=Skywork-OR1-7B, Student Model=DeepSeek-R1-Distill-Qwen-1.5B, Training Steps=200, Learning Rate=5 x 10^-6, Sample Rollouts=4, Max Generation Length=80962026.05 | 61.56 | — | — | |
| LRSModel=OPEN-REASONER-7B, Prompting=0-shot2026.05 | 60 | — | — | |
| CoTModel=OPEN-REASONER-7B2026.05 | 55 | — | — | |
| BaseModel=OPEN-REASONER-7B, Prompting=0-shot2026.05 | 50 | — | — | |
| GRPO + ARROLBackbone=Llama-3.2-1B-Instruct, Method=GRPO + ARROL2026.03 | 47.5 | 17.49 | 1.67 | |
| GRPOBackbone=Llama-3.2-1B-Instruct, Method=GRPO2026.03 | 45 | 14.63 | — | |
| LRS BASICModel=OPEN-REASONER-7B, Prompting=0-shot2026.05 | 45 | — | — | |
| LRSModel=OPEN-REASONER-1.5B, Prompting=0-shot2026.05 | 37.5 | — | — | |
| CoTModel=OPEN-REASONER-1.5B2026.05 | 32.5 | — | — | |
| LRS BASICModel=OPEN-REASONER-1.5B, Prompting=0-shot2026.05 | 32.5 | — | — | |
| Coupled-GRPOGeneration steps=30.0, Unmasking strategy=Static, Base Model Family=Dream2026.05 | 30 | — | — | |
| RLDFGeneration steps=256, Unmasking strategy=Static, Base Model Family=Dream2026.05 | 30 | — | — | |
| RLDFGeneration steps=512, Unmasking strategy=Static, Base Model Family=Dream2026.05 | 30 | — | — | |
| BaseModel=OPEN-REASONER-1.5B, Prompting=0-shot2026.05 | 30 | — | — | |
| few-shotModel=OPEN-REASONER-1.5B2026.05 | 30 | — | — | |
| Coupled-GRPOGeneration steps=512, Unmasking strategy=Static, Base Model Family=LLaDA2026.05 | 22.5 | — | — | |
| BaseBackbone=Llama3.1-8B-Instruct2026.06 | 22.5 | — | — | |
| RPOBackbone=Llama3.1-8B-Instruct2026.06 | 22.5 | — | — | |
| IPOBackbone=Llama3.1-8B-Instruct2026.06 | 22.5 | — | — | |
| TDPOBackbone=Llama3.1-8B-Instruct2026.06 | 22.5 | — | — | |
| RePO_detBackbone=Llama3.1-8B-Instruct2026.06 | 22.5 | — | — | |
| ESPOGeneration steps=512, Unmasking strategy=Static, Base Model Family=LLaDA2026.05 | 20 | — | — | |
| TraceRLGeneration steps=512, Unmasking strategy=Static, Base Model Family=LLaDA2026.05 | 20 | — | — | |
| d1Generation steps=128, Unmasking strategy=Static, Base Model Family=Dream2026.05 | 20 | — | — | |
| KTOBackbone=Llama3.1-8B-Instruct2026.06 | 20 | — | — | |
| d1Generation steps=256, Unmasking strategy=Static, Base Model Family=LLaDA2026.05 | 17.5 | — | — | |
| ESPOGeneration steps=256, Unmasking strategy=Static, Base Model Family=LLaDA2026.05 | 17.5 | — | — | |
| RLDFGeneration steps=128, Unmasking strategy=Static, Base Model Family=LLaDA2026.05 | 17.5 | — | — | |
| RLDFGeneration steps=512, Unmasking strategy=Static, Base Model Family=LLaDA2026.05 | 17.5 | — | — | |
| Dream-7B-InstructGeneration steps=256, Unmasking strategy=Static, Base Model Family=Dream2026.05 | 17.5 | — | — | |
| Dream-7B-InstructGeneration steps=512, Unmasking strategy=Static, Base Model Family=Dream2026.05 | 17.5 | — | — | |
| ESPOGeneration steps=512, Unmasking strategy=Static, Base Model Family=Dream2026.05 | 17.5 | — | — | |
| TraceRLGeneration steps=512, Unmasking strategy=Static, Base Model Family=Dream2026.05 | 17.5 | — | — | |
| Coupled-GRPOGeneration steps=256, Unmasking strategy=Static, Base Model Family=Dream2026.05 | 17.5 | — | — | |
| RLDFGeneration steps=128, Unmasking strategy=Static, Base Model Family=Dream2026.05 | 17.5 | — | — | |
| DPOBackbone=Llama3.1-8B-Instruct2026.06 | 17.5 | — | — | |
| LLaDA-8B-InstructGeneration steps=256, Unmasking strategy=Static, Base Model Family=LLaDA2026.05 | 15 | — | — | |
| LLaDA-1.5Generation steps=128, Unmasking strategy=Static, Base Model Family=LLaDA2026.05 | 15 | — | — | |
| d1Generation steps=512, Unmasking strategy=Static, Base Model Family=LLaDA2026.05 | 15 | — | — | |
| Coupled-GRPOGeneration steps=128, Unmasking strategy=Static, Base Model Family=LLaDA2026.05 | 15 | — | — | |
| RLDFGeneration steps=256, Unmasking strategy=Static, Base Model Family=LLaDA2026.05 | 15 | — | — | |
| Dream-7B-InstructGeneration steps=128, Unmasking strategy=Static, Base Model Family=Dream2026.05 | 15 | — | — | |
| d1Generation steps=256, Unmasking strategy=Static, Base Model Family=Dream2026.05 | 15 | — | — | |
| d1Generation steps=512, Unmasking strategy=Static, Base Model Family=Dream2026.05 | 15 | — | — | |
| ESPOGeneration steps=128, Unmasking strategy=Static, Base Model Family=Dream2026.05 | 15 | — | — | |
| ESPOGeneration steps=256, Unmasking strategy=Static, Base Model Family=Dream2026.05 | 15 | — | — | |
| TraceRLGeneration steps=256, Unmasking strategy=Static, Base Model Family=Dream2026.05 | 15 | — | — | |
| LLaDA-8B-InstructGeneration steps=512, Unmasking strategy=Static, Base Model Family=LLaDA2026.05 | 12.5 | — | — | |
| LLaDA-1.5Generation steps=256, Unmasking strategy=Static, Base Model Family=LLaDA2026.05 | 12.5 | — | — | |
| d1Generation steps=128, Unmasking strategy=Static, Base Model Family=LLaDA2026.05 | 12.5 | — | — | |
| Coupled-GRPOGeneration steps=128, Unmasking strategy=Static, Base Model Family=Dream2026.05 | 12.5 | — | — | |
| KTOBackbone=Llama3.1-8B-Base2026.06 | 12.5 | — | — | |
| LLaDA-1.5Generation steps=512, Unmasking strategy=Static, Base Model Family=LLaDA2026.05 | 10 | — | — | |
| ESPOGeneration steps=128, Unmasking strategy=Static, Base Model Family=LLaDA2026.05 | 10 | — | — | |
| Coupled-GRPOGeneration steps=256, Unmasking strategy=Static, Base Model Family=LLaDA2026.05 | 10 | — | — | |
| TraceRLGeneration steps=128, Unmasking strategy=Static, Base Model Family=Dream2026.05 | 10 | — | — | |
| DPOBackbone=Llama3.1-8B-Base2026.06 | 10 | — | — | |
| RePO_detBackbone=Llama3.1-8B-Base2026.06 | 10 | — | — | |
| LLaDA-8B-InstructGeneration steps=128, Unmasking strategy=Static, Base Model Family=LLaDA2026.05 | 7.5 | — | — | |
| TraceRLGeneration steps=128, Unmasking strategy=Static, Base Model Family=LLaDA2026.05 | 7.5 | — | — | |
| RPOBackbone=Llama3.1-8B-Base2026.06 | 7.5 | — | — | |
| TraceRLGeneration steps=256, Unmasking strategy=Static, Base Model Family=LLaDA2026.05 | 5 | — | — | |
| IPOBackbone=Llama3.1-8B-Base2026.06 | 5 | — | — | |
| TDPOBackbone=Llama3.1-8B-Base2026.06 | 5 | — | — | |
| BaseBackbone=Llama3.1-8B-Base2026.06 | 2.5 | — | — |