Mathematical Reasoning on (AIME24, AIME25, AMC23, HMMT24, HMMT25)
76.25AIME24 ScorePG-OPD
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| PG-OPDStudent Model=OpenMath-1.5B, Teacher Model=JustRL-Nemotron-1.5B, Pruned=50%, Training Time (s)=19757, Speedup=1.61×2026.06 | 76.25 | 67.08 | 95.78 | 36.88 | 45 | 64.2 | |
| PG-OPDStudent Model=OpenMath-1.5B, Teacher Model=JustRL-Nemotron-1.5B, Pruned=25%, Training Time (s)=21905, Speedup=1.45×2026.06 | 70.83 | 61.04 | 97.34 | 33.54 | 37.92 | 60.14 | |
| PG-OPDStudent Model=OpenMath-1.5B, Teacher Model=JustRL-Nemotron-1.5B, Pruned=75%, Training Time (s)=17008, Speedup=1.87×2026.06 | 69.17 | 57.08 | 96.72 | 39.17 | 40.21 | 60.47 | |
| OPDStudent Model=OpenMath-1.5B, Teacher Model=JustRL-Nemotron-1.5B, Pruned=0%, Training Time (s)=31815, Speedup=1.00×2026.06 | 68.96 | 60.21 | 96.56 | 30.83 | 40.42 | 59.4 | |
| PRUNE-OPDStudent Model=OpenMath-1.5B, Teacher Model=JustRL-Nemotron-1.5B, Pruned=–, Training Time (s)=19933, Speedup=1.60×2026.06 | 66.88 | 62.5 | 95 | 34.38 | 40.16 | 59.78 | |
| PG-OPDStudent Model=DeepSeek-R1-Distill-Qwen-1.5B, Teacher Model=JustRL-DeepSeek-1.5B, Pruned=75%, Training Time (s)=17265, Speedup=2.21×2026.06 | 52.08 | 33.96 | 82.81 | 21.67 | 23.54 | 42.81 | |
| OPDStudent Model=DeepSeek-R1-Distill-Qwen-1.5B, Teacher Model=JustRL-DeepSeek-1.5B, Pruned=0%, Training Time (s)=38221, Speedup=1.00×2026.06 | 47.92 | 35.42 | 85.47 | 18.75 | 22.29 | 41.97 | |
| PG-OPDStudent Model=DeepSeek-R1-Distill-Qwen-1.5B, Teacher Model=JustRL-DeepSeek-1.5B, Pruned=25%, Training Time (s)=21883, Speedup=1.75×2026.06 | 47.92 | 34.38 | 82.03 | 21.25 | 19.58 | 41.03 | |
| PRUNE-OPDStudent Model=DeepSeek-R1-Distill-Qwen-1.5B, Teacher Model=JustRL-DeepSeek-1.5B, Pruned=–, Training Time (s)=15838, Speedup=2.41×2026.06 | 47.08 | 31.46 | 82.65 | 20.21 | 18.33 | 39.95 | |
| PG-OPDStudent Model=DeepSeek-R1-Distill-Qwen-1.5B, Teacher Model=JustRL-DeepSeek-1.5B, Pruned=50%, Training Time (s)=20926, Speedup=1.83×2026.06 | 42.71 | 36.25 | 85.47 | 23.75 | 17.92 | 41.22 | |
| PG-OPDStudent Model=DeepSeek-R1-Distill-Qwen-1.5B, Teacher Model=DeepSeek-R1-Distill-Qwen-7B, Pruned=50%, Training Time (s)=20516, Speedup=1.90×2026.06 | 38.96 | 26.25 | 70 | 12.08 | 17.71 | 33 | |
| PG-OPDStudent Model=DeepSeek-R1-Distill-Qwen-1.5B, Teacher Model=DeepSeek-R1-Distill-Qwen-7B, Pruned=25%, Training Time (s)=23181, Speedup=1.68×2026.06 | 33.75 | 27.08 | 74.38 | 13.75 | 22.29 | 34.25 | |
| OPDStudent Model=DeepSeek-R1-Distill-Qwen-1.5B, Teacher Model=DeepSeek-R1-Distill-Qwen-7B, Pruned=0%, Training Time (s)=38896, Speedup=1.00×2026.06 | 30.63 | 24.58 | 70.94 | 10.42 | 16.67 | 30.65 | |
| PRUNE-OPDStudent Model=DeepSeek-R1-Distill-Qwen-1.5B, Teacher Model=DeepSeek-R1-Distill-Qwen-7B, Pruned=–, Training Time (s)=12868, Speedup=3.02×2026.06 | 28.96 | 23.75 | 68.91 | 12.71 | 11.29 | 29.12 | |
| PG-OPDStudent Model=DeepSeek-R1-Distill-Qwen-1.5B, Teacher Model=DeepSeek-R1-Distill-Qwen-7B, Pruned=75%, Training Time (s)=18294, Speedup=2.13×2026.06 | 26.46 | 24.38 | 73.12 | 11.04 | 18.12 | 30.63 | |
| PG-OPDStudent / Teacher model combination=Qwen3-4B-Base / Qwen3-8B, Pruned=50%, Training Time (s)=31017, Speedup=2.14×2026.06 | 18.33 | 12.92 | 51.25 | 4.38 | 4.79 | 18.33 | |
| PG-OPDStudent Model=Qwen3-4B-Base, Teacher Model=Qwen3-4B, Pruned=25%, Training Time (s)=36246, Speedup=1.72×2026.06 | 17.08 | 16.88 | 57.32 | 11.25 | 7.5 | 22.01 | |
| OPDStudent / Teacher model combination=Qwen3-4B-Base / Qwen3-8B, Pruned=0%, Training Time (s)=66486, Speedup=1.00×2026.06 | 16.67 | 14.58 | 53.91 | 3.75 | 1.25 | 18.03 | |
| PG-OPDStudent / Teacher model combination=Qwen3-4B-Base / Qwen3-8B, Pruned=25%, Training Time (s)=36922, Speedup=1.80×2026.06 | 16.67 | 11.25 | 53.91 | 8.33 | 2.5 | 18.53 | |
| OPDStudent Model=Qwen3-4B-Base, Teacher Model=Qwen3-4B, Pruned=0%, Training Time (s)=62476, Speedup=1.00×2026.06 | 16.46 | 16.04 | 48.91 | 5.83 | 4.17 | 18.28 | |
| PG-OPDStudent Model=Qwen3-4B-Base, Teacher Model=Qwen3-4B, Pruned=75%, Training Time (s)=25351, Speedup=2.46×2026.06 | 16.25 | 12.08 | 54.84 | 5 | 8.33 | 19.3 | |
| PRUNE-OPDStudent Model=Qwen3-4B-Base, Teacher Model=Qwen3-4B, Pruned=–, Training Time (s)=15143, Speedup=4.13×2026.06 | 14.58 | 13.54 | 46.56 | 10 | 3.96 | 17.73 | |
| PG-OPDStudent / Teacher model combination=Qwen3-4B-Base / Qwen3-8B, Pruned=75%, Training Time (s)=25906, Speedup=2.57×2026.06 | 14.17 | 18.1 | 52.97 | 7.71 | 1.67 | 18.92 | |
| PG-OPDStudent Model=Qwen3-4B-Base, Teacher Model=Qwen3-4B, Pruned=50%, Training Time (s)=30609, Speedup=2.04×2026.06 | 13.12 | 15 | 53.28 | 7.08 | 7.71 | 19.24 | |
| PG-OPDStudent / Teacher model combination=Qwen3-1.7B-Base / Qwen3-4B, Pruned=50%, Training Time (s)=32282, Speedup=1.44×2026.06 | 13.12 | 7.29 | 43.12 | 4.79 | 0 | 13.67 | |
| PG-OPDStudent / Teacher model combination=Qwen3-1.7B-Base / Qwen3-4B, Pruned=25%, Training Time (s)=38100, Speedup=1.22×2026.06 | 11.46 | 6.25 | 40.62 | 4.17 | 0 | 12.5 | |
| OPDStudent / Teacher model combination=Qwen3-1.7B-Base / Qwen3-4B, Pruned=0%, Training Time (s)=46441, Speedup=1.00×2026.06 | 8.96 | 3.96 | 37.81 | 3.54 | 0 | 10.85 | |
| PG-OPDStudent / Teacher model combination=Qwen3-1.7B-Base / Qwen3-4B, Pruned=75%, Training Time (s)=20341, Speedup=2.28×2026.06 | 7.29 | 4.38 | 34.69 | 4.17 | 0 | 10.1 |