Mathematical Reasoning on OlympiadBench (Accuracy, Tokens)
82.44AccuracyKnowRL-Nemotron-1.5B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| KnowRL-Nemotron-1.5BHint Setting=CSS, Evaluation Protocol=mean@82026.04 | 82.44 | — | |
| KnowRL-Nemotron-1.5BHint Setting=CBRS, Evaluation Protocol=mean@82026.04 | 82.34 | — | |
| KnowRL-Nemotron-1.5BHint Setting=w/o KP, Evaluation Protocol=mean@82026.04 | 80.23 | — | |
| JustRLHint Setting=CSS, Evaluation Protocol=mean@82026.04 | 78.68 | — | |
| QuestAHint Setting=CSS, Evaluation Protocol=mean@82026.04 | 78.53 | — | |
| QuestAHint Setting=CBRS, Evaluation Protocol=mean@82026.04 | 78.45 | — | |
| JustRLHint Setting=CBRS, Evaluation Protocol=mean@82026.04 | 78.41 | — | |
| JustRLHint Setting=w/o KP, Evaluation Protocol=mean@82026.04 | 76.59 | — | |
| Nemotron-1.5BHint Setting=CSS, Evaluation Protocol=mean@82026.04 | 74.09 | — | |
| Nemotron-1.5BHint Setting=CBRS, Evaluation Protocol=mean@82026.04 | 73.89 | — | |
| QuestAHint Setting=w/o KP, Evaluation Protocol=mean@82026.04 | 72.28 | — | |
| Nemotron-1.5BHint Setting=w/o KP, Evaluation Protocol=mean@82026.04 | 71.7 | — | |
| LALPStudent=Qwen2.5-32B-Instruct, Model Configuration=LALP2025.10 | 67.3 | — | |
| Local LowestStudent=Qwen2.5-32B-Instruct, Model Configuration=Local Lowest2025.10 | 64 | — | |
| RandomStudent=Qwen2.5-32B-Instruct, Model Configuration=Random2025.10 | 63.6 | — | |
| GALPStudent=Qwen2.5-32B-Instruct, Model Configuration=GALP2025.10 | 63.6 | — | |
| OriginalBackbone=DeepSeek-R1-Distill-Qwen-7B2025.08 | 56.82 | 8,789 | |
| LAPO-I2026.03 | 56.3 | 4,024 | |
| LAPO-D2026.03 | 56.1 | 4,499 | |
| LCPOBackbone=DeepSeek-R1-Distill-Qwen-7B2025.08 | 56.08 | 4,222 | |
| LEADBackbone=DeepSeek-R1-Distill-Qwen-1.5B, Max Response Length=8K2026.05 | 55.85 | 5,015 | |
| ThinkPrune-4k2026.03 | 55.7 | 4,010 | |
| DASTBackbone=DeepSeek-R1-Distill-Qwen-7B2025.08 | 55.34 | 10,339 | |
| ThinkPrune-I2k2026.03 | 54.7 | 3,498 | |
| DeepScaler-1.5B2026.03 | 54.6 | 5,974 | |
| CoDBackbone=DeepSeek-R1-Distill-Qwen-7B2025.08 | 54.6 | 7,200 | |
| L1-MaxBackbone=DeepSeek-R1-Distill-Qwen-7B2025.08 | 54.3 | 2,510 | |
| L1-ExactBackbone=DeepSeek-R1-Distill-Qwen-7B2025.08 | 53.86 | 3,712 | |
| TrEffBackbone=DeepSeek-R1-Distill-Qwen-7B2025.08 | 53.56 | 6,970 | |
| GDPOBackbone=DeepSeek-R1-Distill-Qwen-1.5B, Max Response Length=8K2026.05 | 53.48 | 3,005 | |
| DRPOBackbone=DeepSeek-R1-Distill-Qwen-1.5B, Max Response Length=8K2026.05 | 53.23 | 4,255 | |
| AutoThink2026.03 | 52.5 | 4,085 | |
| BaseBackbone=DeepSeek-R1-Distill-Qwen-1.5B, Max Response Length=8K2026.05 | 51.75 | 8,370 | |
| HAPO2026.03 | 51.4 | 4,571 | |
| GRPOBackbone=DeepSeek-R1-Distill-Qwen-1.5B, Max Response Length=8K2026.05 | 51.16 | 2,813 | |
| Thinkless2026.03 | 50.2 | 6,057 | |
| EEPOBase Model=Qwen3-14B-Base2025.10 | 50.1 | — | |
| Qwen3-4B + General-ReasonerBase Model=Qwen3-4B, RL/Training Method=General-Reasoner, Cost (USD)=$4,6002026.05 | 49.1 | — | |
| GRPOBase Model=Qwen3-14B-Base2025.10 | 48.6 | — | |
| Original ModelStudent=Qwen2.5-32B-Instruct, Model Configuration=Original Model2025.10 | 47.1 | — | |
| RandomStudent=Qwen2.5-7B-Instruct, Model Configuration=Random2025.10 | 45.6 | — | |
| Qwen2.5-7B + Open-Reasoner-ZeroBase Model=Qwen2.5-7B, RL/Training Method=Open-Reasoner-Zero, Cost (USD)=$6,3002026.05 | 45.4 | — | |
| LCPOBackbone=DeepSeek-R1-Distill-Qwen-1.5B2025.08 | 44.81 | 4,921 | |
| ShorterBetterBackbone=DeepSeek-R1-Distill-Qwen-1.5B, Max Response Length=8K2026.05 | 44.69 | 2,271 | |
| OriginalBackbone=DeepSeek-R1-Distill-Qwen-1.5B2025.08 | 44.66 | 11,715 | |
| Re-Schedule_sigmoidBackbone=Qwen2.5-Math-7B, Evaluation Protocol=avg@32, Weighting Scheme=sigmoid2025.10 | 44.4 | — | |
| GALPStudent=Qwen2.5-7B-Instruct, Model Configuration=GALP2025.10 | 44.1 | — | |
| LALPStudent=Qwen2.5-7B-Instruct, Model Configuration=LALP2025.10 | 44.1 | — | |
| TrEffBackbone=DeepSeek-R1-Distill-Qwen-1.5B2025.08 | 43.62 | 7,652 | |
| Local LowestStudent=Qwen2.5-7B-Instruct, Model Configuration=Local Lowest2025.10 | 43.3 | — | |
| CoDBackbone=DeepSeek-R1-Distill-Qwen-1.5B2025.08 | 43.03 | 11,014 | |
| L1-MaxBackbone=DeepSeek-R1-Distill-Qwen-1.5B2025.08 | 43.03 | 3,452 | |
| L1-ExactBackbone=DeepSeek-R1-Distill-Qwen-1.5B2025.08 | 42.88 | 3,657 | |
| S-BoN (s)n=2562025.06 | 42.5 | — | |
| Re-Schedule_linearBackbone=Qwen2.5-Math-7B, Evaluation Protocol=avg@32, Weighting Scheme=linear2025.10 | 42.5 | — | |
| ACC_sigmoidBackbone=Qwen2.5-Math-7B, Evaluation Protocol=avg@32, Weighting Scheme=sigmoid2025.10 | 42.2 | — | |
| GSIn=642025.06 | 42 | — | |
| S-BoN (s)n=642025.06 | 42 | — | |
| S-BoN (b)n=642025.06 | 41.9 | — | |
| RSDn=642025.06 | 41.7 | — | |
| RSDn=2562025.06 | 41.7 | — | |
| RSDn=162025.06 | 41.5 | — | |
| GSIn=2562025.06 | 41.5 | — | |
| S-BoN (s)n=162025.06 | 41.3 | — | |
| S-BoN (b)n=2562025.06 | 41.2 | — | |
| SFTBase Model=Qwen3-14B-Base2026.05 | 41.2 | — | |
| S-BoN (b)n=162025.06 | 41.1 | — | |
| Qwen2.5-7B + ReasonMaxxerBase Model=Qwen2.5-7B, RL/Training Method=ReasonMaxxer, Cost (USD)=$52026.05 | 41.1 | — | |
| OPOBackbone=Qwen2.5-Math-7B, Evaluation Protocol=avg@322025.10 | 41 | — | |
| GRPOBackbone=Qwen2.5-Math-7B, Evaluation Protocol=avg@322025.10 | 40.9 | — | |
| SCRLBase Model=Qwen3-14B-Base2026.05 | 40.9 | — | |
| GSIn=162025.06 | 40.8 | — | |
| SimpleRL-ZooBackbone=Qwen2.5-Math-7B, Evaluation Protocol=avg@322025.10 | 40.8 | — | |
| Eurus-PRIMEBackbone=Qwen2.5-Math-7B, Evaluation Protocol=avg@322025.10 | 40.6 | — | |
| LPPOBackbone=Qwen2.5-Math-7B, Evaluation Protocol=avg@322025.10 | 40.6 | — | |
| Original ModelStudent=Qwen2.5-7B-Instruct, Model Configuration=Original Model2025.10 | 40.4 | — | |
| QuestABase Model=Qwen3-14B-Base2026.05 | 40.4 | — | |
| Qwen3-4B + ReasonMaxxerBase Model=Qwen3-4B, RL/Training Method=ReasonMaxxer, Cost (USD)=$42026.05 | 40.3 | — | |
| RSDn=42025.06 | 40.1 | — | |
| STRATAGEMModel=STRATAGEM (Ours)2026.04 | 39.9 | — | |
| GSIn=42025.06 | 39.7 | — | |
| S-BoN (s)n=42025.06 | 39.6 | — | |
| SCRLBase Model=Qwen3-4B-Base2026.05 | 39.2 | — | |
| Seed-GRPOBackbone=Qwen2.5-Math-7B, Evaluation Protocol=avg@322025.10 | 38.5 | — | |
| GRPOBase Model=Qwen3-14B-Base2026.05 | 38.4 | — | |
| NuRLBase Model=Qwen3-14B-Base2026.05 | 38.3 | — | |
| S-BoN (b)n=42025.06 | 38.1 | — | |
| KDRL w/ 7BBase Model=Qwen2.5-Math-1.5B, Teacher Model=7B2026.05 | 37.78 | — | |
| SFTTraining Protocol=SFT2025.08 | 37.5 | — | |
| DAPOBase Model=Qwen3-14B-Base2026.05 | 37.5 | — | |
| Qwen2.5-Math-7B + ReasonMaxxerBase Model=Qwen2.5-Math-7B, RL/Training Method=ReasonMaxxer, Cost (USD)=$52026.05 | 36.5 | — | |
| CoDistill-GRPOBase Model=Qwen2.5-Math-1.5B, alpha=1, T=02026.05 | 36.39 | — | |
| CoDistill-GRPOBase Model=Qwen2.5-Math-1.5B, alpha=1, T=502026.05 | 36.12 | — | |
| PSFTTraining Protocol=PSFT2025.08 | 36.02 | — | |
| Qwen2.5-7B + SimpleRL-ZooBase Model=Qwen2.5-7B, RL/Training Method=SimpleRL-Zoo, Cost (USD)=$6002026.05 | 35.8 | — | |
| QueSTBackbone=Qwen3-8B-Base2026.05 | 35.8 | — | |
| GSIn=12025.06 | 35.6 | — | |
| Qwen2.5-32B + ReasonMaxxerBase Model=Qwen2.5-32B, RL/Training Method=ReasonMaxxer, Cost (USD)=$252026.05 | 35.6 | — | |
| DeepSeek-R1-Distill-1.5B + ReasonMaxxerBase Model=DeepSeek-R1-Distill-1.5B, RL/Training Method=ReasonMaxxer, Cost (USD)=$42026.05 | 35.6 | — | |
| NuRLBase Model=Qwen3-4B-Base2026.05 | 35.6 | — |