Mathematical Reasoning on Olympiad
70.9AccuracyDeepSeek-R1-Distill-Qwen-7B + ERC-DAPO
Evaluation Results
| Method | Links | |
|---|---|---|
| DeepSeek-R1-Distill-Qwen-7B + ERC-DAPOModel Backbone=Qwen-7B, Training Algorithm=ERC-DAPO2025.12 | 70.9 | |
| DeepSeek-R1-Distill-Qwen-7B + DAPOModel Backbone=Qwen-7B, Training Algorithm=DAPO2025.12 | 69.9 | |
| DeepSeek-R1-Distill-Qwen-7BModel Backbone=Qwen-7B, Training Algorithm=Baseline2025.12 | 67 | |
| DeepSeek-R1-Distill-Qwen-7B + GRPOModel Backbone=Qwen-7B, Training Algorithm=GRPO2025.12 | 65.6 | |
| DeepSeek-R1-Distill-Qwen-1.5B + ERC-DAPOModel Backbone=Qwen-1.5B, Training Algorithm=ERC-DAPO2025.12 | 61 | |
| SOUPratioBackbone=DeepSeek-R1-Distill-Qwen-1.5B, Truncation=Length-ratio-based2026.01 | 59.05 | |
| DeepSeek-R1-Distill-Qwen-1.5B + DAPOModel Backbone=Qwen-1.5B, Training Algorithm=DAPO2025.12 | 58.6 | |
| VanillaModel=Qwen3-8B, Budget=24002026.03 | 58.22 | |
| VanillaModel=Qwen3-8B, Budget=32002026.03 | 58.22 | |
| On-policyBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.01 | 58 | |
| SPIRALBackbone=DeepSeek-Distill-Qwen-7B, Training Data=Multi-game2025.06 | 57.9 | |
| GSPO + LIEBackbone=Qwen3-4B-Base, Strategy=Length-Incentivized Exploration2026.02 | 57.2 | |
| M2POBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.01 | 56.92 | |
| DeepSeek-Distill-Qwen-7BBackbone=DeepSeek-Distill-Qwen-7B2025.06 | 56.9 | |
| DeepSeek-R1-Distill-Qwen-1.5B + GRPOModel Backbone=Qwen-1.5B, Training Algorithm=GRPO2025.12 | 56.2 | |
| R-KVModel=Qwen3-8B, Budget=32002026.03 | 55.56 | |
| LongFlowModel=Qwen3-8B, Budget=32002026.03 | 55.56 | |
| VATPModel=Qwen3-8B, Budget=32002026.03 | 55.41 | |
| GRPO w/Clip-higherBackbone=Qwen3-4B-Base, Variant=Clip-higher2026.02 | 54.1 | |
| GRPO w/Clip-higher + LIEBackbone=Qwen3-4B-Base, Variant=Clip-higher, Strategy=Length-Incentivized Exploration2026.02 | 54.1 | |
| H2OModel=Qwen3-8B, Budget=32002026.03 | 53.78 | |
| R-KVModel=Qwen3-8B, Budget=24002026.03 | 53.33 | |
| SFTBackbone=DeepSeek-Distill-Qwen-7B, Training Data=Multi-game2025.06 | 52.4 | |
| SOUPentropyBackbone=DeepSeek-R1-Distill-Qwen-1.5B, Truncation=Entropy-based2026.01 | 52.2 | |
| DeepSeek-R1-Distill-Qwen-1.5BModel Backbone=Qwen-1.5B, Training Algorithm=Baseline2025.12 | 51.8 | |
| GSPOBackbone=Qwen3-4B-Base2026.02 | 51.7 | |
| LUFFYBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.01 | 51.61 | |
| LongFlowModel=Qwen3-8B, Budget=24002026.03 | 51.11 | |
| DPO-R1 (HIGH)Backbone=Qwen3-4B-Instruct, Compute Cost (Sampling Budget N)=N = 16, Trainable Pairs=8,964, Processing Time=5.22026.02 | 50.8 | |
| PACEBackbone=Qwen3-4B-Instruct, Compute Cost (Sampling Budget N)=2 < N < 3, Trainable Pairs=10,717, Processing Time=1.02026.02 | 50.8 | |
| DPO-R1 (LOW)Backbone=Qwen3-4B-Instruct, Compute Cost (Sampling Budget N)=N = 2, Trainable Pairs=3,479, Processing Time=0.92026.02 | 50.7 | |
| DPO-R1 (ZHANG ET AL., 2025)Backbone=Qwen3-8B-Instruct, Compute Cost (Sampling Budget N)=N = 8, Trainable Pairs=10,211, Processing Time=4.82026.02 | 50.7 | |
| DPO-R1 (MIDDLE)Backbone=Qwen3-8B-Instruct, Compute Cost (Sampling Budget N)=N = 4, Trainable Pairs=6,073, Processing Time=2.22026.02 | 50.5 | |
| DPO-R1 (HIGH)Backbone=Qwen3-8B-Instruct, Compute Cost (Sampling Budget N)=N = 16, Trainable Pairs=10,543, Processing Time=7.22026.02 | 50.5 | |
| DPO-R1 (MIDDLE)Backbone=Qwen3-4B-Instruct, Compute Cost (Sampling Budget N)=N = 4, Trainable Pairs=5,127, Processing Time=1.22026.02 | 50.4 | |
| LUFFYBackbone=Qwen2.5-Math-7B2026.01 | 49.99 | |
| GRPO + LIEBackbone=Qwen3-4B-Base, Strategy=Length-Incentivized Exploration2026.02 | 49.9 | |
| INSTRUCT (BASE)Backbone=Qwen3-8B-Instruct, Compute Cost (Sampling Budget N)=N/A, Trainable Pairs=N/A, Processing Time=N/A2026.02 | 49.9 | |
| DPO-R1 (LOW)Backbone=Qwen3-8B-Instruct, Compute Cost (Sampling Budget N)=N = 2, Trainable Pairs=3,403, Processing Time=1.42026.02 | 49.9 | |
| PACEBackbone=Qwen3-8B-Instruct, Compute Cost (Sampling Budget N)=2 < N < 3, Trainable Pairs=9,028, Processing Time=1.72026.02 | 49.6 | |
| SPIRALBackbone=Qwen3-8B, Training Data=Multi-game2025.06 | 49.6 | |
| VATPModel=Qwen3-8B, Budget=24002026.03 | 49.33 | |
| H2OModel=Qwen3-8B, Budget=24002026.03 | 49.04 | |
| DPO-R1 (ZHANG ET AL., 2025)Backbone=Qwen3-4B-Instruct, Compute Cost (Sampling Budget N)=N = 8, Trainable Pairs=8,403, Processing Time=2.82026.02 | 48.9 | |
| SOUPratioBackbone=Qwen2.5-Math-7B, Truncation=Length-ratio-based2026.01 | 48.32 | |
| EASDTM=32B, DM=7B2025.12 | 48.15 | |
| Absolute ZeroBackbone=Qwen3-8B-Base, Methodology=Absolute Zero, Supervision Ratio=None2025.12 | 47.8 | |
| General-ReasonerBackbone=Qwen3-4B-Base, Methodology=General-Reasoner, Supervision Ratio=232k WebInstruct2025.12 | 47.7 | |
| INSTRUCT (BASE)Backbone=Qwen3-4B-Instruct, Compute Cost (Sampling Budget N)=N/A, Trainable Pairs=N/A, Processing Time=N/A2026.02 | 47.4 | |
| GRPOBackbone=Qwen3-4B-Base2026.02 | 47.1 | |
| Single ModelTM=32B2025.12 | 46.81 | |
| R-KVModel=DeepSeek-R1-Distill-Llama-8B, Budget=32002026.03 | 46.81 | |
| VATPModel=DeepSeek-R1-Distill-Llama-8B, Budget=32002026.03 | 46.67 | |
| DPO (Full)Base Model=Qwen2.5-7B-Instruct2026.02 | 46.5 | |
| RSDTM=32B, DM=7B, PRM=1.5B2025.12 | 46.5 | |
| R-Few (5%)Backbone=Qwen3-8B-Base, Methodology=R-Few, Supervision Ratio=5%2025.12 | 46.4 | |
| General-ReasonerBackbone=Qwen3-8B-Base, Methodology=General-Reasoner, Supervision Ratio=232k WebInstruct2025.12 | 46.3 | |
| H2OModel=DeepSeek-R1-Distill-Llama-8B, Budget=32002026.03 | 46.22 | |
| SOUPentropyBackbone=Qwen2.5-Math-7B, Truncation=Entropy-based2026.01 | 46.14 | |
| On-policyBackbone=Qwen2.5-Math-7B2026.01 | 45.97 | |
| R-KVModel=DeepSeek-R1-Distill-Llama-8B, Budget=24002026.03 | 45.93 | |
| LongFlowModel=DeepSeek-R1-Distill-Llama-8B, Budget=32002026.03 | 45.93 | |
| SFTBackbone=Qwen3-8B, Training Data=Multi-game2025.06 | 45.9 | |
| RSDTM=72B, DM=7B, PRM=1.5B2025.12 | 45.62 | |
| SAGEBase Model=Qwen2.5-7B-Instruct2026.02 | 45.5 | |
| Majority Voting (N=16)DM=7B2025.12 | 45.33 | |
| SDTM=32B, DM=7B2025.12 | 45.33 | |
| VanillaBase Model=Qwen2.5-7B-Instruct2026.02 | 45.3 | |
| EASDTM=72B, DM=7B2025.12 | 44.74 | |
| Single ModelTM=72B2025.12 | 44.59 | |
| VanillaModel=DeepSeek-R1-Distill-Llama-8B, Budget=24002026.03 | 44.3 | |
| VanillaModel=DeepSeek-R1-Distill-Llama-8B, Budget=32002026.03 | 44.3 | |
| SDTM=72B, DM=7B2025.12 | 44.15 | |
| R-Few (1%)Backbone=Qwen3-8B-Base, Methodology=R-Few, Supervision Ratio=1%2025.12 | 44 | |
| H2OModel=DeepSeek-R1-Distill-Llama-8B, Budget=24002026.03 | 44 | |
| LongFlowModel=DeepSeek-R1-Distill-Llama-8B, Budget=24002026.03 | 43.85 | |
| M2POBackbone=Qwen2.5-Math-7B2026.01 | 43.76 | |
| VATPModel=DeepSeek-R1-Distill-Llama-8B, Budget=24002026.03 | 43.56 | |
| R-ZeroBackbone=Qwen3-8B-Base, Methodology=R-Zero, Supervision Ratio=None2025.12 | 43.4 | |
| DPO (Random)Base Model=Qwen2.5-7B-Instruct2026.02 | 43 | |
| R-Few (5%)Backbone=Qwen3-4B-Base, Methodology=R-Few, Supervision Ratio=5%2025.12 | 42.8 | |
| SPICEBackbone=Qwen3-4B-Base, Methodology=SPICE, Supervision Ratio=None2025.12 | 42.7 | |
| SPICEBackbone=Qwen3-8B-Base, Methodology=SPICE, Supervision Ratio=None2025.12 | 42.5 | |
| R-Few (1%)Backbone=Qwen3-4B-Base, Methodology=R-Few, Supervision Ratio=1%2025.12 | 42.4 | |
| SPIRALBackbone=Qwen3-4B, Training Data=Multi-game2025.06 | 41.8 | |
| Absolute ZeroBackbone=Qwen3-4B-Base, Methodology=Absolute Zero, Supervision Ratio=None2025.12 | 41.5 | |
| R-ZeroBackbone=Qwen3-4B-Base, Methodology=R-Zero, Supervision Ratio=None2025.12 | 40.6 | |
| Base ModelBackbone=Qwen3-8B-Base, Methodology=Base, Supervision Ratio=None2025.12 | 40.4 | |
| Beam Search (N=16)DM=7B2025.12 | 40.15 | |
| SPIRALBackbone=Qwen3-4B, Training Data=Kuhn Poker2025.06 | 38.4 | |
| SFTBackbone=Qwen3-4B, Training Data=Multi-game2025.06 | 37.6 | |
| SFTBackbone=Qwen3-4B, Training Data=Kuhn Poker2025.06 | 36.7 | |
| FIM-TIESModel Scale=1.5B, Avg. Length=411, Random Seeds=42026.03 | 36.3 | |
| Single ModelDM=7B2025.12 | 35.7 | |
| Base ModelBackbone=Qwen3-4B-Base, Methodology=Base, Supervision Ratio=None2025.12 | 34.8 | |
| ACM-TIESModel Scale=1.5B, Avg. Length=1,4892026.03 | 33.8 | |
| SPIRALBackbone=Octothinker-8B, Training Data=Multi-game2025.06 | 33.7 | |
| Qwen3-8B-BaseBackbone=Qwen3-8B2025.06 | 33.5 | |
| Qwen3-4B-BaseBackbone=Qwen3-4B2025.06 | 33.3 | |
| Qwen3-4B-Base2026.02 | 33.2 |