Mathematical Reasoning on DeepMath
66.2AccuracyRLVR (with verifier)
Evaluation Results
| Method | Links | |
|---|---|---|
| RLVR (with verifier)Model size=7B, Reasoning budget=2048 tokens2025.11 | 66.2 | |
| RAROModel size=7B, Reasoning budget=2048 tokens2025.11 | 57.5 | |
| RLVR (with verifier)Model size=3B, Reasoning budget=2048 tokens2025.11 | 55.8 | |
| RLVR (with verifier)Model size=1.5B, Reasoning budget=2048 tokens2025.11 | 50.9 | |
| RL-LogitModel size=7B, Reasoning budget=2048 tokens, Notes=best over 2 variants2025.11 | 49.3 | |
| RAROModel size=3B, Reasoning budget=2048 tokens2025.11 | 49.1 | |
| RationalizationModel size=7B, Reasoning budget=2048 tokens2025.11 | 48.6 | |
| PLR-4Backbone=Qwen2.5-7B-Instruct, k (number of ICL examples)=46.362026.03 | 46.36 | |
| PLR-EMABackbone=Qwen2.5-7B-Instruct, k (number of ICL examples)=322026.03 | 46.16 | |
| PLR-1Backbone=Qwen2.5-7B-Instruct, k (number of ICL examples)=322026.03 | 45.62 | |
| Top-KBackbone=Qwen2.5-7B-Instruct, k (number of ICL examples)=322026.03 | 45.13 | |
| BaseModel size=7B, Reasoning budget=2048 tokens2025.11 | 44.2 | |
| StaticBackbone=Qwen2.5-7B-Instruct, k (number of ICL examples)=322026.03 | 43.27 | |
| RL-LogitModel size=3B, Reasoning budget=2048 tokens, Notes=best over 2 variants2025.11 | 43.1 | |
| PLR-1Backbone=Qwen2.5-7B-Instruct, k (number of ICL examples)=162026.03 | 42.69 | |
| PLR-EMABackbone=Qwen2.5-7B-Instruct, k (number of ICL examples)=162026.03 | 42.52 | |
| PLR-4Backbone=Qwen2.5-7B-Instruct, k (number of ICL examples)=162026.03 | 42.46 | |
| SFTModel size=7B, Reasoning budget=2048 tokens2025.11 | 42.3 | |
| Top-KBackbone=Qwen2.5-7B-Instruct, k (number of ICL examples)=162026.03 | 42.2 | |
| StaticBackbone=Qwen2.5-7B-Instruct, k (number of ICL examples)=162026.03 | 41.85 | |
| PLR-EMABackbone=Qwen2.5-7B-Instruct, k (number of ICL examples)=82026.03 | 41.52 | |
| RAROModel size=1.5B, Reasoning budget=2048 tokens2025.11 | 41.3 | |
| PLR-1Backbone=Qwen2.5-7B-Instruct, k (number of ICL examples)=82026.03 | 39.93 | |
| BaseModel size=3B, Reasoning budget=2048 tokens2025.11 | 39.4 | |
| PLR-4Backbone=Qwen2.5-7B-Instruct, k (number of ICL examples)=82026.03 | 39.17 | |
| SFTModel size=3B, Reasoning budget=2048 tokens2025.11 | 39 | |
| Top-KBackbone=Qwen2.5-7B-Instruct, k (number of ICL examples)=82026.03 | 38.74 | |
| StaticBackbone=Qwen2.5-7B-Instruct, k (number of ICL examples)=82026.03 | 38.7 | |
| RL-LogitModel size=1.5B, Reasoning budget=2048 tokens, Notes=best over 2 variants2025.11 | 37.7 | |
| Iterative DPOModel size=7B, Reasoning budget=2048 tokens, Notes=max over 3 rounds2025.11 | 36.9 | |
| PLR-EMABackbone=Qwen2.5-7B-Instruct, k (number of ICL examples)=42026.03 | 36.64 | |
| PLR-1Backbone=Qwen2.5-7B-Instruct, k (number of ICL examples)=42026.03 | 36.64 | |
| PLR-4Backbone=Qwen2.5-7B-Instruct, k (number of ICL examples)=42026.03 | 36.64 | |
| Top-KBackbone=Qwen2.5-7B-Instruct, k (number of ICL examples)=42026.03 | 36.49 | |
| StaticBackbone=Qwen2.5-7B-Instruct, k (number of ICL examples)=42026.03 | 36.03 | |
| SFTModel size=1.5B, Reasoning budget=2048 tokens2025.11 | 35.7 | |
| RationalizationModel size=1.5B, Reasoning budget=2048 tokens2025.11 | 34.5 | |
| Iterative DPOModel size=3B, Reasoning budget=2048 tokens, Notes=max over 3 rounds2025.11 | 34.2 | |
| Iterative DPOModel size=1.5B, Reasoning budget=2048 tokens, Notes=max over 3 rounds2025.11 | 33 | |
| RationalizationModel size=3B, Reasoning budget=2048 tokens2025.11 | 32.3 | |
| BaseModel size=1.5B, Reasoning budget=2048 tokens2025.11 | 29.6 | |
| TATRA2026.02 | 27.78 | |
| PIASTvariant=(E)2026.02 | 25.33 | |
| PIAST2026.02 | 24.85 | |
| PRL2026.02 | 21.58 | |
| GPSvariant=SR-0.12026.02 | 21.58 | |
| PRL2025.05 | 21.58 | |
| GPSvariant=J2026.02 | 21.4 | |
| GA2026.02 | 18.63 | |
| GA2025.05 | 18.63 | |
| PromptWizard2025.05 | 18.52 | |
| PromptAgent2025.05 | 16.43 | |
| DE2026.02 | 16.1 | |
| DE2025.05 | 16.1 | |
| APE2026.02 | 15.47 | |
| APE2025.05 | 15.47 | |
| GRACE2026.02 | 15.05 | |
| GRACE2025.05 | 15.05 |