Mathematical Reasoning on GSM Hard (Accuracy)
68.6AccuracyRM-Regen
Evaluation Results
| Method | Links | |
|---|---|---|
| RM-RegenModel=GPT-3.5, Verification Setting=Oracle2026.03 | 68.6 | |
| RM-RegenModel=GPT-3.52026.03 | 64 | |
| RM-RegenModel=Gemma2-9B, Verification Setting=Oracle2026.03 | 53 | |
| Reflexion(3 iters)Model=Gemma2-9B, Verification Setting=Oracle2026.03 | 50.8 | |
| RM-RegenModel=Gemma2-9B2026.03 | 49.4 | |
| ProCoModel=Gemma2-9B, iterations=32026.03 | 48.6 | |
| Best-of-N (N=3)Model=Gemma2-9B, Verification Setting=Oracle2026.03 | 48.4 | |
| Best-of-N (N=3)Model=GPT-3.5, Verification Setting=Oracle2026.03 | 46.8 | |
| ST CoTModel=Gemma2-9B, iterations=32026.03 | 46.2 | |
| Reflexion(3 iters)Model=GPT-3.5, Verification Setting=Oracle2026.03 | 45.2 | |
| ReflectEvoModel=Gemma2-9B, Verification Setting=Oracle2026.03 | 44.4 | |
| RM-RegenModel=Llama3.1-8B, Verification Setting=Oracle2026.03 | 41 | |
| ST CoTModel=GPT-3.5, iterations=32026.03 | 39.8 | |
| ProCoModel=GPT-3.5, iterations=32026.03 | 39.6 | |
| NEXATrained on=GSM8K, Evaluation Protocol=Without retraining2026.05 | 37.13 | |
| Reflexion(3 iters)Model=Llama3.1-8B, Verification Setting=Oracle2026.03 | 36.2 | |
| CoTTrained on=GSM8K, Evaluation Protocol=Without retraining2026.05 | 36.07 | |
| RM-RegenModel=Llama3.1-8B2026.03 | 34.8 | |
| SingleTrained on=GSM8K, Evaluation Protocol=Without retraining2026.05 | 34.4 | |
| Best-of-N (N=3)Model=Llama3.1-8B, Verification Setting=Oracle2026.03 | 34.2 | |
| Self-RefineModel=GPT-3.5, iterations=32026.03 | 32.8 | |
| Self-RefineModel=Gemma2-9B, iterations=32026.03 | 32.8 | |
| ReflectEvoModel=Llama3.1-8B, Verification Setting=Oracle2026.03 | 32 | |
| ST CoTModel=Llama3.1-8B, iterations=32026.03 | 31.8 | |
| ProCoModel=Llama3.1-8B, iterations=32026.03 | 31.2 | |
| ALIGNReasoning Strategy=ALIGN2026.01 | 24.6 | |
| SCReasoning Strategy=CoT-SC@maj162026.01 | 24.1 | |
| Self-RefineModel=Llama3.1-8B, iterations=32026.03 | 23 | |
| rStarReasoning Strategy=rStar2026.01 | 22.8 | |
| FS-CoTReasoning Strategy=Few-shot CoT2026.01 | 17.7 | |
| PrincipalReasoning Strategy=Principal2026.01 | 15.6 |