Algorithmic Reasoning on GSM8K (Accuracy)
92.6AccuracyRM-Regen
Evaluation Results
| Method | Links | |
|---|---|---|
| RM-RegenModel=Llama3.1-8B, Verification Setting=Oracle2026.03 | 92.6 | |
| RM-RegenModel=Gemma2-9B, Verification Setting=Oracle2026.03 | 92 | |
| Reflexion(3 iters)Model=Gemma2-9B, Verification Setting=Oracle2026.03 | 91.2 | |
| RM-RegenModel=GPT-3.5, Verification Setting=Oracle2026.03 | 90.2 | |
| RM-RegenModel=Gemma2-9B2026.03 | 89.8 | |
| Reflexion(3 iters)Model=Llama3.1-8B, Verification Setting=Oracle2026.03 | 89.6 | |
| Best-of-N (N=3)Model=Llama3.1-8B, Verification Setting=Oracle2026.03 | 89.6 | |
| ProCoModel=Gemma2-9B, iterations=32026.03 | 89.6 | |
| Best-of-N (N=3)Model=Gemma2-9B, Verification Setting=Oracle2026.03 | 88.8 | |
| ST CoTModel=Gemma2-9B, iterations=32026.03 | 88 | |
| RM-RegenModel=Llama3.1-8B2026.03 | 87.2 | |
| ST CoTModel=Llama3.1-8B, iterations=32026.03 | 87 | |
| ReflectEvoModel=Gemma2-9B, Verification Setting=Oracle2026.03 | 85.8 | |
| Reflexion(3 iters)Model=GPT-3.5, Verification Setting=Oracle2026.03 | 85.6 | |
| ReflectEvoModel=Llama3.1-8B, Verification Setting=Oracle2026.03 | 85.4 | |
| Best-of-N (N=3)Model=GPT-3.5, Verification Setting=Oracle2026.03 | 84.6 | |
| RM-RegenModel=GPT-3.52026.03 | 83.6 | |
| ProCoModel=Llama3.1-8B, iterations=32026.03 | 83.4 | |
| ProCoModel=GPT-3.5, iterations=32026.03 | 80.8 | |
| ST CoTModel=GPT-3.5, iterations=32026.03 | 80.2 | |
| Self-RefineModel=Llama3.1-8B, iterations=32026.03 | 75.6 | |
| Self-RefineModel=GPT-3.5, iterations=32026.03 | 70.4 | |
| Self-RefineModel=Gemma2-9B, iterations=32026.03 | 40 |