Algorithmic Reasoning on MATH (Accuracy)
80.6AccuracyRM-Regen
Evaluation Results
| Method | Links | |
|---|---|---|
| RM-RegenModel=Llama3.1-8B, Verification Setting=Oracle2026.03 | 80.6 | |
| Reflexion(3 iters)Model=Llama3.1-8B, Verification Setting=Oracle2026.03 | 77.8 | |
| Best-of-N (N=3)Model=Llama3.1-8B, Verification Setting=Oracle2026.03 | 76.8 | |
| ReflectEvoModel=Llama3.1-8B, Verification Setting=Oracle2026.03 | 73.6 | |
| ST CoTBase Model=Llama 3.1-8B, Iterations=22026.03 | 72.2 | |
| RM-RegenModel=Gemma2-9B, Verification Setting=Oracle2026.03 | 72.2 | |
| Reflexion(3 iters)Model=Gemma2-9B, Verification Setting=Oracle2026.03 | 72 | |
| RM-RegenBase Model=Llama 3.1-8B2026.03 | 71.4 | |
| RM-RegenModel=Llama3.1-8B2026.03 | 71.4 | |
| ST CoTBase Model=Llama 3.1-8B, Iterations=32026.03 | 69.8 | |
| ST CoTModel=Llama3.1-8B, iterations=32026.03 | 69.8 | |
| ProCoBase Model=Llama 3.1-8B, Iterations=22026.03 | 69.6 | |
| ProCoBase Model=Llama 3.1-8B, Iterations=32026.03 | 69.6 | |
| ProCoBase Model=Llama 3.1-8B, Iterations=52026.03 | 69.6 | |
| ProCoModel=Llama3.1-8B, iterations=32026.03 | 69.6 | |
| ProCoBase Model=Llama 3.1-8B, Iterations=42026.03 | 68.2 | |
| Best-of-N (N=3)Model=Gemma2-9B, Verification Setting=Oracle2026.03 | 67.2 | |
| ST CoTBase Model=Llama 3.1-8B, Iterations=42026.03 | 66.4 | |
| RM-RegenModel=Gemma2-9B2026.03 | 66.2 | |
| ProCoModel=Gemma2-9B, iterations=32026.03 | 64.6 | |
| Best-of-N (N=3)Model=GPT-3.5, Verification Setting=Oracle2026.03 | 64 | |
| ReflectEvoModel=Gemma2-9B, Verification Setting=Oracle2026.03 | 63.8 | |
| RM-RegenModel=GPT-3.5, Verification Setting=Oracle2026.03 | 63.4 | |
| ST CoTBase Model=Llama 3.1-8B, Iterations=52026.03 | 63 | |
| Reflexion(3 iters)Model=GPT-3.5, Verification Setting=Oracle2026.03 | 62.25 | |
| ST CoTModel=Gemma2-9B, iterations=32026.03 | 59.4 | |
| RM-RegenBase Model=GPT-3.52026.03 | 56.8 | |
| RM-RegenModel=GPT-3.52026.03 | 56.8 | |
| ST CoTBase Model=GPT-3.5, Iterations=22026.03 | 54.8 | |
| Self-RefineBase Model=Llama 3.1-8B, Iterations=22026.03 | 54.8 | |
| ProCoBase Model=GPT-3.5, Iterations=22026.03 | 54.6 | |
| ST CoTBase Model=GPT-3.5, Iterations=32026.03 | 53.6 | |
| ST CoTModel=GPT-3.5, iterations=32026.03 | 53.6 | |
| Self-RefineBase Model=Llama 3.1-8B, Iterations=42026.03 | 53.2 | |
| ST CoTBase Model=GPT-3.5, Iterations=42026.03 | 53 | |
| ProCoBase Model=GPT-3.5, Iterations=32026.03 | 52.2 | |
| ProCoModel=GPT-3.5, iterations=32026.03 | 52.2 | |
| ProCoBase Model=GPT-3.5, Iterations=42026.03 | 52.08 | |
| Self-RefineBase Model=Llama 3.1-8B, Iterations=32026.03 | 51 | |
| Self-RefineModel=Llama3.1-8B, iterations=32026.03 | 51 | |
| Self-RefineBase Model=Llama 3.1-8B, Iterations=52026.03 | 49.6 | |
| Self-RefineBase Model=GPT-3.5, Iterations=42026.03 | 48.8 | |
| Self-RefineBase Model=GPT-3.5, Iterations=22026.03 | 48.2 | |
| Self-RefineBase Model=GPT-3.5, Iterations=32026.03 | 47.2 | |
| Self-RefineModel=GPT-3.5, iterations=32026.03 | 47.2 | |
| Self-RefineModel=Gemma2-9B, iterations=32026.03 | 27.8 |