Reasoning on Bamboogle
73AccuracyRM-Regen
Evaluation Results
| Method | Links | |
|---|---|---|
| RM-RegenModel=GPT-3.5, Verification Setting=Oracle2026.03 | 73 | |
| Reflexion(3 iters)Model=GPT-3.5, Verification Setting=Oracle2026.03 | 66 | |
| Best-of-N (N=3)Model=GPT-3.5, Verification Setting=Oracle2026.03 | 63 | |
| RM-RegenModel=Llama3.1-8B, Verification Setting=Oracle2026.03 | 63 | |
| Reflexion(3 iters)Model=Llama3.1-8B, Verification Setting=Oracle2026.03 | 60 | |
| ST CoTBase Model=GPT-3.5, Iterations=32026.03 | 59 | |
| RM-RegenBase Model=GPT-3.52026.03 | 59 | |
| ST CoTModel=GPT-3.5, iterations=32026.03 | 59 | |
| RM-RegenModel=GPT-3.52026.03 | 59 | |
| ReflectEvoModel=Llama3.1-8B, Verification Setting=Oracle2026.03 | 53 | |
| Best-of-N (N=3)Model=Llama3.1-8B, Verification Setting=Oracle2026.03 | 53 | |
| RM-RegenModel=Gemma2-9B, Verification Setting=Oracle2026.03 | 53 | |
| RM-RegenBase Model=Llama 3.1-8B2026.03 | 52 | |
| RM-RegenModel=Llama3.1-8B2026.03 | 52 | |
| ST CoTBase Model=GPT-3.5, Iterations=22026.03 | 51 | |
| ProCoBase Model=GPT-3.5, Iterations=42026.03 | 50 | |
| Reflexion(3 iters)Model=Gemma2-9B, Verification Setting=Oracle2026.03 | 50 | |
| ProCoBase Model=GPT-3.5, Iterations=22026.03 | 49 | |
| Self-RefineBase Model=Llama 3.1-8B, Iterations=32026.03 | 49 | |
| Self-RefineModel=Llama3.1-8B, iterations=32026.03 | 49 | |
| ST CoTBase Model=GPT-3.5, Iterations=42026.03 | 48 | |
| ProCoBase Model=Llama 3.1-8B, Iterations=42026.03 | 48 | |
| ST CoTBase Model=Llama 3.1-8B, Iterations=22026.03 | 48 | |
| ST CoTBase Model=Llama 3.1-8B, Iterations=32026.03 | 48 | |
| ST CoTModel=Llama3.1-8B, iterations=32026.03 | 48 | |
| RM-RegenModel=Gemma2-9B2026.03 | 48 | |
| Self-RefineBase Model=Llama 3.1-8B, Iterations=22026.03 | 47 | |
| ReflectEvoModel=Gemma2-9B, Verification Setting=Oracle2026.03 | 47 | |
| Self-RefineBase Model=Llama 3.1-8B, Iterations=52026.03 | 46 | |
| Best-of-N (N=3)Model=Gemma2-9B, Verification Setting=Oracle2026.03 | 46 | |
| ST CoTModel=Gemma2-9B, iterations=32026.03 | 46 | |
| Self-RefineBase Model=Llama 3.1-8B, Iterations=42026.03 | 45 | |
| ST CoTBase Model=Llama 3.1-8B, Iterations=52026.03 | 45 | |
| Self-RefineBase Model=GPT-3.5, Iterations=42026.03 | 43 | |
| ProCoBase Model=GPT-3.5, Iterations=32026.03 | 43 | |
| ST CoTBase Model=Llama 3.1-8B, Iterations=42026.03 | 43 | |
| ProCoModel=GPT-3.5, iterations=32026.03 | 43 | |
| Self-RefineBase Model=GPT-3.5, Iterations=32026.03 | 42 | |
| ProCoBase Model=Llama 3.1-8B, Iterations=52026.03 | 42 | |
| Self-RefineModel=GPT-3.5, iterations=32026.03 | 42 | |
| Self-RefineBase Model=GPT-3.5, Iterations=22026.03 | 40 | |
| ProCoBase Model=Llama 3.1-8B, Iterations=32026.03 | 39 | |
| ProCoModel=Llama3.1-8B, iterations=32026.03 | 39 | |
| ProCoModel=Gemma2-9B, iterations=32026.03 | 38 | |
| ProCoBase Model=Llama 3.1-8B, Iterations=22026.03 | 35 | |
| Self-RefineModel=Gemma2-9B, iterations=32026.03 | 16 |