Symbolic Reasoning on Coin
100AccuracyMeta-Reasoning Paradigm
Evaluation Results
| Method | Links | |
|---|---|---|
| Meta-Reasoning ParadigmModel=175B GPT-3 (text-davinci-002), Evaluation Protocol=Few-Shot2023.06 | 100 | |
| Meta-Reasoning ParadigmModel=175B GPT-3 (text-davinci-003), Evaluation Protocol=Few-Shot2023.06 | 100 | |
| Meta-Reasoning ParadigmModel=ChatGPT (GPT-3.5-Turbo), Evaluation Protocol=Few-Shot2023.06 | 100 | |
| Chain-of-Thought ParadigmModel=175B GPT-3 (text-davinci-003), Evaluation Protocol=Few-Shot2023.06 | 99.6 | |
| Chain-of-Thought ParadigmModel=ChatGPT (GPT-3.5-Turbo), Evaluation Protocol=Few-Shot2023.06 | 99.2 | |
| Chain-of-Thought ParadigmModel=175B GPT-3 (text-davinci-002), Evaluation Protocol=Few-Shot2023.06 | 97.2 | |
| Chain-of-Thought ParadigmModel=175B GPT-3 (text-davinci-003), Evaluation Protocol=Zero-Shot2023.06 | 96.8 | |
| Chain-of-Thought ParadigmModel=ChatGPT (GPT-3.5-Turbo), Evaluation Protocol=Zero-Shot2023.06 | 96.4 | |
| Chain-of-Thought ParadigmModel=175B GPT-3 (text-davinci-002), Evaluation Protocol=Zero-Shot2023.06 | 91.4 | |
| Reflexion(3 iters)Model=GPT-3.5, Verification Setting=Oracle2026.03 | 88.25 | |
| RM-RegenModel=GPT-3.5, Verification Setting=Oracle2026.03 | 85.75 | |
| RM-RegenModel=Llama3.1-8B, Verification Setting=Oracle2026.03 | 85.75 | |
| Reflexion(3 iters)Model=Llama3.1-8B, Verification Setting=Oracle2026.03 | 81.25 | |
| RM-RegenBase Model=GPT-3.52026.03 | 78.5 | |
| RM-RegenModel=Gemma2-9B, Verification Setting=Oracle2026.03 | 78.25 | |
| RM-RegenBase Model=Llama 3.1-8B2026.03 | 77.75 | |
| Best-of-N (N=3)Model=GPT-3.5, Verification Setting=Oracle2026.03 | 77 | |
| ST CoTBase Model=GPT-3.5, Iterations=22026.03 | 75.5 | |
| ProCoBase Model=GPT-3.5, Iterations=32026.03 | 75.25 | |
| ST CoTBase Model=GPT-3.5, Iterations=32026.03 | 75 | |
| ST CoTBase Model=GPT-3.5, Iterations=42026.03 | 74.5 | |
| ProCoBase Model=GPT-3.5, Iterations=42026.03 | 73.75 | |
| ProCoBase Model=GPT-3.5, Iterations=22026.03 | 73 | |
| Best-of-N (N=3)Model=Llama3.1-8B, Verification Setting=Oracle2026.03 | 72 | |
| Self-RefineBase Model=GPT-3.5, Iterations=42026.03 | 69.75 | |
| ST CoTBase Model=Llama 3.1-8B, Iterations=22026.03 | 69.5 | |
| ST CoTBase Model=Llama 3.1-8B, Iterations=32026.03 | 69.25 | |
| Self-RefineBase Model=GPT-3.5, Iterations=32026.03 | 68.25 | |
| ST CoTBase Model=Llama 3.1-8B, Iterations=42026.03 | 67.75 | |
| ReflectEvoModel=Llama3.1-8B, Verification Setting=Oracle2026.03 | 67.75 | |
| Self-RefineBase Model=GPT-3.5, Iterations=22026.03 | 67.25 | |
| Self-RefineBase Model=Llama 3.1-8B, Iterations=22026.03 | 67.25 | |
| ReflectEvoModel=Gemma2-9B, Verification Setting=Oracle2026.03 | 67.25 | |
| Self-RefineBase Model=Llama 3.1-8B, Iterations=32026.03 | 66 | |
| Self-RefineBase Model=Llama 3.1-8B, Iterations=52026.03 | 66 | |
| ST CoTBase Model=Llama 3.1-8B, Iterations=52026.03 | 65.5 | |
| Self-RefineBase Model=Llama 3.1-8B, Iterations=42026.03 | 64.25 | |
| Reflexion(3 iters)Model=Gemma2-9B, Verification Setting=Oracle2026.03 | 60 | |
| Standard Prompting ParadigmModel=175B GPT-3 (text-davinci-002), Evaluation Protocol=Few-Shot2023.06 | 57.2 | |
| Standard Prompting ParadigmModel=175B GPT-3 (text-davinci-002), Evaluation Protocol=Zero-Shot2023.06 | 53.8 | |
| Best-of-N (N=3)Model=Gemma2-9B, Verification Setting=Oracle2026.03 | 49 | |
| ProCoBase Model=Llama 3.1-8B, Iterations=42026.03 | 48 | |
| ProCoBase Model=Llama 3.1-8B, Iterations=52026.03 | 44.75 | |
| ProCoBase Model=Llama 3.1-8B, Iterations=32026.03 | 44 | |
| ProCoBase Model=Llama 3.1-8B, Iterations=22026.03 | 42 |