Symbolic Reasoning on Letter
92.4AccuracyMeta-Reasoning Paradigm
Evaluation Results
| Method | Links | |
|---|---|---|
| Meta-Reasoning ParadigmModel=ChatGPT (GPT-3.5-Turbo), Evaluation Protocol=Few-Shot2023.06 | 92.4 | |
| Meta-Reasoning ParadigmModel=175B GPT-3 (text-davinci-003), Evaluation Protocol=Few-Shot2023.06 | 91.6 | |
| RM-RegenModel=Llama3.1-8B, Verification Setting=Oracle2026.03 | 89.33 | |
| Self-AgreementBackbone=GPT-3.5-turbo, Scenario=Question type unknown, answer format known2023.11 | 88.9 | |
| RM-RegenModel=GPT-3.5, Verification Setting=Oracle2026.03 | 87.33 | |
| Meta-Reasoning ParadigmModel=175B GPT-3 (text-davinci-002), Evaluation Protocol=Few-Shot2023.06 | 86 | |
| Self-AgreementBackbone=GPT-3.5-turbo, Scenario=Question type and answer format unknown2023.11 | 83.8 | |
| Mixed-Few-Shot CoTBackbone=GPT-3.5-turbo, Scenario=Question type unknown, answer format known2023.11 | 83 | |
| Zero-Shot CoTBackbone=GPT-3.5-turbo, Scenario=Question type and answer format unknown2023.11 | 81 | |
| Reflexion(3 iters)Model=Llama3.1-8B, Verification Setting=Oracle2026.03 | 80.67 | |
| Chain-of-Thought ParadigmModel=ChatGPT (GPT-3.5-Turbo), Evaluation Protocol=Few-Shot2023.06 | 80.2 | |
| RM-RegenBase Model=Llama 3.1-8B2026.03 | 79.33 | |
| Self-ConsistencyBackbone=GPT-3.5-turbo, Scenario=Question type unknown, answer format known2023.11 | 79.1 | |
| Reflexion(3 iters)Model=GPT-3.5, Verification Setting=Oracle2026.03 | 78.67 | |
| RM-RegenBase Model=GPT-3.52026.03 | 77.33 | |
| Chain-of-Thought ParadigmModel=ChatGPT (GPT-3.5-Turbo), Evaluation Protocol=Zero-Shot2023.06 | 75.6 | |
| ProCoBase Model=GPT-3.5, Iterations=22026.03 | 74.67 | |
| Best-of-N (N=3)Model=GPT-3.5, Verification Setting=Oracle2026.03 | 74.67 | |
| ProCoBase Model=GPT-3.5, Iterations=32026.03 | 74 | |
| Self-consistencyModel=GPT-3 (Code-davinci-002)2022.03 | 73.4 | |
| ST CoTBase Model=GPT-3.5, Iterations=22026.03 | 72.67 | |
| ST CoTBase Model=GPT-3.5, Iterations=32026.03 | 72.67 | |
| ST CoTBase Model=GPT-3.5, Iterations=42026.03 | 72.67 | |
| ProCoBase Model=Llama 3.1-8B, Iterations=32026.03 | 71.33 | |
| ProCoBase Model=Llama 3.1-8B, Iterations=52026.03 | 71.33 | |
| Self-consistencyModel=PaLM-540B2022.03 | 70.8 | |
| ProCoBase Model=Llama 3.1-8B, Iterations=22026.03 | 70.67 | |
| Best-of-N (N=3)Model=Llama3.1-8B, Verification Setting=Oracle2026.03 | 70.67 | |
| Chain-of-Thought ParadigmModel=175B GPT-3 (text-davinci-003), Evaluation Protocol=Few-Shot2023.06 | 70.6 | |
| CoT-promptingModel=GPT-3 (Code-davinci-002)2022.03 | 70.4 | |
| ProCoBase Model=GPT-3.5, Iterations=42026.03 | 68.67 | |
| ReflectEvoModel=Llama3.1-8B, Verification Setting=Oracle2026.03 | 68.67 | |
| CoT-promptingModel=PaLM-540B2022.03 | 65.8 | |
| Chain-of-Thought ParadigmModel=175B GPT-3 (text-davinci-003), Evaluation Protocol=Zero-Shot2023.06 | 64.8 | |
| ProCoBase Model=Llama 3.1-8B, Iterations=42026.03 | 64.67 | |
| ST CoTBase Model=Llama 3.1-8B, Iterations=22026.03 | 62 | |
| Multi-Agents (Debate)Backbone=GPT-3.5-turbo, Scenario=Question type unknown, answer format known2023.11 | 61.3 | |
| ST CoTBase Model=Llama 3.1-8B, Iterations=42026.03 | 59.33 | |
| Chain-of-Thought ParadigmModel=175B GPT-3 (text-davinci-002), Evaluation Protocol=Few-Shot2023.06 | 59 | |
| ST CoTBase Model=Llama 3.1-8B, Iterations=32026.03 | 58 | |
| ST CoTBase Model=Llama 3.1-8B, Iterations=52026.03 | 58 | |
| Chain-of-Thought ParadigmModel=175B GPT-3 (text-davinci-002), Evaluation Protocol=Zero-Shot2023.06 | 57.6 | |
| Reflexion(3 iters)Model=Gemma2-9B, Verification Setting=Oracle2026.03 | 51.33 | |
| Self-RefineBase Model=Llama 3.1-8B, Iterations=22026.03 | 47.33 | |
| Self-RefineBase Model=GPT-3.5, Iterations=32026.03 | 46 | |
| Self-AgreementBackbone=Llama-2-13B-Chat, Scenario=Question type and answer format unknown2023.11 | 44.5 | |
| RM-RegenModel=Gemma2-9B, Verification Setting=Oracle2026.03 | 44.29 | |
| ReflectEvoModel=Gemma2-9B, Verification Setting=Oracle2026.03 | 42.14 | |
| Self-RefineBase Model=GPT-3.5, Iterations=42026.03 | 41.33 | |
| Self-RefineBase Model=Llama 3.1-8B, Iterations=42026.03 | 41.33 | |
| Self-RefineBase Model=GPT-3.5, Iterations=22026.03 | 40 | |
| Self-RefineBase Model=Llama 3.1-8B, Iterations=32026.03 | 40 | |
| Best-of-N (N=3)Model=Gemma2-9B, Verification Setting=Oracle2026.03 | 34.67 | |
| Zero-Shot CoTBackbone=Llama-2-13B-Chat, Scenario=Question type and answer format unknown2023.11 | 31 | |
| Self-RefineBase Model=Llama 3.1-8B, Iterations=52026.03 | 30.67 | |
| Multi-Agents (Debate)Backbone=Llama-2-13B-Chat, Scenario=Question type unknown, answer format known2023.11 | 27 | |
| Self-AgreementBackbone=Llama-2-13B-Chat, Scenario=Question type unknown, answer format known2023.11 | 23.1 | |
| Mixed-Few-Shot CoTBackbone=Llama-2-13B-Chat, Scenario=Question type unknown, answer format known2023.11 | 19 | |
| Self-ConsistencyBackbone=Llama-2-13B-Chat, Scenario=Question type unknown, answer format known2023.11 | 16 | |
| Self-consistencyModel=GPT-3 (Code-davinci-001)2022.03 | 10 | |
| CoT-promptingModel=LaMDA-137B2022.03 | 8.2 | |
| Self-consistencyModel=LaMDA-137B2022.03 | 8.2 | |
| CoT-promptingModel=GPT-3 (Code-davinci-001)2022.03 | 7.8 | |
| Standard Prompting ParadigmModel=175B GPT-3 (text-davinci-002), Evaluation Protocol=Zero-Shot2023.06 | 0.2 | |
| Standard Prompting ParadigmModel=175B GPT-3 (text-davinci-002), Evaluation Protocol=Few-Shot2023.06 | 0.2 | |
| CoT-promptingModel=UL2-20B2022.03 | 0 | |
| Self-consistencyModel=UL2-20B2022.03 | 0 |