Logical Reasoning on Track
100Track(Avg)Meta-Reasoning Paradigm
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Meta-Reasoning ParadigmModel=175B GPT-3 (text-davinci-003), Evaluation Protocol=Few-Shot2023.06 | 100 | 100 | 100 | 100 | |
| Meta-Reasoning ParadigmModel=175B GPT-3 (text-davinci-002), Evaluation Protocol=Few-Shot2023.06 | 98.8 | 97.2 | 100 | 99.2 | |
| Meta-Reasoning ParadigmModel=ChatGPT (GPT-3.5-Turbo), Evaluation Protocol=Few-Shot2023.06 | 90.8 | 100 | 88 | 84.4 | |
| Chain-of-Thought ParadigmModel=175B GPT-3 (text-davinci-003), Evaluation Protocol=Few-Shot2023.06 | 76.8 | 68.4 | 80.8 | 81.2 | |
| Chain-of-Thought ParadigmModel=175B GPT-3 (text-davinci-002), Evaluation Protocol=Few-Shot2023.06 | 61.1 | 62.8 | 60.8 | 59.6 | |
| Chain-of-Thought ParadigmModel=ChatGPT (GPT-3.5-Turbo), Evaluation Protocol=Few-Shot2023.06 | 58 | 62.8 | 57.2 | 54 | |
| Chain-of-Thought ParadigmModel=ChatGPT (GPT-3.5-Turbo), Evaluation Protocol=Zero-Shot2023.06 | 50.9 | 55.6 | 54 | 43.2 | |
| Chain-of-Thought ParadigmModel=175B GPT-3 (text-davinci-002), Evaluation Protocol=Zero-Shot2023.06 | 35.5 | 44.8 | 35.6 | 26 | |
| Chain-of-Thought ParadigmModel=175B GPT-3 (text-davinci-003), Evaluation Protocol=Zero-Shot2023.06 | 34.7 | 37.2 | 36 | 30.8 | |
| Standard Prompting ParadigmModel=175B GPT-3 (text-davinci-002), Evaluation Protocol=Few-Shot2023.06 | 25.1 | — | — | — | |
| State-of-the-ArtParadigm=Fine-tuned2023.06 | 24.1 | — | — | — | |
| Standard Prompting ParadigmModel=175B GPT-3 (text-davinci-002), Evaluation Protocol=Zero-Shot2023.06 | 15.7 | 24.4 | 15.2 | 7.6 |