Multi-hop Reasoning on CommaQA-E compositional
80.8Exact MatchChatGPT (SKiC)
Evaluation Results
| Method | Links | |
|---|---|---|
| ChatGPT (SKiC)Model=ChatGPT, Prompting=SKiC†, #-shots=22023.08 | 80.8 | |
| text-davinci-003 (SKiC)Model=text-davinci-003, Prompting=SKiC†, #-shots=22023.08 | 74.8 | |
| ChatGPT (Decomp)Model=ChatGPT, Prompting=Decomp, #-shots=122023.08 | 73.5 | |
| text-davinci-003 (Decomp)Model=text-davinci-003, Prompting=Decomp, #-shots=122023.08 | 66.6 | |
| LLAMA-65B (SKiC)Model=LLAMA-65B, Prompting=SKiC†, #-shots=22023.08 | 52 | |
| ChatGPT (CoT)Model=ChatGPT, Prompting=CoT, #-shots=42023.08 | 46.4 | |
| LLAMA-65B (Decomp)Model=LLAMA-65B, Prompting=Decomp, #-shots=122023.08 | 40.4 | |
| ChatGPT (4-shots)Model=ChatGPT, Prompting=4-shots, #-shots=42023.08 | 40.3 | |
| text-davinci-003 (CoT)Model=text-davinci-003, Prompting=CoT, #-shots=42023.08 | 38.2 | |
| text-davinci-003 (4-shots)Model=text-davinci-003, Prompting=4-shots, #-shots=42023.08 | 33.5 | |
| LLAMA-65B (CoT)Model=LLAMA-65B, Prompting=CoT, #-shots=42023.08 | 30.8 | |
| ChatGPT (zero-shot)Model=ChatGPT, Prompting=zero-shot, #-shots=02023.08 | 30.6 | |
| text-davinci-003 (zero-shot)Model=text-davinci-003, Prompting=zero-shot, #-shots=02023.08 | 26.8 | |
| LLAMA-65B (4-shots)Model=LLAMA-65B, Prompting=4-shots, #-shots=42023.08 | 24.6 | |
| LLAMA-65B (zero-shot)Model=LLAMA-65B, Prompting=zero-shot, #-shots=02023.08 | 16.3 |