Temporal Reasoning on TempQuestions (test)
60.3Exact Match (EM)QAaP
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| QAaPBackbone=gpt-3.5-turbo, Evaluation Protocol=Few-shot2023.05 | 60.3 | 68.1 | |
| CoTBackbone=gpt-3.5-turbo, Evaluation Protocol=Few-shot2023.05 | 49.8 | 60.1 | |
| TEQUILAEvaluation Protocol=Fine-tune2023.05 | 43.8 | 44.6 | |
| Rethinking with retrievalModel=GPT-3 (text-davinci-002), Temperature=0.7, Num samples=102022.12 | 39.05 | — | |
| Self-consistencyModel=GPT-3 (text-davinci-002), Temperature=0.7, Num samples=102022.12 | 37.28 | — | |
| Chain-of-thought promptingModel=GPT-3 (text-davinci-002), Temperature=02022.12 | 33.14 | — | |
| Few-shot promptingModel=GPT-3 (text-davinci-002), Temperature=02022.12 | 29.59 | — | |
| Zero-shot promptingModel=GPT-3 (text-davinci-002), Temperature=02022.12 | 28.4 | — | |
| ReActBackbone=gpt-3.5-turbo, Evaluation Protocol=Few-shot2023.05 | 28 | 33.7 |