Temporal Question Answering on ReasonQA Multi-hop
85Set AccuracyT5-large PIT-SFT
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| T5-large PIT-SFTBackbone=T5-large, Training strategy=PIT-SFT2023.11 | 85 | 89.5 | |
| T5-base PIT-SFTBackbone=T5-base, Training strategy=PIT-SFT2023.11 | 78 | 82.4 | |
| T5-large SFTBackbone=T5-large, Training strategy=Supervised fine-tuning2023.11 | 71 | 76.4 | |
| T5-base SFTBackbone=T5-base, Training strategy=Supervised fine-tuning2023.11 | 59.1 | 65.1 | |
| GPT-4Model version=gpt4-0613, Few-shot=one-shot2023.11 | 51.6 | 65.4 | |
| FLAN-T5-XLParameters=3B, Few-shot=one-shot2023.11 | 35.5 | 49.7 | |
| GPT-3.5Model version=gpt3.5-turbo, Few-shot=one-shot2023.11 | 31.2 | 51.8 |