Temporal Question Answering on TimeQA Easy-mode
85.4Exact Match (EM)DeepSeek-V3-AdapTime
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| DeepSeek-V3-AdapTimeBackbone=DeepSeek-V3-0324, Strategy=AdapTime2026.04 | 85.4 | 86.6 | |
| DeepSeek-V3-CoTBackbone=DeepSeek-V3-0324, Strategy=Chain-of-Thought (CoT)2026.04 | 85.3 | 86.7 | |
| DeepSeek-V3-Step-backBackbone=DeepSeek-V3-0324, Strategy=Step-back2026.04 | 84.4 | 86 | |
| DeepSeek-V3-ICLBackbone=DeepSeek-V3-0324, Strategy=In-context Learning (ICL)2026.04 | 80.8 | 82.9 | |
| DeepSeek-V3-Self-refinementBackbone=DeepSeek-V3-0324, Strategy=Self-refinement2026.04 | 77.6 | 80.4 | |
| Qwen-3-8B-AdapTimeBackbone=Qwen-3-8B, Strategy=AdapTime2026.04 | 72.7 | 74.1 | |
| GPT-4Backbone=GPT-4, Strategy=Vanilla2026.04 | 71.6 | 74.2 | |
| Qwen-3-8B-CoTBackbone=Qwen-3-8B, Strategy=Chain-of-Thought (CoT)2026.04 | 69.4 | 71 | |
| Qwen-3-8B-ICLBackbone=Qwen-3-8B, Strategy=In-context Learning (ICL)2026.04 | 67.5 | 70.3 | |
| TG-LLMStrategy=Temporal Reasoning2026.04 | 66.4 | 69.1 | |
| REMEMO-largeBackbone=REMEMO-large, Strategy=Supervised Fine-tuning2026.04 | 63.7 | 72.3 | |
| T5-largeBackbone=T5-large, Strategy=Supervised Fine-tuning2026.04 | 63.1 | 71.6 | |
| REMEMO-baseBackbone=REMEMO-base, Strategy=Supervised Fine-tuning2026.04 | 61.4 | 70.4 | |
| T5-baseBackbone=T5-base, Strategy=Supervised Fine-tuning2026.04 | 60 | 68.2 | |
| QAaPStrategy=Embedding-based Reasoning2026.04 | 48.2 | 58.3 | |
| LLaMA-3-8B-AdapTimeBackbone=LLaMA-3.1-8B-Instruct, Strategy=AdapTime2026.04 | 41.5 | 47.2 | |
| LLaMA-3-8B-CoTBackbone=LLaMA-3.1-8B-Instruct, Strategy=Chain-of-Thought (CoT)2026.04 | 29.7 | 32.5 | |
| LLaMA-3-8B-ICLBackbone=LLaMA-3.1-8B-Instruct, Strategy=In-context Learning (ICL)2026.04 | 1.1 | 3.1 |