Temporal Question Answering on TempReason OBQA-L3
49.8Exact Match (EM)DeepSeek-V3-AdapTime
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| DeepSeek-V3-AdapTimeBackbone=DeepSeek-V3-0324, Strategy=AdapTime2026.04 | 49.8 | 53.2 | |
| DeepSeek-V3-Step-backBackbone=DeepSeek-V3-0324, Strategy=Step-back2026.04 | 48.8 | 52.3 | |
| DeepSeek-V3-CoTBackbone=DeepSeek-V3-0324, Strategy=Chain-of-Thought (CoT)2026.04 | 47 | 50.4 | |
| DeepSeek-V3-ICLBackbone=DeepSeek-V3-0324, Strategy=In-context Learning (ICL)2026.04 | 43.6 | 48.3 | |
| GPT-4Backbone=GPT-4, Strategy=Vanilla2026.04 | 43.1 | 48.5 | |
| DeepSeek-V3-Self-refinementBackbone=DeepSeek-V3-0324, Strategy=Self-refinement2026.04 | 41.1 | 42.3 | |
| TG-LLMStrategy=Temporal Reasoning2026.04 | 35.6 | 46.9 | |
| REMEMO-largeBackbone=REMEMO-large, Strategy=Supervised Fine-tuning2026.04 | 33.4 | 49.3 | |
| T5-largeBackbone=T5-large, Strategy=Supervised Fine-tuning2026.04 | 28.8 | 46.8 | |
| Qwen-3-8B-CoTBackbone=Qwen-3-8B, Strategy=Chain-of-Thought (CoT)2026.04 | 28.8 | 33.4 | |
| Qwen-3-8B-AdapTimeBackbone=Qwen-3-8B, Strategy=AdapTime2026.04 | 28.8 | 33.8 | |
| REMEMO-baseBackbone=REMEMO-base, Strategy=Supervised Fine-tuning2026.04 | 28.5 | 44.9 | |
| T5-baseBackbone=T5-base, Strategy=Supervised Fine-tuning2026.04 | 23.8 | 41.8 | |
| Qwen-3-8B-ICLBackbone=Qwen-3-8B, Strategy=In-context Learning (ICL)2026.04 | 23.1 | 32.7 | |
| LLaMA-3-8B-CoTBackbone=LLaMA-3.1-8B-Instruct, Strategy=Chain-of-Thought (CoT)2026.04 | 16.5 | 24.1 | |
| LLaMA-3-8B-AdapTimeBackbone=LLaMA-3.1-8B-Instruct, Strategy=AdapTime2026.04 | 14.5 | 22.5 | |
| LLaMA-3-8B-ICLBackbone=LLaMA-3.1-8B-Instruct, Strategy=In-context Learning (ICL)2026.04 | 1.8 | 10.5 |