Temporal Question Answering on TempReason OBQA-L2
48EMDeepSeek-V3-AdapTime
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| DeepSeek-V3-AdapTimeBackbone=DeepSeek-V3-0324, Strategy=AdapTime2026.04 | 48 | 52.1 | |
| DeepSeek-V3-Step-backBackbone=DeepSeek-V3-0324, Strategy=Step-back2026.04 | 45.8 | 50.8 | |
| GPT-4Backbone=GPT-4, Strategy=Vanilla2026.04 | 45.4 | 52.5 | |
| DeepSeek-V3-ICLBackbone=DeepSeek-V3-0324, Strategy=In-context Learning (ICL)2026.04 | 45.1 | 50.8 | |
| DeepSeek-V3-CoTBackbone=DeepSeek-V3-0324, Strategy=Chain-of-Thought (CoT)2026.04 | 44.8 | 49.1 | |
| DeepSeek-V3-Self-refinementBackbone=DeepSeek-V3-0324, Strategy=Self-refinement2026.04 | 44.3 | 47.2 | |
| TG-LLMStrategy=Temporal Reasoning2026.04 | 42.4 | 52.2 | |
| REMEMO-largeBackbone=REMEMO-large, Strategy=Supervised Fine-tuning2026.04 | 37.4 | 54.9 | |
| REMEMO-baseBackbone=REMEMO-base, Strategy=Supervised Fine-tuning2026.04 | 33.6 | 51.6 | |
| T5-largeBackbone=T5-large, Strategy=Supervised Fine-tuning2026.04 | 32.7 | 50.9 | |
| Qwen-3-8B-AdapTimeBackbone=Qwen-3-8B, Strategy=AdapTime2026.04 | 29.1 | 37.9 | |
| T5-baseBackbone=T5-base, Strategy=Supervised Fine-tuning2026.04 | 26 | 45 | |
| Qwen-3-8B-ICLBackbone=Qwen-3-8B, Strategy=In-context Learning (ICL)2026.04 | 23.9 | 33.9 | |
| Qwen-3-8B-CoTBackbone=Qwen-3-8B, Strategy=Chain-of-Thought (CoT)2026.04 | 22.6 | 30.6 | |
| LLaMA-3-8B-AdapTimeBackbone=LLaMA-3.1-8B-Instruct, Strategy=AdapTime2026.04 | 18.7 | 26.8 | |
| LLaMA-3-8B-CoTBackbone=LLaMA-3.1-8B-Instruct, Strategy=Chain-of-Thought (CoT)2026.04 | 18.5 | 26.1 | |
| LLaMA-3-8B-ICLBackbone=LLaMA-3.1-8B-Instruct, Strategy=In-context Learning (ICL)2026.04 | 3.8 | 10 |