Temporal Reasoning on TimeQA Easy 1.0 (test)
93.7EMGPT-4o-mini
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPT-4o-miniStrategy=NeSTR2025.12 | 93.7 | 96.4 | |
| Qwen3-14BStrategy=NeSTR2025.12 | 91.1 | 94.5 | |
| Qwen3-14BStrategy=TISER2025.12 | 90 | 94.3 | |
| Qwen3-8BStrategy=NeSTR2025.12 | 89.5 | 94.2 | |
| Qwen3-8BStrategy=TISER2025.12 | 88.8 | 93.4 | |
| Qwen2.5-7BStrategy=TISER2025.12 | 86.8 | 92.6 | |
| GPT-4o-miniStrategy=TISER2025.12 | 86.7 | 91.9 | |
| Qwen2.5-7BStrategy=NeSTR2025.12 | 85.1 | 90.2 | |
| Qwen3-8BStrategy=Vanilla2025.12 | 79.8 | 87.9 | |
| Qwen3-14BStrategy=Vanilla2025.12 | 78.1 | 87.5 | |
| GPT-4o-miniStrategy=Vanilla2025.12 | 73.4 | 81.7 | |
| TG-LLMStrategy=CoT2025.12 | 66.4 | 69.1 | |
| REMEMO-largeStrategy=Vanilla2025.12 | 63.7 | 72.3 | |
| T5-largeStrategy=Vanilla2025.12 | 63.1 | 71.6 | |
| Event-ALStrategy=Vanilla2025.12 | 63 | 73.8 | |
| FiDEvaluation Protocol=Fine-tune2023.05 | 60.5 | 67.9 | |
| FiDStrategy=Vanilla2025.12 | 60.5 | 67.9 | |
| BigBirdStrategy=Vanilla2025.12 | 51.2 | 71.6 | |
| QAaPBackbone=gpt-3.5-turbo, Evaluation Protocol=Few-shot2023.05 | 48.2 | 58.3 | |
| QAaPStrategy=Few-shot2025.12 | 46.3 | 54.4 | |
| ReActStrategy=Few-shot2025.12 | 45 | 55.1 | |
| ReActBackbone=gpt-3.5-turbo, Evaluation Protocol=Few-shot2023.05 | 34.3 | 41.5 | |
| CoTBackbone=gpt-3.5-turbo, Evaluation Protocol=Few-shot2023.05 | 24.6 | 34.2 | |
| Qwen2.5-7BStrategy=Vanilla2025.12 | 13 | 13 |