Temporal Reasoning Prediction on ICEWS18 (test)
75.78Positive AccuracyGETER
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| GETERBackbone=Llama3-8B-Instruct2025.05 | 75.78 | 74.09 | 87.53 | 78.85 | |
| GETERBackbone=Mistral-7B-Instruct2025.05 | 75.61 | 75.94 | 87.51 | 79.4 | |
| GETERBackbone=Qwen2.5-7B-Instruct2025.05 | 74.77 | 74.41 | 86.79 | 78.37 | |
| LoRABackbone=Qwen2.5-7B-Instruct2025.05 | 69.68 | 60.54 | 63.21 | 64.48 | |
| LoRABackbone=Mistral-7B-Instruct2025.05 | 64.22 | 64.63 | 76.63 | 68.2 | |
| LoRABackbone=Llama3-8B-Instruct2025.05 | 62.3 | 46.24 | 66.46 | 58.23 | |
| GPT-4oType=zero-shot2025.05 | 60.33 | 22.78 | 40.72 | 42.08 | |
| LoRABackbone=Mistral-7B-Instruct, Reasoning Chain=w/o chains text2025.05 | 58.07 | 55.27 | 74.46 | 62.21 | |
| LoRABackbone=Llama3-8B-Instruct, Reasoning Chain=w/o chains text2025.05 | 57.47 | 47.14 | 56.3 | 53.66 | |
| Llama3-8B-InstructType=zero-shot2025.05 | 55.12 | 18.81 | 9.14 | 28.79 | |
| GPT-4oType=zero-shot w/o chains text2025.05 | 51.64 | 36.61 | 23.79 | 38.32 | |
| LoRABackbone=Qwen2.5-7B-Instruct, Reasoning Chain=w/o chains text2025.05 | 45.82 | 59.83 | 66.27 | 56.82 | |
| Qwen2.5-7B-InstructType=zero-shot2025.05 | 44.22 | 48.67 | 10.92 | 35.4 | |
| Qwen2.5-7B-InstructType=zero-shot w/o chains text2025.05 | 30.94 | 40.53 | 25.13 | 32.34 | |
| Llama3-8B-InstructType=zero-shot w/o chains text2025.05 | 7.68 | 24.39 | 38.95 | 22.93 | |
| Mistral-7B-InstructType=zero-shot2025.05 | 4.14 | 33.06 | 41.58 | 25.37 | |
| Mistral-7B-InstructType=zero-shot w/o chains text2025.05 | 1.06 | 34.23 | 47.64 | 26.53 |