Factual Knowledge Evaluation on DyKnow 130 time-sensitive facts Wikidata-derived
80CorrectnessGPT-4
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GPT-4Year=2023, Prompting Prefix=Answer with the name only, Evaluation Protocol=Upper Bound2025.01 | 80 | 13 | 7 | |
| Llama-3 InstructYear=2024, Model Variant=Instruct, Prompting Prefix=Answer with the name only, Evaluation Protocol=Upper Bound2025.01 | 76 | 14 | 10 | |
| Mixtral InstructYear=2023, Model Variant=Instruct, Prompting Prefix=Answer with the name only, Evaluation Protocol=Upper Bound2025.01 | 62 | 29 | 9 | |
| Llama-3Year=2024, Evaluation Protocol=Upper Bound2025.01 | 57 | 36 | 7 | |
| ChatGPTYear=2022, Prompting Prefix=Answer with the name only, Evaluation Protocol=Upper Bound2025.01 | 57 | 35 | 8 | |
| MistralYear=2023, Evaluation Protocol=Upper Bound2025.01 | 53 | 39 | 8 | |
| VicunaYear=2023, Prompting Prefix=Answer with the name only, Evaluation Protocol=Upper Bound2025.01 | 52 | 33 | 15 | |
| Mistral InstructYear=2023, Model Variant=Instruct, Prompting Prefix=Answer with the name only, Evaluation Protocol=Upper Bound2025.01 | 52 | 32 | 16 | |
| Llama-2Year=2023, Evaluation Protocol=Upper Bound2025.01 | 51 | 42 | 7 | |
| Llama-2 ChatYear=2023, Model Variant=Chat, Prompting Prefix=Answer with the name only, Evaluation Protocol=Upper Bound2025.01 | 51 | 37 | 12 | |
| MixtralYear=2023, Evaluation Protocol=Upper Bound2025.01 | 48 | 42 | 10 | |
| Falcon InstructYear=2023, Model Variant=Instruct, Prompting Prefix=Answer with the name only, Evaluation Protocol=Upper Bound2025.01 | 44 | 41 | 15 | |
| GPT-3Year=2020, Evaluation Protocol=Upper Bound2025.01 | 42 | 47 | 12 | |
| FalconYear=2023, Evaluation Protocol=Upper Bound2025.01 | 42 | 47 | 11 | |
| OpenELM 3BYear=2024, Evaluation Protocol=Upper Bound2025.01 | 42 | 42 | 16 | |
| GPT-JYear=2021, Evaluation Protocol=Upper Bound2025.01 | 41 | 46 | 13 | |
| OLMo 1BYear=2024, Evaluation Protocol=Upper Bound2025.01 | 37 | 40 | 23 | |
| BloomYear=2022, Evaluation Protocol=Upper Bound2025.01 | 35 | 49 | 16 | |
| OLMo 7BYear=2024, Evaluation Protocol=Upper Bound2025.01 | 35 | 36 | 29 | |
| OpenELM 1.1BYear=2024, Evaluation Protocol=Upper Bound2025.01 | 35 | 47 | 18 | |
| GPT-2Year=2019, Evaluation Protocol=Upper Bound2025.01 | 26 | 42 | 32 | |
| Flan-T5Year=2022, Evaluation Protocol=Upper Bound2025.01 | 18 | 39 | 43 | |
| OpenELM 270MYear=2024, Evaluation Protocol=Upper Bound2025.01 | 12 | 28 | 61 | |
| T5Year=2020, Evaluation Protocol=Upper Bound2025.01 | 11 | 21 | 68 |