Long-context Question Answering on NarrativeQA (EM)
61.7Exact MatchQwen2.5-OpAmp-72B
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen2.5-OpAmp-72BParameters=72B, Adaptation=OpAmp2025.02 | 61.7 | |
| Llama3.3-70B-instParameters=70B, Type=Instruction-tuned2025.02 | 61.5 | |
| GPT-4o-0806Version=08062025.02 | 61.5 | |
| DeepSeek-V3Version=V32025.02 | 60.5 | |
| Qwen2.5-72B-instParameters=72B, Type=Instruction-tuned2025.02 | 60.2 | |
| Llama3-ChatQA2-70BParameters=70B, Version=ChatQA22025.02 | 59.8 | |
| Llama3.1-OpAmp-8BParameters=8B2025.02 | 57.4 | |
| Llama3.1-8B-instParameters=8B2025.02 | 55.9 | |
| Llama3-ChatQA2-8BParameters=8B2025.02 | 53.1 | |
| Qwen2.5-7B-instParameters=7B2025.02 | 47.7 | |
| Mistral-7B-inst-v0.3Parameters=7B2025.02 | 44.7 |