Long-context Question Answering on Qasper
83.09F1CE-GOCD
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| CE-GOCDLLM=Qwen-plus2026.01 | 83.09 | 82.93 | 83.42 | |
| CE-GOCDLLM=DeepSeek-V32026.01 | 78.84 | 78.1 | 79.84 | |
| CE-GOCDLLM=GPT-42026.01 | 78.69 | 77.44 | 80.38 | |
| Embedding retrievalLLM=GPT-42026.01 | 75.42 | 75.08 | 76.07 | |
| LLM onlyLLM=Qwen-plus2026.01 | 75.39 | 75.59 | 75.53 | |
| LLM onlyLLM=GPT-42026.01 | 75.31 | 76.31 | 74.64 | |
| BM25LLM=GPT-42026.01 | 74.87 | 74.62 | 75.5 | |
| MindMapLLM=GPT-42026.01 | 73.8 | 74.3 | 73.59 | |
| KAGLLM=GPT-42026.01 | 71.17 | 66.29 | 77.03 | |
| LLM onlyLLM=DeepSeek-V32026.01 | 69.08 | 62.77 | 77.06 | |
| PathRAGLLM=GPT-42026.01 | 68.5 | 63.98 | 74.23 | |
| MistralTraining Protocol=Trained from scratch, Training Tokens=>1T, Input Context Length=16K2024.09 | 25.8 | — | — | |
| GSATraining Protocol=Finetuned from Mistral 7B, Training Tokens=20B, Input Context Length=16K, Base Model=Mistral 7B2024.09 | 18.8 | — | — | |
| GLATraining Protocol=Finetuned from Mistral 7B, Training Tokens=20B, Input Context Length=16K, Base Model=Mistral 7B2024.09 | 18.4 | — | — | |
| RetNetTraining Protocol=Finetuned from Mistral 7B, Training Tokens=20B, Input Context Length=16K, Base Model=Mistral 7B2024.09 | 11.1 | — | — | |
| RWKV6Training Protocol=Trained from scratch, Training Tokens=>1T, Input Context Length=16K2024.09 | 9.2 | — | — | |
| MambaTraining Protocol=Trained from scratch, Training Tokens=>1T, Input Context Length=16K2024.09 | 5.6 | — | — |