Long-context Question Answering on HotpotQA
65.49Mean ScoreFull Model
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Full ModelBackbone=Qwen3-VL-8B-Instruct, Pruning Ratio=0%2026.02 | 65.49 | — | — | — | — | |
| WandaBackbone=Qwen3-VL-8B-Instruct, Pruning Ratio=30%2026.02 | 63.19 | — | — | — | — | |
| POPBackbone=Qwen3-VL-8B-Instruct, Pruning Ratio=33.3%2026.02 | 63.13 | — | — | — | — | |
| Full ModelBackbone=Gemma-3-12B-It, Pruning Ratio=0%2026.02 | 59.62 | — | — | — | — | |
| Llama-3.3-70B-InstructParameters=70B2026.01 | 59.6 | — | — | — | — | |
| POPBackbone=Gemma-3-12B-It, Pruning Ratio=33.3%2026.02 | 59.11 | — | — | — | — | |
| WandaBackbone=Gemma-3-12B-It, Pruning Ratio=30%2026.02 | 58.78 | — | — | — | — | |
| Foundation-Sec-8B-InstructParameters=8B2026.01 | 58.4 | — | — | — | — | |
| Full ModelBackbone=Llama-3.1-8B-Instruct, Pruning Ratio=0%2026.02 | 55.66 | — | — | — | — | |
| Foundation-Sec-8B-ReasoningParameters=8B2026.01 | 54.8 | — | — | — | — | |
| Llama-3.1-8B-InstructParameters=8B2026.01 | 54.1 | — | — | — | — | |
| POPBackbone=Llama-3.1-8B-Instruct, Pruning Ratio=31.25%2026.02 | 53.48 | — | — | — | — | |
| WandaBackbone=Llama-3.1-8B-Instruct, Pruning Ratio=30%2026.02 | 53.03 | — | — | — | — | |
| Llama-Primus-Merged2026.01 | 48.3 | — | — | — | — | |
| SliceGPTBackbone=Qwen3-VL-8B-Instruct, Pruning Ratio=25%2026.02 | 38.33 | — | — | — | — | |
| Phi-42026.01 | 34.7 | — | — | — | — | |
| ShortGPTBackbone=Qwen3-VL-8B-Instruct, Pruning Ratio=25%2026.02 | 16.37 | — | — | — | — | |
| SliceGPTBackbone=Llama-3.1-8B-Instruct, Pruning Ratio=25%2026.02 | 8.71 | — | — | — | — | |
| SliceGPTBackbone=Gemma-3-12B-It, Pruning Ratio=25%2026.02 | 4.18 | — | — | — | — | |
| ShortGPTBackbone=Llama-3.1-8B-Instruct, Pruning Ratio=25%2026.02 | 3.81 | — | — | — | — | |
| ShortGPTBackbone=Gemma-3-12B-It, Pruning Ratio=25%2026.02 | 0.34 | — | — | — | — | |
| Claude-3-sonnetCitation Strategy=LAC-S2024.09 | — | 81.3 | 75.3 | 108 | — | |
| GLM-4Citation Strategy=LAC-S2024.09 | — | 76.3 | 76.5 | 100 | — | |
| GLM-4-9B-chatCitation Strategy=LAC-S2024.09 | — | 68.5 | 71.5 | 96 | — | |
| GLM-4.1V-9B-ThinkingRAG Strategy=Direct (None)2026.02 | — | — | — | — | 39.49 | |
| GLM-4.1V-9B-Thinking ColPali RAGRAG Strategy=ColPali RAG2026.02 | — | — | — | — | 43.89 | |
| GLM-4.1V-9B-Thinking Embedding RAGRAG Strategy=Embedding RAG2026.02 | — | — | — | — | 36.03 | |
| GLM-4.1V-9B-Thinking Random RAGRAG Strategy=Random RAG2026.02 | — | — | — | — | 38 | |
| GLM-4.1V-9B-Thinking VERARAG Strategy=Attention-Guided RAG2026.02 | — | — | — | — | 39.56 | |
| GlyphRAG Strategy=Direct (None)2026.02 | — | — | — | — | 38.74 | |
| GPT-4oCitation Strategy=LAC-S2024.09 | — | 74.5 | 80.8 | 92 | — | |
| Llama-3.1-70B-InstructCitation Strategy=LAC-S2024.09 | — | 71.3 | 75.3 | 95 | — | |
| Llama-3.1-8B-InstructCitation Strategy=LAC-S2024.09 | — | 64 | 64.5 | 99 | — | |
| LongCite-8BCitation Strategy=LAC-S2024.09 | — | 70.8 | 69 | 103 | — | |
| LongCite-9BCitation Strategy=LAC-S2024.09 | — | 71.8 | 67.5 | 106 | — | |
| Mistral-Large-InstructCitation Strategy=LAC-S2024.09 | — | 77 | 77.3 | 100 | — | |
| Qwen3-VL-8B-InstructRAG Strategy=Direct (None)2026.02 | — | — | — | — | 28.48 | |
| Qwen3-VL-8B-Instruct OCR RAGRAG Strategy=OCR RAG2026.02 | — | — | — | — | 29.03 | |
| Qwen3-VL-8B-Instruct Random RAGRAG Strategy=Random RAG2026.02 | — | — | — | — | 30.16 | |
| Qwen3-VL-8B-Instruct VERARAG Strategy=Attention-Guided RAG2026.02 | — | — | — | — | 32.32 |