Multi-hop Question Answering on Multi-Hop QA (Aggregated)
89.342Wiki AccuracyGemini-2.5-Flash-Thinking
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Gemini-2.5-Flash-ThinkingBackbone=Gemini-2.5-Flash-Thinking, Category=Frontier Models2026.01 | 89.34 | 81 | 60 | 76.78 | |
| GPT5-NanoBackbone=GPT5-Nano, Category=Frontier Models2026.01 | 89 | 82 | 61.5 | 77.5 | |
| GPT-OSS-120BBackbone=GPT-OSS-120B, Category=Frontier Models2026.01 | 89 | 82 | 66 | 79 | |
| GPT-OSS-20BBackbone=GPT-OSS-20B, Category=Frontier Models2026.01 | 88.5 | 79 | 49.25 | 72.25 | |
| LONGPASBackbone=Qwen3-4B-Instruct, Category=Instruct Models, Training Strategy=LONGPAS2026.01 | 86.5 | 75.38 | 55 | 72.29 | |
| LONGPASBackbone=Qwen3-30B-A3B-Thinking, Category=Reasoning Models, Training Strategy=LONGPAS2026.01 | 86.5 | 77 | 66 | 76.5 | |
| RLVRBackbone=Qwen3-30B-A3B-Instruct, Category=Instruct Models, Training Strategy=RLVR2026.01 | 86 | 76 | 57 | 73 | |
| LONGPASBackbone=Qwen3-4B-Thinking, Category=Reasoning Models, Training Strategy=LONGPAS2026.01 | 85.62 | 78.62 | 54.25 | 72.83 | |
| Qwen3-4B-ThinkingBackbone=Qwen3-4B-Thinking, Category=Reasoning Models, Training Strategy=Vanilla2026.01 | 85.5 | 75.12 | 50.25 | 70.29 | |
| LONGPASBackbone=Qwen3-30B-A3B-Instruct, Category=Instruct Models, Training Strategy=LONGPAS2026.01 | 84.5 | 77.5 | 55.5 | 72.5 | |
| Qwen3-30B-A3B-ThinkingBackbone=Qwen3-30B-A3B-Thinking, Category=Reasoning Models, Training Strategy=Vanilla2026.01 | 84.31 | 76.94 | 63.44 | 74.9 | |
| RLVRBackbone=Qwen3-30B-A3B-Thinking, Category=Reasoning Models, Training Strategy=RLVR2026.01 | 84 | 77 | 64 | 75 | |
| Qwen3-30B-A3B-InstructBackbone=Qwen3-30B-A3B-Instruct, Category=Instruct Models, Training Strategy=Vanilla2026.01 | 83.88 | 77.62 | 51.5 | 71 | |
| RLVRBackbone=Qwen3-4B-Thinking, Category=Reasoning Models, Training Strategy=RLVR2026.01 | 83.5 | 71.5 | 53.5 | 69.5 | |
| RLVRBackbone=Qwen3-4B-Instruct, Category=Instruct Models, Training Strategy=RLVR2026.01 | 82.62 | 75.88 | 49.25 | 69.25 | |
| LONGPASBackbone=Qwen2.5-7B-Instruct, Category=Instruct Models, Training Strategy=LONGPAS2026.01 | 80.88 | 80 | 51.62 | 70.83 | |
| LONGPASBackbone=LLaMA3.1-8B-Instruct, Category=Instruct Models, Training Strategy=LONGPAS2026.01 | 79.38 | 76.38 | 52 | 69.25 | |
| RLVRBackbone=Qwen2.5-7B-Instruct, Category=Instruct Models, Training Strategy=RLVR2026.01 | 68.25 | 67.75 | 40.25 | 58.75 | |
| RLVRBackbone=LLaMA3.1-8B-Instruct, Category=Instruct Models, Training Strategy=RLVR2026.01 | 68 | 69.75 | 46.62 | 61.46 | |
| Qwen3-4B-InstructBackbone=Qwen3-4B-Instruct, Category=Instruct Models, Training Strategy=Vanilla2026.01 | 63.25 | 67.62 | 28.75 | 53.21 | |
| AtomicRAG2026.02 | 56.8 | 70.5 | 50.9 | 59.4 | |
| Qwen2.5-7B-InstructBackbone=Qwen2.5-7B-Instruct, Category=Instruct Models, Training Strategy=Vanilla2026.01 | 54.75 | 69.75 | 32.38 | 52.29 | |
| StructRAG2026.02 | 52.3 | 49.2 | 24.9 | 42.1 | |
| RAPTOR2026.02 | 50.1 | 66.4 | 43.6 | 53.4 | |
| Fast-GraphRAG2026.02 | 47.8 | 62.3 | 42.1 | 50.7 | |
| HippoRAG22026.02 | 47.5 | 67.5 | 41.5 | 52.2 | |
| LightRAG2026.02 | 44.2 | 57 | 23.5 | 41.6 | |
| LLaMA3.1-8B-InstructBackbone=LLaMA3.1-8B-Instruct, Category=Instruct Models, Training Strategy=Vanilla2026.01 | 43.62 | 61.75 | 21 | 42.12 | |
| GFM-RAG2026.02 | 43.6 | 68.6 | 39.9 | 50.7 | |
| Lazy-GraphRAG2026.02 | 40.4 | 54.6 | 39.1 | 44.7 | |
| MS-GraphRAGscope=global2026.02 | 39.5 | 37.6 | 35.4 | 37.5 | |
| KGP2026.02 | 38.6 | 52.3 | 37.2 | 42.7 | |
| MS-GraphRAGscope=local2026.02 | 37.4 | 52.2 | 37.8 | 42.5 | |
| RAGreranking=with2026.02 | 32.1 | 54.5 | 33.3 | 40 | |
| RAGreranking=without2026.02 | 31.8 | 54 | 32.2 | 39.3 | |
| HippoRAG2026.02 | 29.8 | 38.1 | 25.7 | 31.2 | |
| KET-RAG2026.02 | 24.5 | 45.6 | 22.6 | 30.9 |