Multi-hop Question Answering on Multi-hop RAG
86.07F1IndexLM-4B
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| IndexLM-4Banswer model=Gemma3-27B-it, Avg. Tokens=19282025.12 | 86.07 | — | — | — | — | — | |
| IndexLM-1.7Banswer model=Qwen3-4B, Avg. Tokens=20432025.12 | 84.7 | — | — | — | — | — | |
| IndexLM-1.7Banswer model=Gemma3-27B-it, Avg. Tokens=20432025.12 | 84.21 | — | — | — | — | — | |
| IndexLM-0.6Banswer model=Gemma3-27B-it, Avg. Tokens=19662025.12 | 84 | — | — | — | — | — | |
| IndexLM-0.6Banswer model=Qwen3-4B, Avg. Tokens=19662025.12 | 83.31 | — | — | — | — | — | |
| IndexLM-4Banswer model=Qwen3-4B, Avg. Tokens=19282025.12 | 82.75 | — | — | — | — | — | |
| Chunk-Rerankanswer model=Qwen3-4B, Avg. Tokens=40942025.12 | 82.23 | — | — | — | — | — | |
| Chunk-Rerankanswer model=Gemma3-27B-it, Avg. Tokens=40942025.12 | 78.98 | — | — | — | — | — | |
| ReaderLM-v2answer model=Qwen3-4B, Avg. Tokens=40002025.12 | 78.42 | — | — | — | — | — | |
| Qwen3-4B + promptanswer model=Gemma3-27B-it, Avg. Tokens=25632025.12 | 77.85 | — | — | — | — | — | |
| Qwen3-4B + promptanswer model=Qwen3-4B, Avg. Tokens=25632025.12 | 76.88 | — | — | — | — | — | |
| ReaderLM-v2answer model=Gemma3-27B-it, Avg. Tokens=40002025.12 | 75.77 | — | — | — | — | — | |
| HtmlRAGanswer model=Gemma3-27B-it, Avg. Tokens=35622025.12 | 72.37 | — | — | — | — | — | |
| Firecrawl Extractanswer model=Gemma3-27B-it, Avg. Tokens=13192025.12 | 71.12 | — | — | — | — | — | |
| HtmlRAGanswer model=Qwen3-4B, Avg. Tokens=35622025.12 | 70.63 | — | — | — | — | — | |
| Firecrawl Extractanswer model=Qwen3-4B, Avg. Tokens=13192025.12 | 70.47 | — | — | — | — | — | |
| BELLEReasoning Strategy=Agent-based Reasoning2025.05 | 70.4 | 64.7 | 68.5 | — | — | — | |
| Markdown (raw)answer model=Gemma3-27B-it, Avg. Tokens=40962025.12 | 68.52 | — | — | — | — | — | |
| Markdown (raw)answer model=Qwen3-4B, Avg. Tokens=40962025.12 | 68.42 | — | — | — | — | — | |
| LiR3AGBackbone=8B, Reasoning=no-think2025.12 | 67.3 | 65.8 | — | — | — | — | |
| BeamAggRReasoning Strategy=Retrieval-augmented Reasoning2025.05 | 67.2 | 61.9 | 66.8 | — | — | — | |
| Classifier+LLM (Structured)Retrieval=BGE, Note=Input reordered for relevant text at end2025.12 | 67.2 | 62.5 | — | — | — | — | |
| Vanilla RAGBackbone=32B, Reasoning=no-think2025.12 | 66.8 | 65.8 | — | — | — | — | |
| Classifier+LLMRetrieval=BGE2025.12 | 66.6 | 62.5 | — | — | — | — | |
| Classifier+LLMRetrieval=MonoT52025.12 | 65.9 | 62.5 | — | — | — | — | |
| ClassifierRetrieval=BGE2025.12 | 65.5 | 61.9 | — | — | — | — | |
| Vanilla RAGBackbone=32B, Reasoning=think2025.12 | 65.4 | 64.2 | — | — | — | — | |
| ClassifierRetrieval=MonoT52025.12 | 65.3 | 62.1 | — | — | — | — | |
| Classifier+LLMRetrieval=Dense2025.12 | 64.8 | 59.7 | — | — | — | — | |
| ContextPilotHardware=16×H20, Annotations=enabled2025.11 | 64.68 | — | — | — | 17,498.75 | 60.37 | |
| ContextPilotHardware=32×H20, Annotations=enabled2025.11 | 64.68 | — | — | — | 33,072.64 | 58.41 | |
| BaselineRetrieval=BGE2025.12 | 64.5 | 60.1 | — | — | — | — | |
| ClassifierRetrieval=Dense2025.12 | 64.2 | 59.4 | — | — | — | — | |
| VanillaHardware=16×H202025.11 | 64.15 | — | — | — | 9,636.69 | 5.12 | |
| VanillaHardware=32×H202025.11 | 64.15 | — | — | — | 18,406.08 | 4.17 | |
| ContextPilot w/o AnnotationsHardware=16×H202025.11 | 64.09 | — | — | — | 17,498.75 | 60.37 | |
| ContextPilot w/o AnnotationsHardware=32×H202025.11 | 64.09 | — | — | — | 33,072.64 | 58.41 | |
| Vanilla RAGBackbone=14B, Reasoning=no-think2025.12 | 63.7 | 62.8 | — | — | — | — | |
| HTML (raw)answer model=Qwen3-4B, Avg. Tokens=40962025.12 | 63.29 | — | — | — | — | — | |
| BaselineRetrieval=MonoT52025.12 | 62.6 | 59.4 | — | — | — | — | |
| ProbTreeReasoning Strategy=Retrieval-augmented Reasoning2025.05 | 62.5 | 56.5 | 60.1 | — | — | — | |
| Direct OutputBackbone=32B, Reasoning=think2025.12 | 62.1 | 60.8 | — | — | — | — | |
| Vanilla RAGBackbone=14B, Reasoning=think2025.12 | 61.5 | 59.8 | — | — | — | — | |
| Vanilla RAGBackbone=8B, Reasoning=think2025.12 | 61.3 | 59.6 | — | — | — | — | |
| BaselineRetrieval=Dense2025.12 | 61.2 | 56.7 | — | — | — | — | |
| Vanilla RAGBackbone=8B, Reasoning=no-think2025.12 | 61.1 | 60.2 | — | — | — | — | |
| Direct OutputBackbone=32B, Reasoning=no-think2025.12 | 59.7 | 58 | — | — | — | — | |
| IRCOTReasoning Strategy=Retrieval-augmented Reasoning2025.05 | 59.2 | 55.1 | 58.4 | — | — | — | |
| FLAREReasoning Strategy=Retrieval-augmented Reasoning2025.05 | 58.7 | 54.9 | 59.2 | — | — | — | |
| Classifier+LLMRetrieval=BM252025.12 | 58.5 | 53 | — | — | — | — | |
| LONGAGENTReasoning Strategy=Agent-based Reasoning2025.05 | 56.8 | 53.6 | 57.4 | — | — | — | |
| Direct OutputBackbone=14B, Reasoning=think2025.12 | 56.1 | 54.6 | — | — | — | — | |
| Direct OutputBackbone=14B, Reasoning=no-think2025.12 | 55.4 | 55.2 | — | — | — | — | |
| EfficientRAGReasoning Strategy=Retrieval-augmented Reasoning2025.05 | 55.3 | 49.2 | 54.7 | — | — | — | |
| BaselineRetrieval=BM252025.12 | 55 | 49.9 | — | — | — | — | |
| ClassifierRetrieval=BM252025.12 | 54.9 | 50.1 | — | — | — | — | |
| Self-AskReasoning Strategy=Retrieval-augmented Reasoning2025.05 | 54.6 | 49.8 | 52.6 | — | — | — | |
| Direct OutputBackbone=8B, Reasoning=think2025.12 | 54.1 | 52.9 | — | — | — | — | |
| RopMuraReasoning Strategy=Agent-based Reasoning2025.05 | 53.7 | 52.6 | 49.2 | — | — | — | |
| GEARReasoning Strategy=Agent-based Reasoning2025.05 | 52.5 | 50.7 | 51.9 | — | — | — | |
| Single-stepReasoning Strategy=Retrieval-augmented Reasoning2025.05 | 52.3 | 47.2 | 51.3 | — | — | — | |
| HTML (raw)answer model=Gemma3-27B-it, Avg. Tokens=40962025.12 | 51.19 | — | — | — | — | — | |
| Direct OutputBackbone=8B, Reasoning=no-think2025.12 | 50.6 | 49.2 | — | — | — | — | |
| CoTReasoning Strategy=Closed-book Reasoning2025.05 | 50.5 | 43.6 | 49.7 | — | — | — | |
| SPReasoning Strategy=Closed-book Reasoning2025.05 | 47.5 | 39.4 | 44.3 | — | — | — | |
| Self-Correcting RAGCategory=Self-Correcting RAG2026.04 | 44.8 | 35.3 | — | — | — | — | |
| RAG + MMRCategory=Advanced Selection & Reranking2026.04 | 40.5 | 32.1 | — | — | — | — | |
| CRAGCategory=Iterative & Agentic RAG2026.04 | 40.2 | 31.5 | — | — | — | — | |
| FilcoCategory=Advanced Selection & Reranking2026.04 | 37.2 | 29.8 | — | — | — | — | |
| RECOMPCategory=Advanced Selection & Reranking2026.04 | 36.8 | 29.2 | — | — | — | — | |
| Self-RAGCategory=Iterative & Agentic RAG2026.04 | 36.2 | 28.1 | — | — | — | — | |
| IRCoTCategory=Iterative & Agentic RAG2026.04 | 35.2 | 27.5 | — | — | — | — | |
| LongLLMLinguaCategory=Advanced Selection & Reranking2026.04 | 33.5 | 26.2 | — | — | — | — | |
| RRRCategory=Standard RAG Baselines2026.04 | 32.8 | 25.1 | — | — | — | — | |
| HyDECategory=Standard RAG Baselines2026.04 | 32.1 | 24.5 | — | — | — | — | |
| NaiveCategory=Standard RAG Baselines2026.04 | 30.3 | 22.4 | — | — | — | — | |
| DRAGCategory=Iterative & Agentic RAG2026.04 | 30.2 | 29.3 | — | — | — | — | |
| ACEBackbone=LLaMA-3.1-8B-Instruct2026.01 | — | — | 57.9 | 10,653 | — | — | |
| IterDRAGBackbone=LLaMA-3.1-8B-Instruct2026.01 | — | — | 47 | 18,196 | — | — | |
| KGPRetrieval=BM25, Retrieval Units=Passage, Avg # Units=32026.01 | — | — | 62 | — | — | — | |
| KGPRetrieval=BGE, Retrieval Units=Passage, Avg # Units=32026.01 | — | — | 63.4 | — | — | — | |
| KiRAGRetrieval=BM25, Retrieval Units=Passage, Avg # Units=32026.01 | — | — | 26.5 | — | — | — | |
| KiRAGRetrieval=BGE, Retrieval Units=Passage, Avg # Units=32026.01 | — | — | 22.5 | — | — | — | |
| LightRAGRetrieval=BM25, Retrieval Units=Passage, Avg # Units=32026.01 | — | — | 26.65 | — | — | — | |
| LightRAGRetrieval=BGE, Retrieval Units=Passage, Avg # Units=32026.01 | — | — | 20.44 | — | — | — | |
| RAGBackbone=LLaMA-3.1-8B-Instruct2026.01 | — | — | 49.2 | 1,127 | — | — | |
| RankLlamaRetrieval=BM25, Retrieval Units=Passage, Avg # Units=52026.01 | — | — | 42.09 | — | — | — | |
| RankLlamaRetrieval=BGE, Retrieval Units=Passage, Avg # Units=52026.01 | — | — | 43.51 | — | — | — | |
| Retrieval OnlyRetrieval=BM25, Retrieval Units=Passage, Avg # Units=32026.01 | — | — | 37.2 | — | — | — | |
| Retrieval OnlyRetrieval=BM25, Retrieval Units=Sentence, Avg # Units=32026.01 | — | — | 61.6 | — | — | — | |
| Retrieval OnlyRetrieval=BGE, Retrieval Units=Passage, Avg # Units=32026.01 | — | — | 44.4 | — | — | — | |
| Retrieval OnlyRetrieval=BGE, Retrieval Units=Sentence, Avg # Units=32026.01 | — | — | 60.2 | — | — | — | |
| SentGraphRetrieval=BM25, Retrieval Units=Sentence, Avg # Units=2.572026.01 | — | — | 63.4 | — | — | — | |
| SentGraphRetrieval=BGE, Retrieval Units=Sentence, Avg # Units=2.72026.01 | — | — | 65.6 | — | — | — | |
| SetR-CoT & IRIRetrieval=BM25, Retrieval Units=Passage, Avg # Units=2.632026.01 | — | — | 44.13 | — | — | — | |
| SetR-CoT & IRIRetrieval=BGE, Retrieval Units=Passage, Avg # Units=2.912026.01 | — | — | 47.14 | — | — | — | |
| VanillaBackbone=LLaMA-3.1-8B-Instruct2026.01 | — | — | 36.1 | 333 | — | — |