Multi-hop Question Answering on HotpotQA (dev)
85.81Answer F1Dense-aware HKVM controller
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| Dense-aware HKVM controllervariant_type=Dense-aware controller, training_signals=out-of-fold HKVM2026.06 | 85.81 | — | 85.627 | — | — | 100 | — | |
| Learned HKVM calibrationvariant_type=Reference predictor, training_signals=out-of-fold HKVM2026.06 | 82.929 | — | 82.705 | — | — | 96.439 | — | |
| Controller HKVM-onlyvariant_type=Source-only controller control, training_signals=out-of-fold HKVM2026.06 | 82.929 | — | 82.705 | — | — | 100 | — | |
| BM252026.06 | 81.835 | — | 81.648 | — | — | 99.851 | — | |
| SEQGRAPHModel Size=Large2023.07 | 81.62 | 88.28 | 66.51 | 63.24 | — | — | — | |
| ETC-largeInput length=4096, Configuration=lifting from RoBERTa, #Params=558M2020.04 | 81.3 | 89.4 | — | — | — | — | — | |
| Contriever2026.06 | 79.93 | — | 79.73 | — | — | 100 | — | |
| Controller dense-onlyvariant_type=Source-only controller control, training_signals=out-of-fold HKVM2026.06 | 79.909 | — | 79.671 | — | — | 100 | — | |
| ColBERTv2variant_type=Reference predictor2026.06 | 79.844 | — | 79.635 | — | — | 100 | — | |
| ColBERTv22026.06 | 79.844 | — | 79.635 | — | — | 100 | — | |
| ETC-largeInput length=4096, #Params=539M2020.04 | 79.8 | 89 | — | — | — | — | — | |
| FIDModel Size=Large2023.07 | 79.39 | — | 65.59 | — | — | — | — | |
| PATH-FIDModel Size=Large, Implementation Source=Our reimplementation2023.07 | 79 | 86.88 | 65.33 | 61.52 | — | — | — | |
| PATH-FIDModel Size=Large, Implementation Source=Yavuz et al. (2022)2023.07 | 78.9 | 85.7 | 65.8 | 59.3 | — | — | — | |
| Longformer-largeInput length=4096, #Params=435M*2020.04 | 78.8 | 86.06 | — | — | — | — | — | |
| SEQGRAPHModel Size=Base2023.07 | 77.6 | 87.72 | 64.19 | 62.44 | — | — | — | |
| PATH-FIDModel Size=Base2023.07 | 75.69 | 86 | 62.03 | 60.45 | — | — | — | |
| FIDModel Size=Base2023.07 | 75.2 | — | 61.84 | — | — | — | — | |
| ETCInput length=4096, #Params=166M2020.04 | 75.1 | 86.9 | — | — | — | — | — | |
| ETCInput length=4096, Configuration=flat structure, #Params=166M2020.04 | 74.8 | 87 | — | — | — | — | — | |
| ETCInput length=4096, Configuration=no CPC, #Params=166M2020.04 | 74.7 | 86.6 | — | — | — | — | — | |
| LongformerInput length=4096, #Params=149M*2020.04 | 74.3 | 84.4 | — | — | — | — | — | |
| ETCInput length=4096, Configuration=no hard g2l, #Params=166M2020.04 | 74.3 | 86.4 | — | — | — | — | — | |
| Weighted-KG-PPR2026.06 | 73.695 | — | 73.369 | — | — | 95.381 | — | |
| KG-PPR2026.06 | 73.645 | — | 73.315 | — | — | 95.381 | — | |
| Static-HG2026.06 | 73.34 | — | 73.018 | — | — | 98.19 | — | |
| ETCInput length=4096, Configuration=shared, #Params=109M2020.04 | 73.3 | 86.6 | — | — | — | — | — | |
| Weighted-HG-KV2026.06 | 72.957 | — | 72.586 | — | — | 98.19 | — | |
| ETCInput length=4096, Configuration=flat structure, no CPC, no hard g2l, #Params=166M2020.04 | 72.2 | 85.7 | — | — | — | — | — | |
| Co-occurrence-HG2026.06 | 71.074 | — | 70.722 | — | — | 98.204 | — | |
| RNN-RetrievalModel size=BERT base, requires re-ranking=true2020.10 | 65.8 | — | 52.7 | — | — | — | — | |
| Transformer-XHModel size=BERT-base, requires re-ranking=true2020.10 | 62.4 | — | 50.2 | — | — | — | — | |
| Fusion Retriever + Cross-Block ReaderModel size=BERT-base, requires re-ranking=false2020.10 | 61.7 | — | 50.4 | — | — | — | — | |
| Semantic Retriever2020.10 | 58.8 | — | 46.5 | — | — | — | — | |
| ARCOBackbone=Qwen3-4B, Method Group=Process Reward Models (PRM)2026.06 | 55.07 | — | 42.8 | — | — | — | — | |
| multihopModel=GPT-3.5, Compiler=ensemble2023.10 | 54.7 | — | — | — | — | — | — | |
| R1-SearcherBackbone=Qwen3-4B, Method Group=Outcome Reward Models (ORM)2026.06 | 53.42 | — | 41 | — | — | — | — | |
| DocKIT2020.10 | 51.7 | — | 42.1 | — | — | — | — | |
| Search-R1Backbone=Qwen3-4B, Method Group=Outcome Reward Models (ORM)2026.06 | 51.67 | — | 39.4 | — | — | — | — | |
| AgentPRMBackbone=Qwen3-4B, Method Group=Process Reward Models (PRM)2026.06 | 51.49 | — | 38.6 | — | — | — | — | |
| RaRBackbone=Qwen3-4B, Method Group=Outcome Reward Models (ORM)2026.06 | 50.75 | — | 38 | — | — | — | — | |
| multihopModel=Llama2-13b-chat, Compiler=ensemble2023.10 | 50 | — | — | — | — | — | — | |
| Cognitive Graph2020.10 | 49.4 | — | 37.6 | — | — | — | — | |
| multihopModel=GPT-3.5, Compiler=bootstrap2023.10 | 48.7 | 47 | — | — | — | — | — | |
| CARMOBackbone=Qwen3-4B, Method Group=Outcome Reward Models (ORM)2026.06 | 48.52 | — | 34.8 | — | — | — | — | |
| Base ModelBackbone=Qwen3-4B, Prompting=zero-shot2026.06 | 45.93 | — | 34.4 | — | — | — | — | |
| ARCOBackbone=Llama-3.2-3B, Method Group=Process Reward Models (PRM)2026.06 | 45.58 | — | 36.4 | — | — | — | — | |
| Search-R1Backbone=Llama-3.2-3B, Method Group=Outcome Reward Models (ORM)2026.06 | 45.54 | — | 34.6 | — | — | — | — | |
| R1-SearcherBackbone=Llama-3.2-3B, Method Group=Outcome Reward Models (ORM)2026.06 | 44.48 | — | 35 | — | — | — | — | |
| RLERBackbone=Qwen3-4B, Method Group=Outcome Reward Models (ORM)2026.06 | 44.4 | — | 32.2 | — | — | — | — | |
| COT_RAGModel=GPT-3.5, Compiler=bootstrap2023.10 | 42.3 | 36 | — | — | — | — | — | |
| multihopModel=Llama2-13b-chat, Compiler=bootstrap2023.10 | 42 | 48.3 | — | — | — | — | — | |
| CARMOBackbone=Llama-3.2-3B, Method Group=Outcome Reward Models (ORM)2026.06 | 41.74 | — | 30.2 | — | — | — | — | |
| RLERBackbone=Llama-3.2-3B, Method Group=Outcome Reward Models (ORM)2026.06 | 41.31 | — | 28.8 | — | — | — | — | |
| RaRBackbone=Llama-3.2-3B, Method Group=Outcome Reward Models (ORM)2026.06 | 40.83 | — | 29.6 | — | — | — | — | |
| reactModel=Llama2-13b-chat, Compiler=bootstrap × 22023.10 | 40 | — | — | — | — | — | — | |
| reactModel=GPT-3.5, Compiler=bootstrap × 22023.10 | 39 | — | — | — | — | — | — | |
| COT_RAGModel=Llama2-13b-chat, Compiler=bootstrap2023.10 | 38.2 | 36 | — | — | — | — | — | |
| AgentPRMBackbone=Llama-3.2-3B, Method Group=Process Reward Models (PRM)2026.06 | 37.29 | — | 29.4 | — | — | — | — | |
| multihopModel=GPT-3.5, Compiler=fewshot2023.10 | 36.9 | 38.3 | — | — | — | — | — | |
| COT_RAGModel=GPT-3.5, Compiler=fewshot2023.10 | 36.4 | 36 | — | — | — | — | — | |
| multihopModel=Llama2-13b-chat, Compiler=fewshot2023.10 | 34.7 | 32 | — | — | — | — | — | |
| COT_RAGModel=Llama2-13b-chat, Compiler=fewshot2023.10 | 34.5 | 36 | — | — | — | — | — | |
| vanillaModel=GPT-3.5, Compiler=fewshot2023.10 | 34.3 | — | — | — | — | — | — | |
| reactModel=GPT-3.5, Compiler=+human_r2023.10 | 33 | — | — | — | — | — | — | |
| reactModel=GPT-3.5, Compiler=bootstrap2023.10 | 31 | — | — | — | — | — | — | |
| Base ModelBackbone=Llama-3.2-3B, Prompting=zero-shot2026.06 | 29.48 | — | 18.4 | — | — | — | — | |
| reactModel=Llama2-13b-chat, Compiler=+human_r2023.10 | 28.3 | — | — | — | — | — | — | |
| vanillaModel=Llama2-13b-chat, Compiler=fewshot2023.10 | 27.5 | — | — | — | — | — | — | |
| reactModel=Llama2-13b-chat, Compiler=bootstrap2023.10 | 24.7 | — | — | — | — | — | — | |
| reactModel=GPT-3.5, Compiler=none2023.10 | 20.3 | — | — | — | — | — | — | |
| reactModel=Llama2-13b-chat, Compiler=none2023.10 | 20 | — | — | — | — | — | — | |
| CoTRetrieval=no, LLM=code-davinci-002, Reasoning Chains=12023.04 | — | — | — | — | 39.8 | — | — | |
| CoT+SC@40Retrieval=no, LLM=code-davinci-002, Reasoning Chains=402023.04 | — | — | — | — | 44.6 | — | — | |
| DSPRetrieval=yes, LLM=text-davinci-002, Reasoning Chains=202023.04 | — | — | — | — | 62.9 | — | — | |
| GRPOProcess signal=OPD, Standardization operator=Masked-Norm, Actor model=Qwen2.5-7B-Instruct, Teacher model=Qwen2.5-32B-Instruct2026.06 | — | — | — | — | — | — | 62.03 | |
| GRPOProcess signal=G-OPD, Standardization operator=Masked-Norm, Actor model=Qwen2.5-7B-Instruct, Teacher model=Qwen2.5-32B-Instruct2026.06 | — | — | — | — | — | — | 61.96 | |
| IR-COTRetrieval=yes, LLM=code-davinci-002, Reasoning Chains=12023.04 | — | — | — | — | 61.2 | — | — | |
| MCRRetrieval=yes, LLM=code-davinci-002, Reasoning Chains=52023.04 | — | — | — | — | 57 | — | — | |
| MCR+SC@3Retrieval=yes, LLM=code-davinci-002, Reasoning Chains=152023.04 | — | — | — | — | 59.2 | — | — | |
| PASSProcess signal=OPD, Standardization operator=Masked-Norm, Actor model=Qwen2.5-7B-Instruct, Teacher model=Qwen2.5-32B-Instruct2026.06 | — | — | — | — | — | — | 64.14 | |
| PASSProcess signal=G-OPD, Standardization operator=Masked-Norm, Actor model=Qwen2.5-7B-Instruct, Teacher model=Qwen2.5-32B-Instruct2026.06 | — | — | — | — | — | — | 64.34 | |
| Self-Ask (ours)Retrieval=yes, LLM=code-davinci-002, Reasoning Chains=12023.04 | — | — | — | — | 50.2 | — | — |