Multi-hop Question Answering on Bamboogle
75.2AccuracyTC+FM*
Evaluation Results
| Method | Links | |
|---|---|---|
| TC+FM*Model=Qwen-72B2026.02 | 75.2 | |
| TC+FM*Model=Qwen-14B2026.02 | 70.4 | |
| HCAPO (Ours)Type=RL Training, Base Model=Qwen2.5-7B-Instruct2026.03 | 69 | |
| GiGPOType=RL Training, Base Model=Qwen2.5-7B-Instruct2026.03 | 68.9 | |
| TC+FM*Model=Qwen-32B2026.02 | 68.8 | |
| Search-o1Model=Qwen-72B2026.02 | 67.2 | |
| HCAPO (Ours)Type=RL Training, Base Model=Qwen2.5-3B-Instruct2026.03 | 64.5 | |
| GiGPOType=RL Training, Base Model=Qwen2.5-3B-Instruct2026.03 | 64.1 | |
| SkillOrchestra+Routing Strategy=Skill-based Routing, Orchestrator Configuration=Best/Multiple2026.02 | 63.2 | |
| Search-o1Model=Qwen-32B2026.02 | 60.8 | |
| Naive GenModel=Qwen-72B2026.02 | 60 | |
| Naive GenModel=Llama3.1-8B2026.02 | 60 | |
| Standard RAGModel=Qwen-72B2026.02 | 59.2 | |
| TC+FM*Model=Qwen-7B2026.02 | 58.4 | |
| TC+FM*Model=Llama3.1-8B2026.02 | 58.4 | |
| SkillOrchestraRouting Strategy=Skill-based Routing, Orchestrator Configuration=Single2026.02 | 58.4 | |
| Naive GenModel=Qwen-32B2026.02 | 54.4 | |
| Standard RAGModel=Qwen-32B2026.02 | 52.8 | |
| Router-R1Routing Strategy=RL-based Routing2026.02 | 51.2 | |
| FireAct (Multi-task + CoT)Method Category=Fine-tuning, Base Model=GPT-3.5, Training Data=Multi-task + CoT2023.10 | 50.4 | |
| RouterDCRouting Strategy=Heuristic & Discriminative Routing2026.02 | 50.4 | |
| Naive GenModel=Qwen-14B2026.02 | 48.8 | |
| Largest LLMRouting Strategy=Heuristic & Discriminative Routing2026.02 | 48 | |
| Prompt LLM+Routing Strategy=Heuristic & Discriminative Routing, Turn Protocol=multi turn2026.02 | 47.2 | |
| Search-o1Model=Llama3.1-8B2026.02 | 46.4 | |
| Standard RAGModel=Qwen-14B2026.02 | 44.8 | |
| Search-R1++Backbone=Qwen2.5-7B, Agent Type=RL trained Deep Research Agents2026.02 | 44.8 | |
| Prompt LLMRouting Strategy=Heuristic & Discriminative Routing2026.02 | 44.8 | |
| GraphRouterRouting Strategy=Heuristic & Discriminative Routing2026.02 | 44.8 | |
| FireAct (Fine-tuned on HotpotQA)Method Category=Fine-tuning, Base Model=GPT-3.5, Training Data=HotpotQA2023.10 | 44 | |
| Search-o1Model=Qwen-14B2026.02 | 43.2 | |
| Search-o1Model=Qwen-7B2026.02 | 43.2 | |
| FireAct (Multi-task)Method Category=Fine-tuning, Base Model=GPT-3.5, Training Data=Multi-task2023.10 | 43.2 | |
| FrugalGPTRouting Strategy=Heuristic & Discriminative Routing2026.02 | 43 | |
| Standard RAGModel=Qwen-7B2026.02 | 42.4 | |
| CoTMethod Category=Prompting, Base Model=GPT-3.52023.10 | 41.6 | |
| ReActMethod Category=Prompting, Base Model=GPT-3.52023.10 | 40.8 | |
| Search-R1Backbone=Qwen2.5-7B, Agent Type=RL trained Deep Research Agents2026.02 | 40.6 | |
| Standard RAGModel=Llama3.1-8B2026.02 | 39.2 | |
| TC+FM*Model=Qwen-3B2026.02 | 39.2 | |
| KNN Router+Routing Strategy=Heuristic & Discriminative Routing, Turn Protocol=multi turn2026.02 | 38.4 | |
| Search-R1Type=RL Training, Base Model=Qwen2.5-7B-Instruct2026.03 | 36.8 | |
| KNN RouterRouting Strategy=Heuristic & Discriminative Routing2026.02 | 36 | |
| MLP RouterRouting Strategy=Heuristic & Discriminative Routing2026.02 | 36 | |
| Search-o1Model=Qwen-3B2026.02 | 34.4 | |
| Naive GenModel=Qwen-7B2026.02 | 34.4 | |
| BERT RouterRouting Strategy=Heuristic & Discriminative Routing2026.02 | 31.2 | |
| R1-baseBackbone=Qwen2.5-7B, Agent Type=RL Trained LLM, Retrieval Setting=without Retrieval2026.02 | 29.6 | |
| R1-InstructType=RL Training, Base Model=Qwen2.5-3B-Instruct2026.03 | 29.3 | |
| ZeroSearchType=RL Training, Base Model=Qwen2.5-7B-Instruct2026.03 | 27.8 | |
| Search-R1Routing Strategy=No Routing2026.02 | 27.2 | |
| ReActBackbone=Qwen2.5-7B, Agent Type=Training free Deep Research Agent2026.02 | 26.6 | |
| Standard RAGModel=Qwen-3B2026.02 | 26.4 | |
| Search-R1Type=RL Training, Base Model=Qwen2.5-3B-Instruct2026.03 | 26.4 | |
| RAGRouting Strategy=No Routing2026.02 | 22.4 | |
| CoTRouting Strategy=No Routing2026.02 | 22.4 | |
| Naive GenModel=Qwen-3B2026.02 | 20.8 | |
| R1-InstructType=RL Training, Base Model=Qwen2.5-7B-Instruct2026.03 | 19.2 | |
| SFTRouting Strategy=No Routing2026.02 | 11.2 | |
| ZeroSearchType=RL Training, Base Model=Qwen2.5-3B-Instruct2026.03 | 11.1 | |
| IOMethod Category=Prompting, Base Model=GPT-3.52023.10 | 7.2 | |
| VanillaRouting Strategy=No Routing2026.02 | 4 |