Multi-hop QA on Bamboogle (Accuracy %)
74.9Accuracy (%)IGPO
Evaluation Results
| Method | Links | |
|---|---|---|
| IGPOBase Model=Qwen2.5-7B-Instruct, Method Category=RL Methods2026.03 | 74.9 | |
| ProCeedRLBackbone=Qwen3-8B2026.04 | 73.87 | |
| DeepResearcher2026.04 | 72.8 | |
| DAPO/Search-R1Backbone=Qwen3-8B2026.04 | 70.83 | |
| GiGPOBase Model=Qwen2.5-7B-Instruct, Method Category=RL Methods2026.03 | 68.9 | |
| GiGPO2026.04 | 68.9 | |
| Rewinding MoreBackbone=Qwen3-8B, Type=ablation2026.04 | 65.87 | |
| RFTBackbone=Qwen3-8B2026.04 | 64.27 | |
| Qwen3-8B-v3-SFTBackbone=Qwen3-8B2026.04 | 62.13 | |
| ReAct PromptingBackbone=Qwen3-8B2026.04 | 50.93 | |
| EvolveRBase Model=Qwen2.5-7B-Instruct, Method Category=Prompting2026.03 | 42 | |
| SkillRLBase Model=Qwen2.5-7B-Instruct, Method Category=RL Methods2026.03 | 40.3 | |
| Search-R1Base Model=Qwen2.5-7B-Instruct, Method Category=Prompting2026.03 | 40.1 | |
| ZeroSearchBase Model=Qwen2.5-7B-Instruct, Method Category=Prompting2026.03 | 35.2 | |
| SLEA-RLBase Model=Qwen2.5-7B-Instruct, Method Category=RL Methods2026.03 | 33.2 | |
| R1-InstructBase Model=Qwen2.5-7B-Instruct, Method Category=Prompting2026.03 | 27.5 | |
| Search-o1Base Model=Qwen2.5-7B-Instruct, Method Category=Prompting2026.03 | 27 | |
| PPOBase Model=Qwen2.5-7B-Instruct, Method Category=RL Methods2026.03 | 26.2 | |
| GSPOBase Model=Qwen2.5-7B-Instruct, Method Category=RL Methods2026.03 | 25.4 | |
| GRPOBase Model=Qwen2.5-7B-Instruct, Method Category=RL Methods2026.03 | 25 | |
| RLOOBase Model=Qwen2.5-7B-Instruct, Method Category=RL Methods2026.03 | 24.8 | |
| Reinforce++Base Model=Qwen2.5-7B-Instruct, Method Category=RL Methods2026.03 | 23.7 | |
| RAGBase Model=Qwen2.5-7B-Instruct, Method Category=Prompting2026.03 | 23.2 | |
| CoTBase Model=Qwen2.5-7B-Instruct, Method Category=Prompting2026.03 | 22.6 | |
| Qwen2.5Base Model=Qwen2.5-7B-Instruct, Method Category=Prompting2026.03 | 22.2 |