Single-Hop Question Answering on NQ, TriviaQA, and PopQA
48.5NQ ScoreSLEA-RL
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| SLEA-RLBase Model=Qwen2.5-7B-Instruct, Type=RL Methods, Training Context=Trained on NQ and HotpotQA2026.03 | 48.5 | 81.8 | 55.2 | |
| IGPOBase Model=Qwen2.5-7B-Instruct, Type=RL Methods2026.03 | 46.7 | 80.1 | 52.5 | |
| GiGPOBase Model=Qwen2.5-7B-Instruct, Type=RL Methods2026.03 | 46.4 | 64.7 | 46.1 | |
| SkillRLBase Model=Qwen2.5-7B-Instruct, Type=RL Methods2026.03 | 45.9 | 63.3 | 73.8 | |
| Search-R1++Base Model=Qwen2.5-3B, Method Category=RL trained Deep Research Agents2026.02 | 42.7 | 60.8 | 43.2 | |
| GSPOBase Model=Qwen2.5-7B-Instruct, Type=RL Methods2026.03 | 41.5 | 77.7 | 45.4 | |
| RLOOBase Model=Qwen2.5-7B-Instruct, Type=RL Methods2026.03 | 40.7 | 72.5 | 43.1 | |
| GRPOBase Model=Qwen2.5-7B-Instruct, Type=RL Methods2026.03 | 40.3 | 77 | 49.6 | |
| Search-R1Base Model=Qwen2.5-3B, Method Category=RL trained Deep Research Agents2026.02 | 39.6 | 57 | 38.1 | |
| PPOBase Model=Qwen2.5-7B-Instruct, Type=RL Methods2026.03 | 38.7 | 75.4 | 48.7 | |
| Reinforce++Base Model=Qwen2.5-7B-Instruct, Type=RL Methods2026.03 | 34.3 | 67.5 | 44.3 | |
| R1-baseBase Model=Qwen2.5-3B, Method Category=RL Trained LLM without Retrieval2026.02 | 22.6 | 45.5 | 17.3 | |
| ReActBase Model=Qwen2.5-3B, Method Category=Training free Deep Research Agent2026.02 | 6.3 | 13 | 7.2 |