Question Answering on NQ (Natural Questions) in-domain (test)
48.8Exact MatchSearch-R1-7B-Base-PPO w/ Format
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Search-R1-7B-Base-PPO w/ FormatModel Scale=7B, Training Algorithm=PPO, Input Formatting=Included2025.10 | 48.8 | — | |
| EM+Info-ThinkFramework=VERITAS-R1, Reward Components=EM+Info-Think2025.10 | 48.6 | — | |
| EM+Info-Think+Think-AnsFramework=VERITAS-R1, Reward Components=EM+Info-Think+Think-Ans2025.10 | 48.4 | — | |
| EM+Think-AnsFramework=VERITAS-R1, Reward Components=EM+Think-Ans2025.10 | 48.2 | — | |
| Search-R1-7B-Base-PPOModel Scale=7B, Training Algorithm=PPO2025.10 | 48 | — | |
| DeSA-7B-Instrct-GRPOModel Scale=7B, Training Algorithm=GRPO2025.10 | 46.8 | — | |
| Search-R1Base LLM=Qwen2.5-7B-Instruct2026.01 | 42.11 | 55.96 | |
| ReSearch-7B-InstructModel Scale=7B, Model Variant=Instruct2025.10 | 41.5 | — | |
| D2PLAN-3BBase LLM=Qwen2.5-3B-Instruct2026.01 | 41.3 | 55.6 | |
| AutoRefineBase LLM=Qwen2.5-3B-Instruct2026.01 | 41.05 | 57.01 | |
| ZeroSearchBase LLM=Qwen2.5-7B-Instruct2026.01 | 41 | 54.02 | |
| R1-SearcherBase LLM=Qwen2.5-7B-Instruct2026.01 | 40 | 53.43 | |
| ReSearchBase LLM=Qwen2.5-7B-Instruct2026.01 | 39.81 | 53.85 | |
| D2PLAN-7BBase LLM=Qwen2.5-7B-Instruct2026.01 | 39.78 | 58.03 | |
| ReSearch-7B-BaseModel Scale=7B, Model Variant=Base2025.10 | 39.6 | — | |
| StepSearchBase LLM=Qwen2.5-7B-Instruct2026.01 | 38.64 | 53.07 | |
| ZeroSearchBase LLM=Qwen2.5-3B-Instruct2026.01 | 38.42 | 51.36 | |
| DeSA-3B-Instruct-GRPOModel Scale=3B, Training Algorithm=GRPO2025.10 | 37.5 | — | |
| Rejection SamplingReasoning Strategy=Rejection Sampling2025.10 | 36 | — | |
| RAGReasoning Strategy=Retrieval-Augmented Generation2025.10 | 34.9 | — | |
| StepSearchBase LLM=Qwen2.5-3B-Instruct2026.01 | 32.91 | 45.6 | |
| SFTTraining Algorithm=Supervised Fine-Tuning2025.10 | 31.8 | — | |
| R1-baseModel Variant=Base2025.10 | 29.7 | — | |
| Search-R1Base LLM=Qwen2.5-3B-Instruct2026.01 | 28.7 | 44.21 | |
| R1-instructModel Variant=Instruct2025.10 | 27 | — | |
| IRCoTReasoning Strategy=Iterative Retrieval Chain-of-Thought2025.10 | 22.4 | — | |
| Search-o1Reasoning Strategy=Search-o12025.10 | 15.1 | — | |
| Direct InferenceReasoning Strategy=Direct Inference2025.10 | 13.4 | — | |
| CoTReasoning Strategy=Chain-of-Thought2025.10 | 4.8 | — |