General Question Answering on TriviaQA out-of-domain (val test)
70.9EMSearch-R2
Evaluation Results
| Method | Links | |
|---|---|---|
| Search-R2Backbone=Qwen2.5-32B2026.02 | 70.9 | |
| Search-R1Backbone=Qwen2.5-32B2026.02 | 68 | |
| Search-R2Backbone=Qwen3-8B2026.02 | 67.6 | |
| Search-R2Backbone=Qwen2.5-7B2026.02 | 65.9 | |
| EM+Think-AnsFramework=VERITAS-R1, Reward Components=EM+Think-Ans2025.10 | 65.8 | |
| EM+Info-ThinkFramework=VERITAS-R1, Reward Components=EM+Info-Think2025.10 | 65 | |
| EM+Info-Think+Think-AnsFramework=VERITAS-R1, Reward Components=EM+Info-Think+Think-Ans2025.10 | 64.5 | |
| Search-R1-7B-Base-PPO w/ FormatModel Scale=7B, Training Algorithm=PPO, Input Formatting=Included2025.10 | 64.4 | |
| ReSearch-7B-InstructModel Scale=7B, Model Variant=Instruct2025.10 | 64 | |
| Search-R1-7B-Base-PPOModel Scale=7B, Training Algorithm=PPO2025.10 | 63.8 | |
| Search-R1Backbone=Qwen3-8B2026.02 | 63.1 | |
| DeSA-7B-Instrct-GRPOModel Scale=7B, Training Algorithm=GRPO2025.10 | 63.1 | |
| ReSearch-7B-BaseModel Scale=7B, Model Variant=Base2025.10 | 60.6 | |
| Rejection SamplingBackbone=Qwen2.5-7B2026.02 | 59.2 | |
| Rejection SamplingReasoning Strategy=Rejection Sampling2025.10 | 59.2 | |
| RAGBackbone=Qwen2.5-7B2026.02 | 58.5 | |
| RAGReasoning Strategy=Retrieval-Augmented Generation2025.10 | 58.5 | |
| DeSA-3B-Instruct-GRPOModel Scale=3B, Training Algorithm=GRPO2025.10 | 57.5 | |
| Search-R1Backbone=Qwen2.5-7B2026.02 | 56 | |
| R1-baseBackbone=Qwen2.5-7B2026.02 | 53.9 | |
| R1-baseModel Variant=Base2025.10 | 53.9 | |
| R1-instructBackbone=Qwen2.5-7B2026.02 | 53.7 | |
| R1-instructModel Variant=Instruct2025.10 | 53.7 | |
| IRCoTBackbone=Qwen2.5-7B2026.02 | 47.8 | |
| IRCoTReasoning Strategy=Iterative Retrieval Chain-of-Thought2025.10 | 47.8 | |
| Search-o1Backbone=Qwen2.5-7B2026.02 | 44.3 | |
| Search-o1Reasoning Strategy=Search-o12025.10 | 44.3 | |
| Direct InferenceBackbone=Qwen2.5-7B2026.02 | 40.8 | |
| Direct InferenceReasoning Strategy=Direct Inference2025.10 | 40.8 | |
| SFTBackbone=Qwen2.5-7B2026.02 | 35.4 | |
| SFTTraining Algorithm=Supervised Fine-Tuning2025.10 | 35.4 | |
| CoTBackbone=Qwen2.5-7B2026.02 | 18.5 | |
| CoTReasoning Strategy=Chain-of-Thought2025.10 | 18.5 |