Single-hop Question Answering on Natural Questions (NQ) (test)
47.5EMSE-Search-3B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SE-Search-3BRetrieval=Agent w/ Retrieval, Backbone=Qwen2.5-3B2026.02 | 47.5 | — | |
| AutoRefine-BaseRetrieval=Agent w/ Retrieval, Backbone=Qwen2.5-3B2026.02 | 46.7 | — | |
| O2-SearcherRetrieval=Agent w/ Retrieval, Backbone=Qwen2.5-3B2026.02 | 44.4 | — | |
| AutoRefine-InstructRetrieval=Agent w/ Retrieval, Backbone=Qwen2.5-3B2026.02 | 43.6 | — | |
| ZeroSearch-BaseRetrieval=Agent w/ Retrieval, Backbone=Qwen2.5-3B2026.02 | 43 | — | |
| ReSearch-BaseRetrieval=Agent w/ Retrieval, Backbone=Qwen2.5-3B2026.02 | 42.7 | — | |
| Search-R1-BaseRetrieval=Agent w/ Retrieval, Backbone=Qwen2.5-3B2026.02 | 42.1 | — | |
| InForageRetrieval=Agent w/ Retrieval, Backbone=Qwen2.5-3B2026.02 | 42.1 | — | |
| Search-R1-InstructRetrieval=Agent w/ Retrieval, Backbone=Qwen2.5-3B2026.02 | 39.7 | — | |
| ReSearch-InstructRetrieval=Agent w/ Retrieval, Backbone=Qwen2.5-3B2026.02 | 36.5 | — | |
| ProGraph-R1Base Model=Qwen2.5-7B-Instruct, Optimization Method=Training, Knowledge Interaction=Graph-based knowledge2026.01 | 35.94 | 50.71 | |
| Naive RAGRetrieval=Workflow w/ Retrieval, Backbone=Qwen2.5-3B2026.02 | 34.8 | — | |
| ProGraph-R1Base Model=Qwen2.5-3B-Instruct, Optimization Method=Training, Knowledge Interaction=Graph-based knowledge2026.01 | 34.38 | 47.34 | |
| Graph-R1Base Model=Qwen2.5-7B-Instruct, Optimization Method=Training, Knowledge Interaction=Graph-based knowledge2026.01 | 33.59 | 49.87 | |
| Search-R1Base Model=Qwen2.5-7B-Instruct, Optimization Method=Training, Knowledge Interaction=Chunk-based knowledge2026.01 | 32.03 | 45.88 | |
| R1-SearcherBase Model=Qwen2.5-7B-Instruct, Optimization Method=Training, Knowledge Interaction=Chunk-based knowledge2026.01 | 32.03 | 44.93 | |
| Graph-R1Base Model=Qwen2.5-3B-Instruct, Optimization Method=Training, Knowledge Interaction=Graph-based knowledge2026.01 | 30.47 | 44.75 | |
| SFTRetrieval=w/o Retrieval, Backbone=Qwen2.5-3B2026.02 | 24.9 | — | |
| Search-R1Base Model=Qwen2.5-3B-Instruct, Optimization Method=Training, Knowledge Interaction=Chunk-based knowledge2026.01 | 24.22 | 37.96 | |
| R1-SearcherBase Model=Qwen2.5-3B-Instruct, Optimization Method=Training, Knowledge Interaction=Chunk-based knowledge2026.01 | 24.22 | 36.53 | |
| Search-o1Retrieval=Agent w/ Retrieval, Backbone=Qwen2.5-3B2026.02 | 23.8 | — | |
| R1-BaseRetrieval=w/o Retrieval, Backbone=Qwen2.5-3B2026.02 | 22.6 | — | |
| R1-InstructRetrieval=w/o Retrieval, Backbone=Qwen2.5-3B2026.02 | 21 | — | |
| R1Base Model=Qwen2.5-7B-Instruct, Optimization Method=Training, Knowledge Interaction=No knowledge interaction2026.01 | 16.41 | 28.45 | |
| R1Base Model=Qwen2.5-3B-Instruct, Optimization Method=Training, Knowledge Interaction=No knowledge interaction2026.01 | 11.72 | 21.51 | |
| IRCoTRetrieval=Workflow w/ Retrieval, Backbone=Qwen2.5-3B2026.02 | 11.1 | — | |
| Direct GenerationRetrieval=w/o Retrieval, Backbone=Qwen2.5-3B2026.02 | 10.6 | — | |
| SFTBase Model=Qwen2.5-7B-Instruct, Optimization Method=Training, Knowledge Interaction=No knowledge interaction2026.01 | 5.12 | 19.02 | |
| SFTBase Model=Qwen2.5-3B-Instruct, Optimization Method=Training, Knowledge Interaction=No knowledge interaction2026.01 | 3.12 | 11.23 | |
| NaiveGenerationBase Model=Qwen2.5-3B-Instruct, Optimization Method=Prompt engineering, Knowledge Interaction=No knowledge interaction2026.01 | 2.34 | 8.9 | |
| NaiveGenerationBase Model=Qwen2.5-7B-Instruct, Optimization Method=Prompt engineering, Knowledge Interaction=No knowledge interaction2026.01 | 1.56 | 13 | |
| StandardRAGBase Model=Qwen2.5-7B-Instruct, Optimization Method=Prompt engineering, Knowledge Interaction=Chunk-based knowledge2026.01 | 1.56 | 15.97 | |
| StandardRAGBase Model=Qwen2.5-3B-Instruct, Optimization Method=Prompt engineering, Knowledge Interaction=Chunk-based knowledge2026.01 | 0 | 10.69 |