Question Answering on NQ (Exact Match)
74.2Exact MatchFreshPER
Evaluation Results
| Method | Links | |
|---|---|---|
| FreshPERType=LLM2026.04 | 74.2 | |
| Neuro-RITRanking Strategy=Reranker2026.04 | 72.57 | |
| PA-RAGRanking Strategy=Reranker2026.04 | 72.01 | |
| Neuro-RIT + RerankerGenerator=Mistral-7B-Instruct-v0.22026.04 | 69.79 | |
| Neuro-RITRanking Strategy=None2026.04 | 69.68 | |
| RALM + RerankerGenerator=Mistral-7B-Instruct-v0.22026.04 | 68.48 | |
| PA-RAGRanking Strategy=None2026.04 | 67.92 | |
| Neuro-RITGenerator=Mistral-7B-Instruct-v0.22026.04 | 66.47 | |
| PA-RAGRanking Strategy=RankCoT2026.04 | 66.3 | |
| RALMGenerator=Mistral-7B-Instruct-v0.22026.04 | 64.85 | |
| InstructRAGRanking Strategy=Reranker2026.04 | 63.58 | |
| Neuro-RIT + RankCoTGenerator=Mistral-7B-Instruct-v0.22026.04 | 63.48 | |
| RALMRanking Strategy=None2026.04 | 62.21 | |
| Neuro-RITRanking Strategy=RankCoT2026.04 | 61.93 | |
| RALMRanking Strategy=Reranker2026.04 | 61.65 | |
| InstructRAGRanking Strategy=None2026.04 | 60.62 | |
| RALM + RankCoTGenerator=Mistral-7B-Instruct-v0.22026.04 | 60.31 | |
| InstructRAGRanking Strategy=RankCoT2026.04 | 60.17 | |
| RALMRanking Strategy=RankCoT2026.04 | 59.43 | |
| RetRobustRanking Strategy=Reranker2026.04 | 54.42 | |
| RetRobustRanking Strategy=RankCoT2026.04 | 53.26 | |
| RetRobustRanking Strategy=None2026.04 | 53.23 | |
| On-PolicyType=LLM2026.04 | 50.8 | |
| SLATEBackbone=Qwen2.5-7B-Base, training_data=NQ+HotpotQA2026.02 | 49.7 | |
| SEARCH-R1Backbone=Qwen2.5-7B-Base, training_data=NQ+HotpotQA2026.02 | 48 | |
| Search-E1-InstructRetrieval Strategy=Multi-Hop Retrieval, Supervision Paradigm=Process-supervision RL, Base Model=Qwen2.5-3B2026.05 | 47.4 | |
| AutoRefine-BaseRetrieval Strategy=Multi-Hop Retrieval, Supervision Paradigm=Outcome-reward RL, Base Model=Qwen2.5-3B2026.05 | 46.7 | |
| EvolveRBackbone Model=Qwen2.5-7B2025.10 | 46.2 | |
| PiCABackbone=Qwen2.5-7B Instruct2026.05 | 46 | |
| ZeroSearchBackbone=Qwen2.5-7B Instruct2026.05 | 43.6 | |
| AutoRefine-InstructRetrieval Strategy=Multi-Hop Retrieval, Supervision Paradigm=Outcome-reward RL, Base Model=Qwen2.5-3B2026.05 | 43.6 | |
| TIPSBackbone=Qwen2.5-3B Instruct2026.05 | 43.5 | |
| EvolveRBackbone Model=Qwen2.5-3B2025.10 | 43.4 | |
| TIPSBackbone=Qwen2.5-7B Instruct2026.05 | 43.4 | |
| ReSearch-BaseRetrieval Strategy=Multi-Hop Retrieval, Supervision Paradigm=Outcome-reward RL, Base Model=Qwen2.5-3B2026.05 | 42.7 | |
| PiCABackbone=Qwen2.5-3B Instruct2026.05 | 42.6 | |
| SLATEBackbone=Qwen2.5-3B-Base, Training Data=NQ+HotpotQA2026.02 | 42.5 | |
| MT-PPOBackbone=Qwen2.5-7B Instruct2026.05 | 42.4 | |
| Search-R1-BaseRetrieval Strategy=Multi-Hop Retrieval, Supervision Paradigm=Outcome-reward RL, Base Model=Qwen2.5-3B2026.05 | 42.1 | |
| GiGPO-InstructRetrieval Strategy=Multi-Hop Retrieval, Supervision Paradigm=Process-supervision RL, Base Model=Qwen2.5-3B2026.05 | 42 | |
| ZeroSearchBackbone=Qwen2.5-3B Instruct2026.05 | 41.4 | |
| ZeroSearchBackbone=Qwen2.5-7B-Base2026.02 | 41.1 | |
| SEARCH-R1Backbone=Qwen2.5-3B-Base, Training Data=NQ+HotpotQA2026.02 | 40.6 | |
| Search-R1-baseBackbone Model=Qwen2.5-3B2025.10 | 40.6 | |
| ReSearchBackbone=Qwen2.5-7B-Base2026.02 | 39.8 | |
| MT-PPOBackbone=Qwen2.5-3B Instruct2026.05 | 39.7 | |
| Search-R1-InstructRetrieval Strategy=Multi-Hop Retrieval, Supervision Paradigm=Outcome-reward RL, Base Model=Qwen2.5-3B2026.05 | 39.7 | |
| Search-R1-instructBackbone Model=Qwen2.5-7B2025.10 | 39.3 | |
| Search-R1Backbone=Qwen2.5-7B Instruct2026.05 | 39.3 | |
| RAGBackbone=Qwen2.5-7B-Base2026.02 | 37.1 | |
| ReSearch-InstructRetrieval Strategy=Multi-Hop Retrieval, Supervision Paradigm=Outcome-reward RL, Base Model=Qwen2.5-3B2026.05 | 36.5 | |
| Rejection SamplingBackbone Model=Qwen2.5-7B2025.10 | 36 | |
| SFTBackbone=Qwen2.5-7B-Base2026.02 | 35.3 | |
| RAGBackbone Model=Qwen2.5-7B2025.10 | 34.9 | |
| RAGBackbone=Qwen2.5-7B Instruct2026.05 | 34.9 | |
| ZeroSearchBackbone=Qwen2.5-3B-Base2026.02 | 34.8 | |
| RAGBackbone Model=Qwen2.5-3B2025.10 | 34.8 | |
| RAGBackbone=Qwen2.5-3B Instruct2026.05 | 34.8 | |
| Naive RAGRetrieval Strategy=Single-Hop Retrieval, Base Model=Qwen2.5-3B2026.05 | 34.8 | |
| IRCoTBackbone=Qwen2.5-7B-Base2026.02 | 34.2 | |
| Search-R1-instructBackbone Model=Qwen2.5-3B2025.10 | 34.1 | |
| Search-R1Backbone=Qwen2.5-3B Instruct2026.05 | 34.1 | |
| ReSearchBackbone=Qwen2.5-3B-Base2026.02 | 33.9 | |
| Standard PERType=LLM2026.04 | 33.6 | |
| Search-o1Backbone=Qwen2.5-7B-Base2026.02 | 32.1 | |
| SFTBackbone Model=Qwen2.5-7B2025.10 | 31.8 | |
| RAGBackbone=Qwen2.5-3B-Base2026.02 | 31.2 | |
| SFTBackbone=Qwen2.5-3B-Base2026.02 | 30.1 | |
| R1 (no search)Backbone=Qwen2.5-7B-Base, search=disabled2026.02 | 29.8 | |
| R1-baseBackbone Model=Qwen2.5-7B2025.10 | 29.7 | |
| Rejection SamplingBackbone Model=Qwen2.5-3B2025.10 | 29.4 | |
| Qwen2.5-VL 3BText Source=Pure-Text, CEPE protocol=true2025.10 | 29.31 | |
| IRCoTBackbone=Qwen2.5-3B-Base2026.02 | 28.6 | |
| R1-instructBackbone Model=Qwen2.5-7B2025.10 | 27 | |
| Search-o1Backbone=Qwen2.5-3B-Base2026.02 | 26.7 | |
| R1 (no search)Backbone=Qwen2.5-3B-Base2026.02 | 25.8 | |
| SFTBackbone Model=Qwen2.5-3B2025.10 | 24.9 | |
| SFTRetrieval Strategy=None, Base Model=Qwen2.5-3B2026.05 | 24.9 | |
| SEETOKText Source=Visual-Text, CEPE protocol=true2025.10 | 24.14 | |
| Search-o1Backbone Model=Qwen2.5-3B2025.10 | 23.8 | |
| Search-o1Backbone=Qwen2.5-3B Instruct2026.05 | 23.8 | |
| Search-o1Retrieval Strategy=Multi-Hop Retrieval, Supervision Paradigm=Prompting, Base Model=Qwen2.5-3B2026.05 | 23.8 | |
| CoTBackbone=Qwen2.5-7B-Base2026.02 | 22.8 | |
| R1-baseBackbone Model=Qwen2.5-3B2025.10 | 22.6 | |
| R1-BaseRetrieval Strategy=None, Base Model=Qwen2.5-3B2026.05 | 22.6 | |
| IRCoTBackbone Model=Qwen2.5-7B2025.10 | 22.4 | |
| DirectBackbone=Qwen2.5-7B-Base2026.02 | 21.3 | |
| Qwen2.5-VL 3BText Source=Visual-Text, CEPE protocol=true2025.10 | 21.13 | |
| R1-instructBackbone Model=Qwen2.5-3B2025.10 | 21 | |
| R1-InstructRetrieval Strategy=None, Base Model=Qwen2.5-3B2026.05 | 21 | |
| CoTBackbone=Qwen2.5-3B-Base2026.02 | 16.8 | |
| DirectBackbone=Qwen2.5-3B-Base2026.02 | 15.2 | |
| Search-o1Backbone Model=Qwen2.5-7B2025.10 | 15.1 | |
| Search-o1Backbone=Qwen2.5-7B Instruct2026.05 | 15.1 | |
| Direct InferenceBackbone Model=Qwen2.5-7B2025.10 | 13.4 | |
| IRCoTBackbone Model=Qwen2.5-3B2025.10 | 11.1 | |
| IRCoTRetrieval Strategy=Multi-Hop Retrieval, Supervision Paradigm=Prompting, Base Model=Qwen2.5-3B2026.05 | 11.1 | |
| Direct InferenceBackbone Model=Qwen2.5-3B2025.10 | 10.6 | |
| Direct GenerationRetrieval Strategy=None, Base Model=Qwen2.5-3B2026.05 | 10.6 | |
| CoTBackbone Model=Qwen2.5-7B2025.10 | 4.8 |