Retrieval-Augmented Generation on NQ (accuracy)
77.1AccuracyReFeedL
Evaluation Results
| Method | Links | |
|---|---|---|
| ReFeedLModel=Llama-3.1-8B, Retrieval Selection=top-12026.02 | 77.1 | |
| OursModel=Qwen3-8B, Retrieval Selection=top-12026.02 | 76.9 | |
| OursModel=Llama-3.1-8B, Retrieval Selection=top-12026.02 | 76.8 | |
| InstructRAGModel=Llama-3.1-8B, Retrieval Selection=top-12026.02 | 76.3 | |
| DPRModel=Llama-3.1-8B, Retrieval Selection=top-12026.02 | 75.6 | |
| ReFeedLModel=Qwen3-8B, Retrieval Selection=top-12026.02 | 75.6 | |
| InstructRAGModel=Qwen3-8B, Retrieval Selection=top-12026.02 | 75.1 | |
| BM25Model=Llama-3.1-8B, Retrieval Selection=top-12026.02 | 74.7 | |
| DPRModel=Qwen3-8B, Retrieval Selection=top-12026.02 | 74.5 | |
| BM25Model=Qwen3-8B, Retrieval Selection=top-12026.02 | 73.8 | |
| Raw MetricGenerator=Llama3-8B-Instruct, Reward Model=Raw Metric2025.02 | 48.96 | |
| ConsJudgeGenerator=Llama3-8B-Instruct, Reward Model Backbone=Llama3-8B-Instruct2025.02 | 48.78 | |
| ConsJudgeGenerator=MiniCPM-2.4B, Reward Model Backbone=Qwen2.5-14B-Instruct2025.02 | 48.01 | |
| Vanilla LLMGenerator=MiniCPM-2.4B, Reward Model Backbone=Qwen2.5-14B-Instruct2025.02 | 47.3 | |
| Vanilla LLMGenerator=MiniCPM-2.4B, Reward Model Backbone=Llama3-8B-Instruct2025.02 | 47.23 | |
| SFTGenerator=Llama3-8B-Instruct, Reward Model Backbone=Llama3-8B-Instruct2025.02 | 47.16 | |
| SFTGenerator=MiniCPM-2.4B, Reward Model Backbone=Llama3-8B-Instruct2025.02 | 47.02 | |
| ConsJudgeGenerator=MiniCPM-2.4B, Reward Model Backbone=Llama3-8B-Instruct2025.02 | 47.02 | |
| SFTGenerator=MiniCPM-2.4B, Reward Model Backbone=Qwen2.5-14B-Instruct2025.02 | 47.02 | |
| Vanilla LLMGenerator=Llama3-8B-Instruct, Reward Model Backbone=Llama3-8B-Instruct2025.02 | 46.63 | |
| Raw MetricGenerator=MiniCPM-2.4B, Reward Model=Raw Metric2025.02 | 46.14 | |
| ZeroModel=Qwen3-8B, Retrieval Selection=top-12026.02 | 44.2 | |
| ZeroModel=Llama-3.1-8B, Retrieval Selection=top-12026.02 | 38.5 |