Retrieval-Augmented Generation on TriviaQA
88.26AccuracyConsJudge
Evaluation Results
| Method | Links | |
|---|---|---|
| ConsJudgeGenerator=Llama3-8B-Instruct, Reward Model Backbone=Llama3-8B-Instruct2025.02 | 88.26 | |
| Raw MetricGenerator=Llama3-8B-Instruct, Reward Model=Raw Metric2025.02 | 86.83 | |
| SFTGenerator=Llama3-8B-Instruct, Reward Model Backbone=Llama3-8B-Instruct2025.02 | 86.1 | |
| Vanilla LLMGenerator=Llama3-8B-Instruct, Reward Model Backbone=Llama3-8B-Instruct2025.02 | 85.13 | |
| ConsJudgeGenerator=MiniCPM-2.4B, Reward Model Backbone=Llama3-8B-Instruct2025.02 | 80.8 | |
| ConsJudgeGenerator=MiniCPM-2.4B, Reward Model Backbone=Qwen2.5-14B-Instruct2025.02 | 80.69 | |
| Vanilla LLMGenerator=MiniCPM-2.4B, Reward Model Backbone=Llama3-8B-Instruct2025.02 | 80.4 | |
| SFTGenerator=MiniCPM-2.4B, Reward Model Backbone=Llama3-8B-Instruct2025.02 | 80.33 | |
| Raw MetricGenerator=MiniCPM-2.4B, Reward Model=Raw Metric2025.02 | 80.03 | |
| Vanilla LLMGenerator=MiniCPM-2.4B, Reward Model Backbone=Qwen2.5-14B-Instruct2025.02 | 80.03 | |
| SFTGenerator=MiniCPM-2.4B, Reward Model Backbone=Qwen2.5-14B-Instruct2025.02 | 79.8 |