Retrieval-Augmented Generation on MARCOQA
88.25LLM ScoreConsJudge
Evaluation Results
| Method | Links | |
|---|---|---|
| ConsJudgeGenerator=Llama3-8B-Instruct, Reward Model Backbone=Llama3-8B-Instruct2025.02 | 88.25 | |
| Vanilla LLMGenerator=Llama3-8B-Instruct, Reward Model Backbone=Llama3-8B-Instruct2025.02 | 88.15 | |
| SFTGenerator=Llama3-8B-Instruct, Reward Model Backbone=Llama3-8B-Instruct2025.02 | 87.46 | |
| SFTGenerator=MiniCPM-2.4B, Reward Model Backbone=Llama3-8B-Instruct2025.02 | 86.35 | |
| ConsJudgeGenerator=MiniCPM-2.4B, Reward Model Backbone=Llama3-8B-Instruct2025.02 | 86.16 | |
| Vanilla LLMGenerator=MiniCPM-2.4B, Reward Model Backbone=Llama3-8B-Instruct2025.02 | 86 | |
| SFTGenerator=MiniCPM-2.4B, Reward Model Backbone=Qwen2.5-14B-Instruct2025.02 | 85.9 | |
| ConsJudgeGenerator=MiniCPM-2.4B, Reward Model Backbone=Qwen2.5-14B-Instruct2025.02 | 85.73 | |
| Vanilla LLMGenerator=MiniCPM-2.4B, Reward Model Backbone=Qwen2.5-14B-Instruct2025.02 | 85.59 | |
| Raw MetricGenerator=Llama3-8B-Instruct, Reward Model=Raw Metric2025.02 | 84.8 | |
| Raw MetricGenerator=MiniCPM-2.4B, Reward Model=Raw Metric2025.02 | 84.75 |