Factuality Evaluation on HotpotQA
0.686Average ScoreRLFH
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| RLFHBase Model=Llama-3.1-8B-Instruct2024.06 | 0.686 | 6.23 | 2.1 | 100 | 0.714 | |
| RLFHBase Model=Qwen2.5-7B-Instruct2024.06 | 0.668 | 7.3 | 3.66 | 90 | 0.651 | |
| FACTBase Model=Llama-3.1-8B-Instruct, Algorithm=SFT2024.06 | 0.653 | 2.49 | 1.31 | 100 | 0.635 | |
| ITIBase Model=Llama-3.1-8B-Instruct2024.06 | 0.646 | 4.48 | 1.91 | 99 | 0.649 | |
| FACTBase Model=Llama-3.1-8B-Instruct, Algorithm=DPO2024.06 | 0.645 | 4.9 | 2.18 | 99 | 0.652 | |
| Llama-3.1-8BCategory=Open-source Models2024.06 | 0.639 | 4.57 | 2.44 | 99 | 0.652 | |
| Qwen2.5-7BCategory=Open-source Models2024.06 | 0.638 | 9.13 | 4.8 | 93 | 0.634 | |
| DeepSeekV2-LiteCategory=Open-source Models2024.06 | 0.618 | 15.4 | 9.22 | 96 | 0.642 | |
| Falcon3-10BCategory=Open-source Models2024.06 | 0.593 | 5.14 | 3.06 | 90 | 0.608 | |
| Ministral-8BCategory=Open-source Models2024.06 | 0.591 | 7.36 | 3.81 | 96 | 0.633 | |
| DOLABase Model=Llama-3.1-8B-Instruct2024.06 | 0.546 | 6.61 | 6 | 90 | 0.524 | |
| Yi-1.5-9BCategory=Open-source Models2024.06 | 0.536 | 12.7 | 12 | 100 | 0.533 |