Factuality Evaluation on SQUAD v2
28.9Correct CountYi-1.5-9B
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Yi-1.5-9BCategory=Open-source Models2024.06 | 28.9 | 10 | 100 | 0.734 | |
| DeepSeekV2-LiteCategory=Open-source Models2024.06 | 23.6 | 7.09 | 98 | 0.754 | |
| Llama-3.1-8BCategory=Open-source Models2024.06 | 22.8 | 6.02 | 98 | 0.777 | |
| DOLABase Model=Llama-3.1-8B-Instruct2024.06 | 22.6 | 8.61 | 97 | 0.713 | |
| FACTBase Model=Llama-3.1-8B-Instruct, Algorithm=DPO2024.06 | 22.6 | 6.31 | 99 | 0.778 | |
| RLFHBase Model=Llama-3.1-8B-Instruct2024.06 | 21.2 | 5.32 | 100 | 0.786 | |
| Qwen2.5-7BCategory=Open-source Models2024.06 | 21.1 | 4.82 | 97 | 0.813 | |
| ITIBase Model=Llama-3.1-8B-Instruct2024.06 | 19.2 | 5.03 | 98 | 0.776 | |
| RLFHBase Model=Qwen2.5-7B-Instruct2024.06 | 17.3 | 3.55 | 96 | 0.83 | |
| FACTBase Model=Llama-3.1-8B-Instruct, Algorithm=SFT2024.06 | 17.2 | 4.6 | 100 | 0.783 | |
| Ministral-8BCategory=Open-source Models2024.06 | 15.7 | 4.26 | 82 | 0.761 | |
| Falcon3-10BCategory=Open-source Models2024.06 | 11.1 | 2.18 | 96 | 0.813 |