Hallucination Detection on FaithDial
81.7AccuracyVISTA
Evaluation Results
| Method | Links | |
|---|---|---|
| VISTAModel=Deepseek2025.10 | 81.7 | |
| VISTAModel=GPT-4o2025.10 | 79.95 | |
| VISTAModel=GPT-52025.10 | 79.54 | |
| VISTAModel=Qwen-32B2025.10 | 75.73 | |
| VISTAModel=Qwen-8B2025.10 | 75.1 | |
| LLM-as JudgeModel=Llama-70B2025.10 | 72.9 | |
| VISTAModel=Llama-70B2025.10 | 72.36 | |
| VISTAModel=Mistral-7B2025.10 | 72.01 | |
| Fact ScoreModel=GPT-52025.10 | 67.03 | |
| VISTAModel=Llama-8B2025.10 | 65.19 | |
| LLM-as JudgeModel=GPT-52025.10 | 64.65 | |
| Fact ScoreModel=Deepseek2025.10 | 63.75 | |
| Fact ScoreModel=GPT-4o2025.10 | 62.81 | |
| LLM-as JudgeModel=GPT-4o2025.10 | 60.43 | |
| Fact ScoreModel=Qwen-32B2025.10 | 58.41 | |
| Fact ScoreModel=Qwen-8B2025.10 | 58.19 | |
| LLM-as JudgeModel=Qwen-8B2025.10 | 56.47 | |
| LLM-as JudgeModel=Deepseek2025.10 | 55.45 | |
| Fact ScoreModel=Llama-70B2025.10 | 54.6 | |
| Fact ScoreModel=Llama-8B2025.10 | 54.6 | |
| Fact ScoreModel=Mistral-7B2025.10 | 48.41 | |
| LLM-as JudgeModel=Llama-8B2025.10 | 48.23 | |
| LLM-as JudgeModel=Mistral-7B2025.10 | 46.43 | |
| LLM-as JudgeModel=Qwen-32B2025.10 | 35.89 |