Hallucination detection on BEGIN
87.2AccuracyVISTA
Evaluation Results
| Method | Links | |
|---|---|---|
| VISTAModel=GPT-52025.10 | 87.2 | |
| VISTAModel=Deepseek2025.10 | 84.6 | |
| VISTAModel=GPT-4o2025.10 | 83.2 | |
| VISTAModel=Qwen-32B2025.10 | 80.6 | |
| VISTAModel=Qwen-8B2025.10 | 80.6 | |
| LLM-as JudgeModel=Llama-70B2025.10 | 79 | |
| VISTAModel=Llama-70B2025.10 | 77.4 | |
| VISTAModel=Llama-8B2025.10 | 73.8 | |
| VISTAModel=Mistral-7B2025.10 | 72 | |
| Fact ScoreModel=GPT-52025.10 | 71 | |
| LLM-as JudgeModel=Deepseek2025.10 | 70.8 | |
| LLM-as JudgeModel=GPT-4o2025.10 | 70.4 | |
| LLM-as JudgeModel=Qwen-8B2025.10 | 70.2 | |
| LLM-as JudgeModel=GPT-52025.10 | 70 | |
| Fact ScoreModel=GPT-4o2025.10 | 65.8 | |
| Fact ScoreModel=Qwen-32B2025.10 | 64.4 | |
| Fact ScoreModel=Qwen-8B2025.10 | 64.4 | |
| LLM-as JudgeModel=Llama-8B2025.10 | 61 | |
| Fact ScoreModel=Deepseek2025.10 | 59.8 | |
| LLM-as JudgeModel=Mistral-7B2025.10 | 57.4 | |
| Fact ScoreModel=Llama-8B2025.10 | 53.8 | |
| Fact ScoreModel=Mistral-7B2025.10 | 53.8 | |
| Fact ScoreModel=Llama-70B2025.10 | 53 | |
| LLM-as JudgeModel=Qwen-32B2025.10 | 46.8 |