Hallucination Detection on AIS
0.644AccuracyVISTA
Evaluation Results
| Method | Links | |
|---|---|---|
| VISTAModel=Qwen-32B2025.10 | 0.644 | |
| VISTAModel=Qwen-8B2025.10 | 0.636 | |
| VISTAModel=GPT-4o2025.10 | 0.63 | |
| VISTAModel=Llama-70B2025.10 | 0.624 | |
| VISTAModel=Llama-8B2025.10 | 0.622 | |
| VISTAModel=GPT-52025.10 | 0.602 | |
| VISTAModel=Deepseek2025.10 | 0.596 | |
| Fact ScoreModel=GPT-52025.10 | 0.592 | |
| Fact ScoreModel=Deepseek2025.10 | 0.588 | |
| VISTAModel=Mistral-7B2025.10 | 0.586 | |
| LLM-as JudgeModel=GPT-52025.10 | 0.574 | |
| Fact ScoreModel=GPT-4o2025.10 | 0.568 | |
| LLM-as JudgeModel=GPT-4o2025.10 | 0.568 | |
| Fact ScoreModel=Qwen-8B2025.10 | 0.562 | |
| Fact ScoreModel=Llama-70B2025.10 | 0.56 | |
| LLM-as JudgeModel=Llama-70B2025.10 | 0.558 | |
| LLM-as JudgeModel=Deepseek2025.10 | 0.532 | |
| Fact ScoreModel=Qwen-32B2025.10 | 0.532 | |
| LLM-as JudgeModel=Qwen-8B2025.10 | 0.532 | |
| Fact ScoreModel=Mistral-7B2025.10 | 0.526 | |
| LLM-as JudgeModel=Llama-8B2025.10 | 0.518 | |
| Fact ScoreModel=Llama-8B2025.10 | 0.516 | |
| LLM-as JudgeModel=Mistral-7B2025.10 | 0.488 | |
| LLM-as JudgeModel=Qwen-32B2025.10 | 0.464 |