Hallucination Detection on HaluEval (Component Metrics)
72.2Dialogue ScoreThree-Stage Fine-Tuning Method
Evaluation Results
| Method | Links | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Three-Stage Fine-Tuning MethodInstruction Model=Qwen2.5-14B2025.04 | 72.2 | 76.3 | 73.2 | 71.3 | — | — | — | — | — | — | |
| Three-Stage Fine-Tuning MethodInstruction Model=Qwen2-7B2025.04 | 72.1 | 68.1 | 68.9 | 68 | — | — | — | — | — | — | |
| Supervised fine-tuning (SFT)Instruction Model=Qwen2-7B, Processing=Direct fine-tuning2025.04 | 71.9 | 54.2 | 62.3 | 58.6 | — | — | — | — | — | — | |
| Supervised fine-tuning (SFT)Instruction Model=Qwen2.5-14B, Processing=Direct fine-tuning2025.04 | 71.3 | 73.8 | 68.2 | 60.7 | — | — | — | — | — | — | |
| OriginalInstruction Model=Qwen2.5-14B, Processing=None2025.04 | 70.5 | 73.2 | 62.7 | 54.3 | — | — | — | — | — | — | |
| OriginalInstruction Model=Qwen2-7B, Processing=None2025.04 | 70.3 | 51.5 | 54.5 | 51.1 | — | — | — | — | — | — | |
| Three-Stage Fine-Tuning MethodInstruction Model=Llama3-8B2025.04 | 68.6 | 62.7 | 67.7 | 66.9 | — | — | — | — | — | — | |
| Three-Stage Fine-Tuning MethodInstruction Model=Llama3.2-3B2025.04 | 60.8 | 57.4 | 53.9 | 52.5 | — | — | — | — | — | — | |
| Supervised fine-tuning (SFT)Instruction Model=Llama3-8B, Processing=Direct fine-tuning2025.04 | 59.1 | 53.8 | 57.3 | 53.2 | — | — | — | — | — | — | |
| OriginalInstruction Model=Llama3-8B, Processing=None2025.04 | 57.3 | 50 | 51.8 | 49.6 | — | — | — | — | — | — | |
| Self-Improving PretrainingTraining Data=SlimPajama, Pretraining for=Factuality2026.01 | 54.6 | — | 58.5 | — | 84.7 | — | — | — | — | — | |
| Supervised fine-tuning (SFT)Instruction Model=Llama3.2-3B, Processing=Direct fine-tuning2025.04 | 53.2 | 35 | 51.9 | 48.1 | — | — | — | — | — | — | |
| Llama Pretrain BaselineTraining Data=SlimPajama, Pretraining for=Factuality2026.01 | 50.8 | — | 51.4 | — | 61.5 | — | — | — | — | — | |
| Llama BaseStage=Base Model2026.01 | 50 | — | 50.1 | — | 50 | — | — | — | — | — | |
| OriginalInstruction Model=Llama3.2-3B, Processing=None2025.04 | 49.9 | 27.2 | 49.5 | 46.2 | — | — | — | — | — | — | |
| Claude Opus 4.6Evaluation Mode=Single2026.04 | — | — | — | — | — | 16.2 | 20.1 | 13.8 | 16.7 | — | |
| Council ModeEvaluation Mode=Multi-agent Consensus2026.04 | — | — | — | — | — | 10.1 | 13.6 | 8.4 | 10.7 | -35.9 | |
| DeepSeek V3.2Evaluation Mode=Single2026.04 | — | — | — | — | — | 21.3 | 26.5 | 18.7 | 22.2 | — | |
| Gemini 3.1 ProEvaluation Mode=Single2026.04 | — | — | — | — | — | 19.4 | 24.8 | 16.1 | 20.1 | — | |
| GPT-5.4Evaluation Mode=Single2026.04 | — | — | — | — | — | 18.7 | 23.4 | 15.3 | 19.1 | — | |
| Seed 2.0 ProEvaluation Mode=Single2026.04 | — | — | — | — | — | 22.8 | 28.2 | 19.5 | 23.5 | — |