Clinical Summarization on Hallucination-Generated-DI
13Hallucination CountHDSR-PL
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| HDSR-PLBase Model=LLaMA-3.2-3B-Instruct, Hallucination Detector=MedCat, Variant=best, Evaluation Protocol=automatic LLM-as-Judge2026.05 | 13 | 4.23 | 4.58 | 4.68 | 4.3 | 4.44 | |
| HDSR-PLBase Model=Gemma-3-4b-it, Hallucination Detector=MedCat, Variant=best, Evaluation Protocol=automatic LLM-as-Judge2026.05 | 13 | 4.33 | 4.65 | 4.7 | 4.63 | 4.58 | |
| HDSR-PLBase Model=LLaMA-3.1-8B-Instruct, Hallucination Detector=MedCat, Variant=best2026.05 | 15 | 4.4 | 4.28 | 4.05 | 3.9 | 4.16 | |
| PromptingBase Model=Gemma-3-4b-it, Evaluation Protocol=automatic LLM-as-Judge2026.05 | 15 | 4.23 | 4.65 | 4.68 | 4.48 | 4.51 | |
| HDSRBase Model=LLaMA-3.1-8B-Instruct, Hallucination Detector=MedAlign, Variant=best2026.05 | 22 | 4.13 | 4.48 | 4.53 | 3.95 | 4.27 | |
| PromptingBase Model=LLaMA-3.2-3B-Instruct, Evaluation Protocol=automatic LLM-as-Judge2026.05 | 26 | 4.13 | 4.58 | 4.68 | 4.2 | 4.39 | |
| PromptingBase Model=LLaMA-3.1-8B-Instruct2026.05 | 29 | 4.08 | 3.83 | 4.05 | 3.23 | 3.79 | |
| GPT-5 PromptingMethod=Prompting2026.05 | 36 | 3.55 | 4.73 | 4.73 | 4.08 | 4.27 | |
| SFTBase Model=LLaMA-3.1-8B-Instruct2026.05 | 57 | 3.03 | 4.43 | 4.53 | 3.03 | 3.75 |