Hallucination Detection on HotpotQA
0.928AUROCPR
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| PRModel=LLama-3.2-1B2026.01 | 0.928 | — | |
| LatentAuditModel Backbone=Llama-3 (8B)2026.04 | 0.928 | — | |
| LatentAuditModel Backbone=Qwen-3 (8B)2026.04 | 0.922 | — | |
| LatentAuditModel Backbone=Qwen-2.5 (7B)2026.04 | 0.918 | — | |
| HIVEBase LLM=Dream-7B-Instruct2026.04 | 0.9176 | 0.9737 | |
| MoPModel=LLama-3.2-1B2026.01 | 0.9151 | — | |
| LatentAuditModel Backbone=Mistral (7B)2026.04 | 0.91 | — | |
| HIVEBase LLM=LLaDA-8B-Instruct2026.04 | 0.9094 | 0.9648 | |
| MoP-VanillaExpertsModel=LLama-3.2-1B2026.01 | 0.9091 | — | |
| PRModel=LLama-3.2-3B2026.01 | 0.9066 | — | |
| LatentAuditModel Backbone=Llama-2 (7B)2026.04 | 0.905 | — | |
| Probing BaselineModel=LLama-3.2-1B2026.01 | 0.9025 | — | |
| CCSBase LLM=Dream-7B-Instruct2026.04 | 0.896 | 0.9665 | |
| PRModel=Mistral-7B-v0.12026.01 | 0.8904 | — | |
| PRBackbone=Mistral-7B-v0.32026.01 | 0.8903 | — | |
| PRModel=Mistral-7B-v0.32026.01 | 0.8903 | — | |
| MoPModel=LLama-3.2-3B2026.01 | 0.8816 | — | |
| PRModel=LLama-3-8B2026.01 | 0.8781 | — | |
| PRBackbone=Llama-3-8B2026.01 | 0.8781 | — | |
| CCSBase LLM=LLaDA-8B-Instruct2026.04 | 0.8779 | 0.9298 | |
| PRModel=LLama-3-70B2026.01 | 0.8769 | — | |
| OSCARBackbone=LLaDA-8B-Instruct, Method Category=Traj., Sample Count=64, Evaluation Protocol=LLM-as-Judge2026.04 | 0.875 | — | |
| TraceDetBase LLM=Dream-7B-Instruct2026.04 | 0.874 | 0.9564 | |
| OSCARBackbone=Dream-7B-Instruct, Method Category=Traj., Sample Count=64, Evaluation Protocol=LLM-as-Judge2026.04 | 0.871 | — | |
| OSCARBackbone=LLaDA-8B-Instruct, Method Category=Traj., Sample Count=128, Evaluation Protocol=LLM-as-Judge2026.04 | 0.867 | — | |
| MoPModel=LLama-3-70B2026.01 | 0.8665 | — | |
| Probing BaselineModel=LLama-3.2-3B2026.01 | 0.8654 | — | |
| MoP-VanillaExpertsModel=LLama-3.2-3B2026.01 | 0.8641 | — | |
| TraceDetBase LLM=LLaDA-8B-Instruct2026.04 | 0.8605 | 0.9428 | |
| MoPBackbone=Mistral-7B-v0.32026.01 | 0.8582 | — | |
| MoPModel=Mistral-7B-v0.32026.01 | 0.8582 | — | |
| DynHDBackbone=Dream-7B-Instruct, Method Category=Traj., Sample Count=64, Evaluation Protocol=LLM-as-Judge, Requires trained classifier=true2026.04 | 0.856 | — | |
| MoPModel=LLama-3-8B2026.01 | 0.8545 | — | |
| MoPBackbone=Llama-3-8B2026.01 | 0.8545 | — | |
| DynHDBackbone=LLaDA-8B-Instruct, Method Category=Traj., Sample Count=64, Evaluation Protocol=LLM-as-Judge, Requires trained classifier=true2026.04 | 0.853 | — | |
| MoPModel=Mistral-7B-v0.12026.01 | 0.8463 | — | |
| MoP-VanillaExpertsModel=LLama-3-8B2026.01 | 0.8457 | — | |
| MoP-VanillaExpertsBackbone=Llama-3-8B2026.01 | 0.8457 | — | |
| Probing BaselineModel=LLama-3-70B2026.01 | 0.8445 | — | |
| DynHDBackbone=LLaDA-8B-Instruct, Method Category=Traj., Sample Count=128, Evaluation Protocol=LLM-as-Judge, Requires trained classifier=true2026.04 | 0.842 | — | |
| MoP-RandomGateModel=LLama-3.2-1B2026.01 | 0.8411 | — | |
| Probing BaselineModel=Mistral-7B-v0.12026.01 | 0.8376 | — | |
| MoP-VanillaExpertsModel=Mistral-7B-v0.12026.01 | 0.8354 | — | |
| Probing BaselineBackbone=Mistral-7B-v0.32026.01 | 0.8319 | — | |
| Probing BaselineModel=Mistral-7B-v0.32026.01 | 0.8319 | — | |
| MoP-VanillaExpertsBackbone=Mistral-7B-v0.32026.01 | 0.8293 | — | |
| MoP-VanillaExpertsModel=Mistral-7B-v0.32026.01 | 0.8293 | — | |
| Val LossBackbone=LLaMA-3.1-8B-Instruct, First-sentence truncation=true, Forward horizon (w)=72026.05 | 0.8287 | — | |
| FEPoIDBackbone=LLaMA-3.1-8B-Instruct, First-sentence truncation=true, Forward horizon (w)=72026.05 | 0.8287 | — | |
| MoP-VanillaExpertsModel=LLama-3-70B2026.01 | 0.8248 | — | |
| Probing BaselineModel=LLama-3-8B2026.01 | 0.8223 | — | |
| Probing BaselineBackbone=Llama-3-8B2026.01 | 0.8223 | — | |
| OSCARBackbone=Dream-7B-Instruct, Method Category=Traj., Sample Count=128, Evaluation Protocol=LLM-as-Judge2026.04 | 0.819 | — | |
| Semantic EntropyBase LLM=Dream-7B-Instruct2026.04 | 0.8184 | 0.931 | |
| FEPoIDBackbone=Mistral-7B-Instruct-v0.3, First-sentence truncation=true, Forward horizon (w)=72026.05 | 0.8153 | — | |
| CurvatureBackbone=LLaMA-3.1-8B-Instruct, First-sentence truncation=true, Forward horizon (w)=72026.05 | 0.8119 | — | |
| CurvatureBackbone=Mistral-7B-Instruct-v0.3, First-sentence truncation=true, Forward horizon (w)=72026.05 | 0.8088 | — | |
| Val LossBackbone=Mistral-7B-Instruct-v0.3, First-sentence truncation=true, Forward horizon (w)=72026.05 | 0.8087 | — | |
| RankMEBackbone=LLaMA-3.1-8B-Instruct, First-sentence truncation=true, Forward horizon (w)=72026.05 | 0.8013 | — | |
| DynHDBackbone=Dream-7B-Instruct, Method Category=Traj., Sample Count=128, Evaluation Protocol=LLM-as-Judge, Requires trained classifier=true2026.04 | 0.801 | — | |
| MoP-RandomGateModel=LLama-3-8B2026.01 | 0.7988 | — | |
| MoP-RandomGateBackbone=Llama-3-8B2026.01 | 0.7988 | — | |
| FEPoIDLLM Backbone=Mistral-7B-Instruct-v0.3, Method Category=Hidden-State Probing2026.05 | 0.7982 | — | |
| OSCARBackbone=Dream-7B-Instruct, Method Category=Traj., Sample Count=64, Evaluation Protocol=EM2026.04 | 0.798 | — | |
| FEPoIDBackbone=LlaMA-3.1-8B, Forward horizon (w)=7, Representation extraction method=FST, Model tuning status=non-instruction-tuned (base model)2026.05 | 0.7972 | — | |
| Lexical SimilarityBackbone=LLaMA-3.1-8B-Instruct, First-sentence truncation=true, Forward horizon (w)=72026.05 | 0.7924 | — | |
| ICR ProbeModel=Qwen2.5-3B2025.07 | 0.7905 | — | |
| TSVBase LLM=Dream-7B-Instruct2026.04 | 0.7847 | 0.9232 | |
| CurvatureBackbone=LlaMA-3.1-8B, Forward horizon (w)=7, Representation extraction method=FST, Model tuning status=non-instruction-tuned (base model)2026.05 | 0.7839 | — | |
| FEPoIDLLM Backbone=LlaMA-3.1-8B-Instruct, Method Category=Hidden-State Probing2026.05 | 0.7807 | — | |
| Semantic EntropyBase LLM=LLaDA-8B-Instruct2026.04 | 0.78 | 0.9004 | |
| Attention DivergenceSingle Generation=true, Backbone=Mistral-7B2026.05 | 0.78 | — | |
| ICR ProbeModel=Qwen2.5-14B2025.07 | 0.7751 | — | |
| SAPLMAModel=Qwen2.5-3B2025.07 | 0.7747 | — | |
| OSCARBackbone=Dream-7B-Instruct, Method Category=Traj., Sample Count=128, Evaluation Protocol=EM2026.04 | 0.774 | — | |
| Val LossBackbone=LlaMA-3.1-8B, Forward horizon (w)=7, Representation extraction method=FST, Model tuning status=non-instruction-tuned (base model)2026.05 | 0.7735 | — | |
| SNRBackbone=LLaMA-3.1-8B-Instruct, First-sentence truncation=true, Forward horizon (w)=72026.05 | 0.7723 | — | |
| IDBackbone=LLaMA-3.1-8B-Instruct, First-sentence truncation=true, Forward horizon (w)=72026.05 | 0.7723 | — | |
| Val LossBackbone=LlaMA-3.2-1B, Forward horizon (w)=3, FST extraction status=without FST2026.05 | 0.7641 | — | |
| FEPoIDBackbone=LlaMA-3.2-1B, Forward horizon (w)=3, FST extraction status=without FST2026.05 | 0.7641 | — | |
| RGNBackbone=Mistral-7B-Instruct-v0.3, First-sentence truncation=true, Forward horizon (w)=72026.05 | 0.7631 | — | |
| IDBackbone=Mistral-7B-Instruct-v0.3, First-sentence truncation=true, Forward horizon (w)=72026.05 | 0.7631 | — | |
| TSVBase LLM=LLaDA-8B-Instruct2026.04 | 0.7616 | 0.8953 | |
| TraceDetBackbone=Dream-7B-Instruct, Method Category=Traj., Sample Count=64, Evaluation Protocol=LLM-as-Judge, Requires trained classifier=true2026.04 | 0.76 | — | |
| CurvatureBackbone=LlaMA-3.2-1B, Forward horizon (w)=3, FST extraction status=without FST2026.05 | 0.7587 | — | |
| Lexical SimilarityBackbone=Mistral-7B-Instruct-v0.3, First-sentence truncation=true, Forward horizon (w)=72026.05 | 0.7534 | — | |
| MoP-RandomGateModel=LLama-3.2-3B2026.01 | 0.7513 | — | |
| TraceDetBackbone=Dream-7B-Instruct, Method Category=Traj., Sample Count=128, Evaluation Protocol=LLM-as-Judge, Requires trained classifier=true2026.04 | 0.751 | — | |
| EigenScoreBase LLM=LLaDA-8B-Instruct2026.04 | 0.7507 | 0.8926 | |
| RGNBackbone=LlaMA-3.1-8B, Forward horizon (w)=7, Representation extraction method=FST, Model tuning status=non-instruction-tuned (base model)2026.05 | 0.7499 | — | |
| Attention probeRegime=Pre-gen., Model=Qwen2.5-32B2026.06 | 0.7484 | — | |
| Attention probe, soft targetRegime=Pre-gen., Model=Qwen2.5-32B2026.06 | 0.7484 | — | |
| Attention probeRegime=Post-gen., Model=Qwen2.5-32B2026.06 | 0.7476 | — | |
| Attention probe, soft targetRegime=Post-gen., Model=Qwen2.5-32B2026.06 | 0.7476 | — | |
| MoP-RandomGateModel=Mistral-7B-v0.12026.01 | 0.7451 | — | |
| Pred. EntropyBackbone=Mistral-7B-Instruct-v0.3, First-sentence truncation=true, Forward horizon (w)=72026.05 | 0.7449 | — | |
| EigenScoreBase LLM=Dream-7B-Instruct2026.04 | 0.7446 | 0.9033 | |
| Val LossBackbone=LlaMA-3.2-3B, Forward horizon (w)=7, FST extraction status=without FST2026.05 | 0.7439 | — | |
| SNRBackbone=LlaMA-3.2-3B, Forward horizon (w)=7, FST extraction status=without FST2026.05 | 0.7439 | — | |
| Lexical SimilarityBase LLM=Dream-7B-Instruct2026.04 | 0.7429 | 0.9042 |