ResearchDatasetsBenchmark-500FollowBenchmarksTask NameDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyTask NameDataset NameSOTA ResultTrendResultsPrefill-stage hallucination risk detectionBenchmark-500 Relaxed Consensus (Pvote ≥ 0.8)0.696AUROC (Mean)4Prefill-stage hallucination risk detectionBenchmark-500 Strict Consensus Pvote = 1.0 vs. Clean0.694AUROC (Mean)4
Prefill-stage hallucination risk detectionBenchmark-500 Relaxed Consensus (Pvote ≥ 0.8)0.696AUROC (Mean)4
Prefill-stage hallucination risk detectionBenchmark-500 Strict Consensus Pvote = 1.0 vs. Clean0.694AUROC (Mean)4