Hallucination Detection on LLM-generated screenplays Aggregate
91.4PrecisionAtlas
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| Atlas2026.07 | 91.4 | 80 | 85.3 | — | — | — | — | |
| LLM-as-a-Judge2026.07 | 83.8 | 77.5 | 80.5 | — | — | — | — | |
| AtlasEvaluator=GPT-5.42026.07 | 0.914 | 0.8 | 0.853 | 40 | 32 | 3 | 8 | |
| LLM-as-a-JudgeEvaluator=GPT-5.42026.07 | 0.838 | 0.775 | 0.805 | 40 | 31 | 6 | 9 |