Hallucination Detection on LLM-generated screenplays Story 1
100PrecisionAtlas
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| Atlas2026.07 | 100 | 75 | 85.7 | — | — | — | — | |
| LLM-as-a-Judge2026.07 | 100 | 68.8 | 81.5 | — | — | — | — | |
| AtlasEvaluator=GPT-5.42026.07 | 100 | 75 | 85.7 | 16 | 12 | 0 | 4 | |
| LLM-as-a-JudgeEvaluator=GPT-5.42026.07 | 100 | 68.8 | 81.5 | 16 | 11 | 0 | 5 |