Adversarial detection on Expanded benchmark Total
90Caught CountTheoria
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| TheoriaModel=GPT-5.5, Web search=Enabled, Evaluation protocol=Structured (step-level typed judges with BEFORE→AFTER diffs)2026.07 | 90 | 94.7 | |
| Holistic judgeModel=GPT-5.5, Web search=Enabled, Evaluation protocol=Holistic (single-call global assessment)2026.07 | 79 | 83.2 |