Adversarial Detection on Expanded benchmark Hidden premise
29Caught RateTheoria
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| TheoriaModel=GPT-5.5, Web search=Enabled, Evaluation protocol=Structured (step-level typed judges with BEFORE→AFTER diffs)2026.07 | 29 | 90.6 | |
| Holistic judgeModel=GPT-5.5, Web search=Enabled, Evaluation protocol=Holistic (single-call global assessment)2026.07 | 20 | 62.5 |