Interactive Social Privacy and Theory of Mind on Social Intelligence in Adversarial Dialogue (test)
26.7Fooling % (Hard)GPT-5.4
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| GPT-5.4Defender Model=GPT-5.4, Attacker Type=Base Attacker2026.04 | 26.7 | 51.1 | 49.8 | 49.9 | 3.53 | |
| GPT-5.4Defender Model=GPT-5.4, Attacker Type=Cross-Examiner2026.04 | 16.7 | 45.3 | 44.7 | 49.9 | 4.2 | |
| GPT-5.4Defender Model=GPT-5.4, Attacker Type=Bluffing Attacker2026.04 | 15.6 | 24.7 | 22 | 45.6 | 4.99 | |
| GPT-5.4Defender Model=GPT-5.4, Attacker Type=Deception-Aware2026.04 | 6.2 | 38.7 | 40.7 | 46.2 | 4.15 |