Adversarial Theory of Mind on TOM-SB
42.4Fooling Rate (Hard)ADA (Fooling + ToM)
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| ADA (Fooling + ToM)Method Type=Training-based, Defender Model=Qwen3-14B, Attacker Model=Gemini3-flash, Reward Configuration=Fooling + ToM2026.04 | 42.4 | 51.3 | 58.7 | 64.9 | 4.58 | |
| ADA (ToM Only)Method Type=Training-based, Defender Model=Qwen3-14B, Attacker Model=Gemini3-flash, Reward Configuration=ToM Only2026.04 | 40.6 | 43.6 | 53.3 | 65.5 | 4.66 | |
| Base Prompt (Gemini3-Pro)Method Type=Prompt-based, Defender Model=Gemini3-Pro, Attacker Model=Gemini3-flash2026.04 | 34.4 | 57.8 | 48.9 | 47 | 3.6 | |
| ADA (Fooling only)Method Type=Training-based, Defender Model=Qwen3-14B, Attacker Model=Gemini3-flash, Reward Configuration=Fooling only2026.04 | 34.4 | 46.7 | 49 | 62.4 | 4.04 | |
| Base Prompt (GPT-5.4)Method Type=Prompt-based, Defender Model=GPT-5.4, Attacker Model=Gemini3-flash2026.04 | 26.7 | 51.1 | 49.8 | 49.9 | 3.53 | |
| Base PromptMethod Type=Prompt-based, Defender Model=Qwen3-14B, Attacker Model=Gemini3-flash2026.04 | 13.2 | 36 | 36 | 49.3 | 3.12 | |
| Online SFTMethod Type=Training-based, Defender Model=Qwen3-14B, Attacker Model=Gemini3-flash2026.04 | 12.5 | 35.6 | 33.8 | 48.3 | 3.13 | |
| Mislead PromptMethod Type=Prompt-based, Defender Model=Qwen3-14B, Attacker Model=Gemini3-flash2026.04 | 4.2 | 37.3 | — | — | 2.69 | |
| Refuse PromptMethod Type=Prompt-based, Defender Model=Qwen3-14B, Attacker Model=Gemini3-flash2026.04 | 0.3 | 0.2 | — | — | 5.94 |