Theory-of-Mind Reasoning on FANToM (400-question stratified split)
67.5BeliefQA AccuracyLiteral
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| LiteralCondition=single-call, no tools, Backbone=gpt-4o-mini, seed=422026.06 | 67.5 | 72.5 | 37.5 | 88.7 | 42.5 | 61.7 | |
| NoIntentCondition=ReAct-style tool use up to budget 3, Backbone=gpt-4o-mini, seed=422026.06 | 66.2 | 52.5 | 36.2 | 86.3 | 56.2 | 59.5 | |
| AURA (Intent)Condition=full AURA pipeline with IntentInferrer, Backbone=gpt-4o-mini, seed=422026.06 | 66.2 | 52.5 | 41.2 | 87.5 | 62.5 | 62 |