Behavioral simulation on 226-example (held-out)
77Modal AccuracyClaude Opus 4.7
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Claude Opus 4.7Role=Frontier model2026.06 | 77 | 2.87 | |
| GPT-5.5Role=Frontier model2026.06 | 76.2 | 3.99 | |
| Claude Opus 4.6Role=Frontier model2026.06 | 76.1 | 4.18 | |
| TL-Twin DeltaRole=Direct scenario model2026.06 | 75 | 2.36 | |
| TL-Twin GammaRole=Ensemble scenario model2026.06 | 74.9 | 2.28 | |
| GPT-5.4Role=Frontier model2026.06 | 74.9 | 2.59 | |
| Gemini 3.1 ProRole=Frontier model2026.06 | 74.2 | 4.19 | |
| TL-Twin BetaRole=Calibrated behavior model2026.06 | 72.8 | 2.2 | |
| TL-Twin AlphaRole=Population-movement model2026.06 | 70.5 | 1.16 |