Agent Trace Evaluation on TRAIL-annotated SWE-bench 1.0 (test)
69.8Category F1Holistic Agent Evaluation Framework
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Holistic Agent Evaluation FrameworkExcl. (%)=0, Judge=GPT-5.42026.05 | 69.8 | 86 | 63.8 | |
| GPT-5.4Excl. (%)=22.62026.05 | 50.4 | 7 | 1.4 | |
| Claude-Sonnet-4.6Excl. (%)=29.0, Reasoning-effort settings=Best of2026.05 | 41 | 0 | 0 | |
| Gemini-3.1-ProExcl. (%)=25.8, Reasoning-effort settings=Best of2026.05 | 33 | 7.2 | 0.3 | |
| FAGI-AgentCompassExcl. (%)=–2026.05 | 23.2 | 25 | 5.1 | |
| Gemini-2.5-Flash-PreviewExcl. (%)=–2026.05 | 21.3 | 6 | 0 | |
| GPT-4.1Excl. (%)=–2026.05 | 16.6 | 0 | 0 | |
| Gemini-2.5-Pro-PreviewExcl. (%)=–2026.05 | 14.8 | 23.8 | 5 |