Agent Trace Evaluation on TRAIL-annotated GAIA 1.0 (test)
54.7Category F1Holistic Agent Evaluation Framework
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Holistic Agent Evaluation FrameworkExcl. (%)=0, Judge=GPT-5.42026.05 | 54.7 | 82.3 | 61.6 | |
| GPT-5.4Excl. (%)=6.02026.05 | 43.4 | 29.2 | 14.6 | |
| Gemini-2.5-Pro-PreviewExcl. (%)=–2026.05 | 38.9 | 54.6 | 18.3 | |
| Claude-Sonnet-4.6Excl. (%)=8.5, Reasoning-effort settings=Best of2026.05 | 35.2 | 33.6 | 10.6 | |
| Gemini-2.5-Flash-PreviewExcl. (%)=–2026.05 | 33.7 | 37.2 | 10 | |
| FAGI-AgentCompassExcl. (%)=–2026.05 | 30.9 | 65.7 | 23.9 | |
| OpenAI o3Excl. (%)=–2026.05 | 29.6 | 53.5 | 9.2 | |
| Claude-3.7-SonnetExcl. (%)=–2026.05 | 25.4 | 20.4 | 4.7 | |
| Gemini-3.1-ProExcl. (%)=6.0, Reasoning-effort settings=Best of2026.05 | 23.8 | 25.9 | 8.7 | |
| GPT-4.1Excl. (%)=–2026.05 | 21.8 | 10.7 | 2.8 | |
| OpenAI o1Excl. (%)=–2026.05 | 13.8 | 4 | 1.3 |