AI Agent Reasoning and Tool-use on GAIA
78.49Level 1 Scoreh2oGPTe
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| h2oGPTe2025.02 | 78.49 | 64.78 | 40.82 | 65.12 | — | |
| AgenticReasoningPrimary reasoning model=DeepSeek-R12025.02 | 74.36 | 69.21 | 45.46 | 66.13 | — | |
| OpenAI Deep Research2025.02 | 74.29 | 69.06 | 47.6 | 67.36 | — | |
| InspectReAct2025.02 | 67.92 | 59.3 | 30.77 | 57.58 | — | |
| GPT-5Params=-2026.03 | 67.9 | 58.1 | 46.1 | — | 59.3 | |
| GPT-4.1Params=-2026.03 | 58.5 | 50 | 34.6 | — | 50.3 | |
| Langfun2025.02 | 58.06 | 51.57 | 24.49 | 49.17 | — | |
| TraceR1Params=8B2026.03 | 55.9 | 35.8 | 24.4 | — | 40.2 | |
| HF AgentController=GPT-4o2026.06 | 47.17 | 31.4 | 11.54 | — | 33.4 | |
| GPT-4oParams=-2026.03 | 47.1 | 31.4 | 11.5 | — | 33.4 | |
| Qwen3-VL-8BParams=8B2026.03 | 46.2 | 27.6 | 16.3 | — | 31.5 | |
| ReGRPO (default, λval = 0)Controller=MAT-Qwen2-VL-7B2026.06 | 39.02 | 18.71 | 4.89 | — | 23.35 | |
| SPORT AgentController=Tuned-Qwen2-VL-7B2026.06 | 35.85 | 16.28 | 3.84 | — | 20.61 | |
| HF AgentController=GPT-4o-mini2026.06 | 33.96 | 27.91 | 3.84 | — | 26.06 | |
| Warm-up AgentController=GPT-4-turbo2026.06 | 30.2 | 15.1 | 0 | — | 17.6 | |
| T3-AgentController=MAT-MiniCPM-V-8.5B2026.06 | 26.42 | 11.63 | 3.84 | — | 15.15 | |
| T3-AgentController=MAT-Qwen2-VL-7B2026.06 | 26.42 | 15.12 | 3.84 | — | 16.97 | |
| T3-AgentParams=7B2026.03 | 26.4 | 15.1 | 3.8 | — | 16.9 | |
| Qwen2.5-VL-14BParams=14B2026.03 | 24.5 | 11.6 | 3.8 | — | 15.2 | |
| DeepSeek-VL2Params=72B2026.03 | 19.3 | 12.4 | 10.3 | — | 14.2 | |
| HF AgentController=Qwen2-VL-7B2026.06 | 16.98 | 8.14 | 0 | — | 9.7 | |
| Qwen2.5-VL-7BParams=7B2026.03 | 16.9 | 9.3 | 0 | — | 10.3 | |
| HF AgentController=MiniCPM-V-8.5B2026.06 | 13.21 | 5.81 | 0 | — | 7.27 | |
| HF AgentController=LLaVA-NeXT-8B2026.06 | 9.43 | 1.16 | 0 | — | 3.64 | |
| LLAVA-NeXT-8BParams=8B2026.03 | 9.4 | 1.2 | 0 | — | 3.6 | |
| HF AgentController=InternVL2-8B2026.06 | 7.55 | 4.65 | 0 | — | 4.85 |