Expert-Level Reasoning on GAIA text-only (val)
81.6Inference AccuracyReThinker
Evaluation Results
| Method | Links | |
|---|---|---|
| ReThinkerModel Category=Inference Frameworks, Backbone=Gemini-3-Pro, Evaluation Protocol=gpt-4.1-2025-04-14 judge2026.02 | 81.6 | |
| Gemini-3-ProModel Category=Foundation Models with Tools, Evaluation Protocol=gpt-4.1-2025-04-14 judge2026.02 | 79 | |
| GPT-5-highModel Category=Foundation Models with Tools, Evaluation Protocol=gpt-4.1-2025-04-14 judge2026.02 | 76.4 | |
| MiroThinker-v1.0Model Category=Inference Frameworks, Backbone=30B, Evaluation Protocol=gpt-4.1-2025-04-14 judge2026.02 | 73.5 | |
| ReThinkerModel Category=Inference Frameworks, Backbone=OpenPangu-72B, Evaluation Protocol=gpt-4.1-2025-04-14 judge2026.02 | 72.8 | |
| GLM-4.6Model Category=Foundation Models with Tools, Evaluation Protocol=gpt-4.1-2025-04-14 judge2026.02 | 71.9 | |
| Claude-4.5-SonnetModel Category=Foundation Models with Tools, Evaluation Protocol=gpt-4.1-2025-04-14 judge2026.02 | 71.2 | |
| Tongyi DeepResearchModel Category=Inference Frameworks, Backbone=30B-A3B, Evaluation Protocol=gpt-4.1-2025-04-14 judge2026.02 | 70.9 | |
| OpenAI DeepResearchModel Category=Inference Frameworks, Evaluation Protocol=gpt-4.1-2025-04-14 judge2026.02 | 67.4 | |
| DeepSeek-V3.2Model Category=Foundation Models with Tools, Evaluation Protocol=gpt-4.1-2025-04-14 judge2026.02 | 63.5 | |
| Kimi K2Model Category=Foundation Models with Tools, Evaluation Protocol=gpt-4.1-2025-04-14 judge2026.02 | 57.7 | |
| WebExplorerModel Category=Inference Frameworks, Evaluation Protocol=gpt-4.1-2025-04-14 judge2026.02 | 50 |