General AI Assistant Tasks on GAIA (avg@8)
88.5Avg@8 ScoreMiroThinker-H1
Evaluation Results
| Method | Links | |
|---|---|---|
| MiroThinker-H1Agent Style=ReAct, Max interaction turns=200, Judge LLM=gpt-4.1-2025-04-14, Max episode retries=5, Context management retention=52026.03 | 88.5 | |
| MiroThinker-1.7Agent Style=ReAct, Max interaction turns=200, Judge LLM=gpt-4.1-2025-04-14, Max episode retries=5, Context management retention=52026.03 | 82.7 | |
| MiroThinker-v1.0-72BParameters=72B2025.11 | 81.9 | |
| MiroThinker-1.7-miniAgent Style=ReAct, Max interaction turns=200, Judge LLM=gpt-4.1-2025-04-14, Max episode retries=5, Context management retention=52026.03 | 80.3 | |
| OpenAI-GPT-5Agent Style=ReAct, Max interaction turns=200, Judge LLM=gpt-4.1-2025-04-14, Max episode retries=5, Context management retention=52026.03 | 76.4 | |
| OpenAI-GPT-5-highType=Foundation Models with Tools2025.11 | 76.4 | |
| Minimax-M2Type=Foundation Models with Tools2025.11 | 75.7 | |
| MiroThinker-v1.0-30BParameters=30B2025.11 | 73.5 | |
| GLM-4.6Type=Foundation Models with Tools2025.11 | 71.9 | |
| Claude-4.5-SonnetType=Foundation Models with Tools2025.11 | 71.2 | |
| Tongyi-DeepResearch-30BAgent Style=ReAct, Max interaction turns=200, Judge LLM=gpt-4.1-2025-04-14, Max episode retries=5, Context management retention=52026.03 | 70.9 | |
| Tongyi-DeepResearch-30BType=Research Agents2025.11 | 70.9 | |
| Claude-4-SonnetType=Foundation Models with Tools2025.11 | 68.3 | |
| OpenAI DeepResearchType=Research Agents2025.11 | 67.4 | |
| MiroThinker-v1.0-8BParameters=8B2025.11 | 66.4 | |
| SFR-DeepResearch-20BType=Research Agents2025.11 | 66 | |
| DeepSeek-V3.2Type=Foundation Models with Tools2025.11 | 63.5 | |
| DeepSeek-V3.1Type=Foundation Models with Tools2025.11 | 63.1 | |
| Kimi-K2-0905Type=Foundation Models with Tools2025.11 | 60.2 | |
| DeepMiner-32B-RLType=Research Agents2025.11 | 58.7 | |
| AFM-32B-RLType=Research Agents2025.11 | 55.3 | |
| WebExplorer-8B-RLType=Research Agents2025.11 | 50 |