General Assistant on GAIA (Pass@1 L1/L2/L3)
88.68Pass@1 (L1)Mem2Evolve
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Mem2EvolveMethod Category=Ours, Backbone=GPT-5-Chat2026.04 | 88.68 | 82.56 | 57.69 | 76.31 | |
| AlitaMethod Category=Capability-Centric Evolving, Backbone=GPT-5-Chat2026.04 | 81.13 | 75.58 | 46.15 | 72.73 | |
| OpenAI-DeepResearchMethod Category=Naive-Large Language Model, Source=Original paper2026.04 | 74.29 | 69.06 | 47.6 | 67.36 | |
| AutoAgentsMethod Category=Capability-Centric Evolving, Backbone=GPT-5-Chat2026.04 | 35.85 | 24.42 | 19.23 | 26.5 | |
| DSPyMethod Category=Experience-Centric Evolving, Backbone=GPT-5-Chat2026.04 | 30.19 | 15.12 | 11.54 | 18.95 | |
| AgentVerseMethod Category=Capability-Centric Evolving, Backbone=GPT-5-Chat2026.04 | 30.19 | 16.28 | 19.23 | 21.9 | |
| SwarmAgenticMethod Category=Capability-Centric Evolving, Backbone=GPT-5-Chat2026.04 | 28.3 | 18.6 | 13.46 | 20.4 | |
| GPT-5-Chat (ReAct)Method Category=Naive-Large Language Model, Backbone=GPT-5-Chat, Prompting Strategy=ReAct2026.04 | 26.42 | 17.44 | 11.54 | 18.47 | |
| AFLOWMethod Category=Experience-Centric Evolving, Backbone=GPT-5-Chat2026.04 | 26.42 | 17.44 | 15.38 | 19.75 | |
| GPT-5-Chat (CoT)Method Category=Naive-Large Language Model, Backbone=GPT-5-Chat, Prompting Strategy=CoT2026.04 | 24.53 | 17.44 | 11.54 | 17.84 | |
| DyLANMethod Category=Experience-Centric Evolving, Backbone=GPT-5-Chat2026.04 | 24.53 | 19.78 | 11.54 | 18.62 | |
| EvoAgentMethod Category=Experience-Centric Evolving, Backbone=GPT-5-Chat2026.04 | 22.64 | 19.78 | 11.54 | 17.99 | |
| GPT-5-Chat (Direct)Method Category=Naive-Large Language Model, Backbone=GPT-5-Chat, Prompting Strategy=Direct2026.04 | 16.98 | 12.79 | 7.69 | 12.49 |