Automated Machine Learning on MLE-Bench
96.89Valid Submission RateFM Agent
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| FM AgentBase Model=Gemini-2.5-pro2025.10 | 96.89 | — | 51.56 | 8.44 | 12.44 | 22.67 | 43.56 | — | — | |
| InternAgentBase Model=deepseek-r12025.10 | 96.44 | — | 48.44 | 7.11 | 10.67 | 18.67 | 36.44 | — | — | |
| ML-MasterBase Model=deepseek-r12025.10 | 93.3 | — | 44.9 | 4.4 | 7.6 | 17.3 | 29.3 | — | — | |
| Neo multi-agentBase Model=undisclosed2025.10 | 85.78 | — | 40 | 10.22 | 10.22 | 13.78 | 34.22 | — | — | |
| iML2026.02 | 85 | 85 | 80 | 0 | 10 | 35 | 45 | 0.77 | — | |
| AIDEBase Model=o1-preview2025.10 | 82.8 | — | 29.4 | 3.4 | 4.1 | 9.4 | 16.9 | — | — | |
| AutoGluon2026.02 | 80 | 80 | 55 | 0 | 10 | 30 | 40 | 0.66 | — | |
| MLZero2026.02 | 60 | 70 | 45 | 5 | 5 | 20 | 30 | 0.54 | — | |
| Operand ensembleBase Model=gpt-5 (low verbosity/effort)†2025.10 | 55.11 | — | 40.89 | 20.89 | 7.11 | 11.56 | 39.56 | — | — | |
| MLE-STAR2026.02 | 55 | 65 | 40 | 0 | 0 | 15 | 15 | 0.43 | — | |
| R&D-AgentBase Model=gpt-52025.10 | 53.33 | — | 40.44 | 6.67 | 12 | 16.44 | 35.11 | — | — | |
| OpenHandsBase Model=gpt-4o-2024-08-062025.10 | 52 | — | 7.1 | 0.4 | 1.3 | 2.7 | 5.1 | — | — | |
| MLABBase Model=gpt-4o-2024-08-062025.10 | 44.3 | — | 1.9 | 0 | 0 | 0.8 | 1.3 | — | — | |
| AutoML-Agent2026.02 | 40 | 50 | 25 | 0 | 0 | 5 | 5 | 0.22 | — | |
| AIDE2025.08 | — | — | — | — | — | — | — | — | 32.5 | |
| KompeteAI2025.08 | — | — | — | — | — | — | — | — | 47.6 | |
| KompeteAI (w/o merging)merging=disabled2025.08 | — | — | — | — | — | — | — | — | 38.5 | |
| KompeteAI (w/o RAG)RAG=disabled2025.08 | — | — | — | — | — | — | — | — | 45 | |
| KompeteAI (w/o scoring model)scoring model=disabled2025.08 | — | — | — | — | — | — | — | — | 41.6 | |
| ML-Master2025.08 | — | — | — | — | — | — | — | — | 41.2 | |
| MLE-STAR2025.08 | — | — | — | — | — | — | — | — | 37.3 | |
| RD-Agent2025.08 | — | — | — | — | — | — | — | — | 35 |