Data Prediction on MLEBench lite
100Valid ScoreQwen3-Coder
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Qwen3-Coderparameters=480B2026.01 | 100 | 9.09 | 22.72 | |
| GPT-5.1reasoning=high2026.01 | 90.91 | 22.73 | 45.45 | |
| GPT-5.1reasoning=medium2026.01 | 90.91 | 22.73 | 31.82 | |
| Claude Sonnet 42026.01 | 90.9 | 13.63 | 22.73 | |
| Kimi K2 Instruct2026.01 | 86.37 | 13.64 | 27.27 | |
| Deepseek v3.12026.01 | 86.37 | 13.64 | 27.27 | |
| Claude Sonnet 4.52026.01 | 86.36 | 22.73 | 36.36 | |
| Qwen3 Instructparameters=235B2026.01 | 81.82 | 4.55 | 13.64 | |
| GPT-5reasoning=medium2026.01 | 77.27 | 9.09 | 27.27 | |
| GPT-5.1reasoning=none2026.01 | 72.72 | 13.64 | 22.73 | |
| Qwen3 Instructparameters=4B2026.01 | 50 | 4.55 | 9.09 | |
| Qwen2.5 Instructparameters=7B2026.01 | 0 | 0 | 0 |