Agentic Coding on SWE-Bench Multilingual
71.7AccuracyMiMo-V2-Flash
Evaluation Results
| Method | Links | |
|---|---|---|
| MiMo-V2-FlashVariant=Flash2026.01 | 71.7 | |
| DeepSeek-V3.2 ThinkingThinking Mode=true2026.01 | 70.2 | |
| Claude Sonnet 4.5Variant=Sonnet 4.52026.01 | 68 | |
| Qwen3.6# Total Params=35B, # Active Params=3B2026.05 | 67.2 | |
| Kimi-K2 ThinkingThinking Mode=true2026.01 | 61.1 | |
| Qwen3.5# Total Params=35B, # Active Params=3B2026.05 | 60.3 | |
| LAGUNA XS.2# Total Params=33.4B, # Active Params=3B2026.05 | 57.7 | |
| Devstral Small 2# Total Params=24B, # Active Params=24B2026.05 | 55.7 | |
| GPT-5 HighVariant=High2026.01 | 55.3 | |
| Gemma 4# Total Params=31B, # Active Params=31B2026.05 | 51.7 | |
| LongCat-Flash-LiteArchitecture=MoE + NE, # Total Params=68.5B, # Activated Params=2.9B~4.5B2026.01 | 38.1 | |
| Kimi-Linear-48B-A3BArchitecture=MoE, # Total Params=48B, # Activated Params=3B2026.01 | 37.2 | |
| Qwen3-Next-80B-A3B-InstructArchitecture=MoE, # Total Params=80B, # Activated Params=3B2026.01 | 31.3 |