Agentic Coding on SWE-bench Verified
87.6Percentage ResolvedClaude Opus-4.7
Evaluation Results
| Method | Links | |
|---|---|---|
| Claude Opus-4.72026.06 | 87.6 | |
| Claude-Opus-4.62026.05 | 80.8 | |
| Gemini-3.1-Pro2026.05 | 80.6 | |
| DS-V4-Pro2026.06 | 80.6 | |
| OpenAI GPT-5.42026.06 | 80.6 | |
| Gemini 3.1-Pro2026.06 | 80.6 | |
| Kimi-K2.62026.06 | 80.2 | |
| OpenAI-GPT-5.2-Thinking2026.05 | 80 | |
| Claude-Sonnet-4.5Model Type=Proprietary, OpenHands scaffold=true, Maximum iterations=1502025.12 | 77.2 | |
| GPT-5Model Type=Proprietary, OpenHands scaffold=true, Maximum iterations=1502025.12 | 74.9 | |
| Ring-2.6-1Tvariant=high2026.06 | 74 | |
| Kimi-K2-thinkingModel Type=Open Source, Parameter Scale=> 100B, OpenHands scaffold=true, Maximum iterations=1502025.12 | 71.3 | |
| DeepSeek-V3.1-Nex-N1Model Type=Open Source, Parameter Scale=> 100B, OpenHands scaffold=true, Maximum iterations=1502025.12 | 70.6 | |
| Minimax-M2Model Type=Open Source, Parameter Scale=> 100B, OpenHands scaffold=true, Maximum iterations=1502025.12 | 69.4 | |
| GLM-4.6Model Type=Open Source, Parameter Scale=> 100B, OpenHands scaffold=true, Maximum iterations=1502025.12 | 68 | |
| DeepSeek-V3.1Model Type=Open Source, Parameter Scale=> 100B, OpenHands scaffold=true, Maximum iterations=1502025.12 | 66 | |
| Gemini-2.5-proModel Type=Proprietary, OpenHands scaffold=true, Maximum iterations=1502025.12 | 59.6 | |
| REAPModel=Kimi-K2-Instruct-W4A16, Compression=25%2025.10 | 58 | |
| EANModel=Kimi-K2-Instruct-W4A16, Compression=50%2025.10 | 57.6 | |
| REAPModel=Kimi-K2-Instruct-W4A16, Compression=50%2025.10 | 57.6 | |
| EANModel=Kimi-K2-Instruct-W4A16, Compression=25%2025.10 | 56.2 | |
| BaselineModel=Kimi-K2-Instruct-W4A16, Compression=None2025.10 | 55.4 | |
| SWE-World-32B-RLParams=32B, Training=SFT + RL2026.05 | 55 | |
| BaselineModel=Qwen3-Coder-480B-A35B-Instruct-FP8, Compression=None2025.10 | 54 | |
| REAPModel=Qwen3-Coder-480B-A35B-Instruct-FP8, Compression=25%2025.10 | 54 | |
| EANModel=Qwen3-Coder-480B-A35B-Instruct-FP8, Compression=50%2025.10 | 53.6 | |
| EANModel=Qwen3-Coder-480B-A35B-Instruct-FP8, Compression=25%2025.10 | 53.4 | |
| SWE-Lego-Qwen3-32BParams=32B, Training=SFT2026.05 | 52.6 | |
| REAPModel=Qwen3-Coder-480B-A35B-Instruct-FP8, Compression=50%2025.10 | 52.2 | |
| M2A-Agent-8BParams=8B, Training=Training-free, Avg. reasoning length=327.4, Avg. Step=178.02026.05 | 51.2 | |
| Qwen3-32B-Nex-N1Model Type=Open Source, Parameter Scale=< 100B, OpenHands scaffold=true, Maximum iterations=1502025.12 | 50.5 | |
| Kimi-Dev-72BParams=72B, Training=SFT + RL2026.05 | 48.6 | |
| Task ArithmeticParams=8B, Training=Training-free, Avg. reasoning length=836.9, Avg. Step=87.12026.05 | 47.6 | |
| SLERPParams=8B, Training=Training-free, Avg. reasoning length=747.4, Avg. Step=87.62026.05 | 47.2 | |
| Agent-8BParams=8B, Training=SFT, Avg. reasoning length=253.3, Avg. Step=175.32026.05 | 44 | |
| RAIN-MergingParams=8B, Training=Training-free, Avg. reasoning length=262.7, Avg. Step=170.42026.05 | 43.2 | |
| SWE-Lego-Qwen3-8BParams=8B, Training=SFT2026.05 | 42.2 | |
| DeepSWE-32B-PreviewParams=32B, Training=SFT + RL2026.05 | 42.2 | |
| Multi-Task-8BParams=8B, Training=SFT, Avg. reasoning length=347.4, Avg. Step=150.12026.05 | 41.1 | |
| Klear-Agent-8B-RLParams=8B, Training=SFT + RL2026.05 | 40.4 | |
| TIES-MergingParams=8B, Training=Training-free, Avg. reasoning length=494.6, Avg. Step=112.12026.05 | 39 | |
| FrequencyModel=Qwen3-Coder-480B-A35B-Instruct-FP8, Compression=25%2025.10 | 37.8 | |
| SERA-8BParams=8B, Training=SFT2026.05 | 37.1 | |
| CONTEXTRLTraining Source=Klear-AgentForge-8B, Training Strategy=CONTEXTRL2026.06 | 30.2 | |
| Qwen3-30B-A3B-Nex-N1Model Type=Open Source, Parameter Scale=< 100B, OpenHands scaffold=true, Maximum iterations=1502025.12 | 29.7 | |
| Qwen3-Coder-30BTraining Source=Off-the-shelf reference models2026.06 | 28.8 | |
| RL baselineTraining Source=Klear-AgentForge-8B, Training Strategy=RL baseline2026.06 | 28 | |
| Base modelTraining Source=Klear-AgentForge-8B, Training Strategy=Base model2026.06 | 26.6 | |
| SWE-AGILEParams=8B, Training=SFT + RL2026.05 | 24.1 | |
| SWE-Dev-7BParams=7B, Training=SFT + RL2026.05 | 23.4 | |
| SWE-Mirror-LM-7BParams=7B, Training=SFT2026.05 | 22.8 | |
| DAREParams=8B, Training=Training-free, Avg. reasoning length=1219.1, Avg. Step=68.62026.05 | 22 | |
| InternLM3-8B-Nex-N1Model Type=Open Source, Parameter Scale=< 100B, OpenHands scaffold=true, Maximum iterations=1502025.12 | 20.3 | |
| ALIVE-OracleBackbone=Qwen3-30B-Instruct2026.02 | 17.6 | |
| ALIVE-SelfBackbone=Qwen3-30B-Instruct2026.02 | 17.2 | |
| SWE-agent-LM-7BParams=7B, Training=SFT2026.05 | 15.2 | |
| GRPO (Scalar Reward)Backbone=Qwen3-30B-Instruct2026.02 | 14.8 | |
| FCP (Verbal Only)Backbone=Qwen3-30B-Instruct2026.02 | 14 | |
| SFTBackbone=Qwen3-30B-Instruct2026.02 | 13.6 | |
| Qwen3-32BModel Type=Open Source, Parameter Scale=< 100B, OpenHands scaffold=true, Maximum iterations=1502025.12 | 12.9 | |
| Base ModelBackbone=Qwen3-30B-Instruct2026.02 | 11.8 | |
| Qwen3-30B-A3BModel Type=Open Source, Parameter Scale=< 100B, OpenHands scaffold=true, Maximum iterations=1502025.12 | 9.6 | |
| Qwen3-14BTraining Source=Off-the-shelf reference models2026.06 | 8.4 | |
| Qwen3-32BTraining Source=Off-the-shelf reference models2026.06 | 8.4 | |
| CONTEXTRLTraining Source=Qwen3-8B, Training Strategy=CONTEXTRL2026.06 | 7 | |
| RL baselineTraining Source=Qwen3-8B, Training Strategy=RL baseline2026.06 | 6.2 | |
| Base modelTraining Source=Qwen3-8B, Training Strategy=Base model2026.06 | 5 | |
| Reasoning-8BParams=8B, Training=SFT, Avg. reasoning length=1800.7, Avg. Step=9.22026.05 | 0.2 | |
| FrequencyModel=Qwen3-Coder-480B-A35B-Instruct-FP8, Compression=50%2025.10 | 0 | |
| FrequencyModel=Kimi-K2-Instruct-W4A16, Compression=25%2025.10 | 0 | |
| FrequencyModel=Kimi-K2-Instruct-W4A16, Compression=50%2025.10 | 0 |