Software Engineering on SWE-Bench (val)
28.8AccClaude 3.7 Sonnet
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Claude 3.7 SonnetModel Size=~100B+, Training Strategy=Baseline Model2026.02 | 28.8 | 77.5 | |
| DeepSeek-R1Model Size=37B, Training Strategy=Baseline Model2026.02 | 8.8 | 30 | |
| Rubric-Augmented ClassifierModel Size=4B, Training Strategy=RL w/o GT2026.02 | 2.5 | 20 | |
| Claude 3.5 HaikuModel Size=~20B+, Training Strategy=Baseline Model2026.02 | 1.3 | 2.5 | |
| Mistral-7BModel Size=7B, Training Strategy=Baseline Model2026.02 | 0 | 15 | |
| Qwen3-4B (No RL)Model Size=4B, Training Strategy=Baseline (RL w/o GT)2026.02 | 0 | 0 | |
| Baseline ClassifierModel Size=4B, Training Strategy=RL w/o GT2026.02 | 0 | 2.5 |