Software Engineering on SWE-bench Verified (Pass@1 and Time Metrics)
72Pass@1GPT-5 mini (2025-08-07)
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| GPT-5 mini (2025-08-07)Peak Len.=400k, Evaluation Mode=Single (×1), Benchmark Split=Full2026.07 | 72 | — | — | — | — | — | — | |
| GLM-4.5-Air (106B-A30B) + CompactionRLPeak Len.=80k, Evaluation Mode=Compacted (×4), Scaffold=Terminus-KIRA, Benchmark Split=200-instance subset2026.07 | 66.8 | — | — | — | — | — | — | |
| Qwen3-Coder-480B-A35B-InstructPeak Len.=256k, Evaluation Mode=Single (×1), Benchmark Split=Full2026.07 | 66.5 | — | — | — | — | — | — | |
| GLM-4.5-Air (106B-A30B) + RL (w/o compaction)Peak Len.=80k, Evaluation Mode=Compacted (×4), Scaffold=Terminus-KIRA, Benchmark Split=200-instance subset2026.07 | 62.5 | — | — | — | — | — | — | |
| gpt-oss-120bPeak Len.=128k, Evaluation Mode=Single (×1), Benchmark Split=Full2026.07 | 62 | — | — | — | — | — | — | |
| GLM-4.5-Air (106B-A30B)Peak Len.=80k, Evaluation Mode=Compacted (×4), Scaffold=Terminus-KIRA, Benchmark Split=200-instance subset2026.07 | 59.8 | — | — | — | — | — | — | |
| GLM-4.5-Air (106B-A30B) + RL (w/o compaction)Peak Len.=80k, Evaluation Mode=Single (×1), Scaffold=Terminus-KIRA, Benchmark Split=200-instance subset2026.07 | 58.3 | — | — | — | — | — | — | |
| Qwen3.5-35B-A3BPeak Len.=64k, Evaluation Mode=Compacted (×4), Benchmark Split=200-instance subset2026.07 | 58 | — | — | — | — | — | — | |
| GLM-4.5-Air (106B-A30B)Peak Len.=80k, Evaluation Mode=Single (×1), Scaffold=Terminus-KIRA, Benchmark Split=200-instance subset2026.07 | 57.8 | — | — | — | — | — | — | |
| GLM-4.5-Air (106B-A30B) + CompactionRLPeak Len.=80k, Evaluation Mode=Single (×1), Scaffold=Terminus-KIRA, Benchmark Split=200-instance subset2026.07 | 57.3 | — | — | — | — | — | — | |
| GLM-4.7-Flash (30B-A3B) + CompactionRLPeak Len.=64k, Evaluation Mode=Compacted (×4), Scaffold=Terminus-KIRA, Benchmark Split=200-instance subset2026.07 | 56 | — | — | — | — | — | — | |
| Qwen3-Coder-30B-A3B-InstructPeak Len.=256k, Evaluation Mode=Single (×1), Benchmark Split=Full2026.07 | 51.9 | — | — | — | — | — | — | |
| GLM-4.7-Flash (30B-A3B)Peak Len.=64k, Evaluation Mode=Compacted (×4), Scaffold=Terminus-KIRA, Benchmark Split=200-instance subset2026.07 | 50.5 | — | — | — | — | — | — | |
| GLM-4.7-Flash (30B-A3B) + RL (w/o compaction)Peak Len.=64k, Evaluation Mode=Single (×1), Scaffold=Terminus-KIRA, Benchmark Split=200-instance subset2026.07 | 50 | — | — | — | — | — | — | |
| GLM-4.7-Flash (30B-A3B) + RL (w/o compaction)Peak Len.=64k, Evaluation Mode=Compacted (×4), Scaffold=Terminus-KIRA, Benchmark Split=200-instance subset2026.07 | 48 | — | — | — | — | — | — | |
| GLM-4.7-Flash (30B-A3B)Peak Len.=64k, Evaluation Mode=Single (×1), Scaffold=Terminus-KIRA, Benchmark Split=200-instance subset2026.07 | 47.5 | — | — | — | — | — | — | |
| Qwen3-235B-A22B-Instruct-2507Peak Len.=256k, Evaluation Mode=Single (×1), Benchmark Split=Full2026.07 | 45.2 | — | — | — | — | — | — | |
| GLM-4.7-Flash (30B-A3B) + CompactionRLPeak Len.=64k, Evaluation Mode=Single (×1), Scaffold=Terminus-KIRA, Benchmark Split=200-instance subset2026.07 | 43.7 | — | — | — | — | — | — | |
| PS-adaModel=Qwen3-32B2026.05 | 42.2 | 172.9 | 168.1 | 290 | 282 | 2,146 | 39.8 | |
| BaselineModel=Qwen3-32B2026.05 | 36.8 | 268 | 258.2 | 410 | 395 | 2,353 | 30.9 | |
| PS-adaModel=Qwen3-14B2026.05 | 29.5 | 66.3 | 80.4 | 170 | 206 | 1,405 | 36.7 | |
| Qwen3-30B-A3B-Instruct-2507Peak Len.=256k, Evaluation Mode=Single (×1), Benchmark Split=Full2026.07 | 25.2 | — | — | — | — | — | — | |
| BaselineModel=Qwen3-14B2026.05 | 24.7 | 133.1 | 138.4 | 300 | 312 | 1,597 | 25.9 |