Autonomous Software Engineering Evaluation on Trap + SWE-bench Aggregate Lite
293Success Count (TRUE)StateFlow
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| StateFlown=1,050, Audit mode=ensemble (static then LLM judge), Wall-clock cap=600 s, Models=F2 (frontier small), M1 (mid-tier code-tuned), M2 (mid-tier reasoning), Seeds=52026.06 | 293 | 263 | 488 | 6 | 25.05 | 22.48 | |
| Reflexionn=1,050, Audit mode=ensemble (static then LLM judge), Wall-clock cap=600 s, Models=F2 (frontier small), M1 (mid-tier code-tuned), M2 (mid-tier reasoning), Seeds=52026.06 | 92 | 85 | 674 | 199 | 8.1 | 6.48 | |
| Autopilotn=1,050, Audit mode=ensemble (static then LLM judge), Wall-clock cap=600 s, Models=F2 (frontier small), M1 (mid-tier code-tuned), M2 (mid-tier reasoning), Seeds=52026.06 | 85 | 10 | 928 | 32 | 0.95 | 0.38 |