Agent Success Prediction on SWE-bench Pro (held-out)
0.909AUC-ROCOracle
Evaluation Results
| Method | Links | |
|---|---|---|
| Oracledescription=standard IRT with agent features trained on all data2026.04 | 0.909 | |
| LLM-as-a-judgefeatures=LLM-as-a-judge2026.04 | 0.696 | |
| Combinedfeatures=Combined (Embedding + LLM-as-a-judge)2026.04 | 0.677 | |
| Embeddingfeatures=Embedding2026.04 | 0.668 | |
| Baselinedescription=predicts the LLM's empirical success rate ignoring the scaffold2026.04 | 0.571 |