Agent success prediction on GSO (held-out)
91.1AUC-ROCOracle
Evaluation Results
| Method | Links | |
|---|---|---|
| Oracledescription=standard IRT with agent features trained on all data2026.04 | 91.1 | |
| LLM-as-a-judgefeatures=LLM-as-a-judge2026.04 | 73.5 | |
| Embeddingfeatures=Embedding2026.04 | 72 | |
| Combinedfeatures=Combined (Embedding + LLM-as-a-judge)2026.04 | 71.9 | |
| Baselinedescription=predicts the LLM's empirical success rate ignoring the scaffold2026.04 | 63.7 |