Failure Attribution on Magentic
81.2Agent AccuracyOur Baseline
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Our BaselineBackbone LLM=GPT-52026.02 | 81.2 | 56.3 | |
| Who&When*Backbone LLM=GPT-5, Prompt-modified=true, Variant=Best-performing2026.02 | 6.2 | 56.3 |
| Method | Links | ||
|---|---|---|---|
| Our BaselineBackbone LLM=GPT-52026.02 | 81.2 | 56.3 | |
| Who&When*Backbone LLM=GPT-5, Prompt-modified=true, Variant=Best-performing2026.02 | 6.2 | 56.3 |