Error category prediction on TRAIL Planning and Reasoning categories (117 traces)
49.7Micro F1GPT-5
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPT-5mapping_strategy=full+partial, judge=GPT-52026.05 | 49.7 | 45.9 | |
| GPT-5mapping_strategy=full, judge=GPT-52026.05 | 46.7 | 36.8 | |
| Always top-4baseline_type=top-4 most frequent categories2026.05 | 45.9 | 19.9 | |
| OSS-120Bmapping_strategy=full+partial, judge=OSS-120B2026.05 | 42.7 | 37.4 | |
| OSS-120Bmapping_strategy=full, judge=OSS-120B2026.05 | 37.7 | 26.1 | |
| Random (GT freq)baseline_type=random predictor weighted by true category frequency2026.05 | 34.2 | 28.8 |