LLM Agent Optimization on Real-API 300 turns (evaluation)
263.5Token UsageTask-aware raw
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Task-aware rawRole=low-cost2026.06 | 263.5 | 82.4 | 3.127 | |
| scalar-state repaired controllerRole=sweet spot, Shorthand=Scalar+R2026.06 | 581.1 | 89.4 | 1.694 | |
| ConservativeRole=baseline2026.06 | 703.8 | 89.9 | 1.278 | |
| task-aware repaired controllerRole=quality, Shorthand=Task-aware+R2026.06 | 817.2 | 90.4 | 1.106 | |
| MiddleRole=dominated2026.06 | 933.7 | 88.8 | 0.951 |