Autonomous Agent Planning and Execution on Task Performance Environment (test)
0.387RewardSFT+RL
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SFT+RLSize=8B, Strategy=Dynamic2025.09 | 0.387 | 1,714.3 | |
| Zero-shotSize=70B, Strategy=4 steps2025.09 | 0.379 | 11,510.9 | |
| Zero-shotSize=70B, Strategy=Always2025.09 | 0.349 | 38,836.2 | |
| Zero-shotSize=70B, Strategy=Never2025.09 | 0.343 | 559.6 | |
| SFTSize=8B, Strategy=Dynamic2025.09 | 0.343 | 1,869.9 | |
| SFT+RLSize=8B, Strategy=Never2025.09 | 0.298 | 878 | |
| SFTSize=8B, Strategy=Never2025.09 | 0.286 | 991.5 | |
| Base+RLSize=8B, Strategy=Never2025.09 | 0.274 | 505.1 | |
| Base+RLSize=8B, Strategy=Dynamic2025.09 | 0.21 | 10,818.7 |