Agentic Reasoning and Task Execution on Controlled-domain Long-memory
0.453G (Mean Task-Rubric Score)Base
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| BasePolicy Variant=Base, Seeds=3, Rollouts per task=3, Number of tasks=20, Total trials (N)=1802026.07 | 0.453 | — | — | — | — | — | |
| Offline AW ControllerPolicy Variant=AW, Seeds=3, Rollouts per task=3, Number of tasks=20, Total trials (N)=1802026.07 | 0.451 | 0.003 | 0.011 | 0.008 | 0.449 | 10 |