Agentic Task Success on ALFWorld (val)
82.8Success RateTREK (DeepSeek-V4)
Evaluation Results
| Method | Links | |
|---|---|---|
| TREK (DeepSeek-V4)Model=Qwen2.5-7B-Instruct, Evaluation Protocol=avg@1, Proposal Source=External DeepSeek-V4 verified rollouts, Consolidation Objective=Forward-KL2026.07 | 82.8 | |
| TREK (self-context)Model=Qwen2.5-7B-Instruct, Evaluation Protocol=avg@1, Proposal Source=Self-context with failure-lesson memory, Consolidation Objective=Forward-KL2026.07 | 80.4 | |
| OPD (self-context)Model=Qwen2.5-7B-Instruct, Evaluation Protocol=avg@1, Proposal Source=Self-context with failure-lesson memory, Consolidation Objective=On-policy distillation2026.07 | 78.3 | |
| GRPOModel=Qwen2.5-7B-Instruct, Evaluation Protocol=avg@1, Proposal Source=N/A2026.07 | 75.8 |