Agentic Task Success on ScienceWorld full (val)
26.7Success RateTREK (DeepSeek-V4)
Evaluation Results
| Method | Links | |
|---|---|---|
| TREK (DeepSeek-V4)Model=Qwen2.5-7B-Instruct, Evaluation Protocol=avg@1, Proposal Source=External DeepSeek-V4 verified rollouts, Consolidation Objective=Forward-KL2026.07 | 26.7 | |
| TREK (self-context)Model=Qwen2.5-7B-Instruct, Evaluation Protocol=avg@1, Proposal Source=Self-context with failure-lesson memory, Consolidation Objective=Forward-KL2026.07 | 23.4 | |
| OPD (self-context)Model=Qwen2.5-7B-Instruct, Evaluation Protocol=avg@1, Proposal Source=Self-context with failure-lesson memory, Consolidation Objective=On-policy distillation2026.07 | 21.6 | |
| GRPOModel=Qwen2.5-7B-Instruct, Evaluation Protocol=avg@1, Proposal Source=N/A2026.07 | 12.5 |