Reinforcement Learning on Sokoban
0.87RewardERL
Evaluation Results
| Method | Links | |
|---|---|---|
| ERLBackbone=Qwen3-4B-Instruct-2507, Method variant=Full2026.02 | 0.87 | |
| ERL w/o Mem.Backbone=Qwen3-4B-Instruct-2507, Method variant=Without memory reuse2026.02 | 0.87 | |
| ERL w/o Refl.Backbone=Qwen3-4B-Instruct-2507, Method variant=Without structured reflection2026.02 | 0.59 | |
| ERL w/o Mem.Backbone=Olmo3-7B-Instruct, Method variant=Without memory reuse2026.02 | 0.24 | |
| ERLBackbone=Olmo3-7B-Instruct, Method variant=Full2026.02 | 0.2 | |
| RLVRBackbone=Qwen3-4B-Instruct-25072026.02 | 0.06 | |
| ERL w/o Refl.Backbone=Olmo3-7B-Instruct, Method variant=Without structured reflection2026.02 | 0.06 | |
| RLVRBackbone=Olmo3-7B-Instruct2026.02 | 0.04 |