Planning on Sokoban
42.4p@1LAMER
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| LAMERTraining Strategy=Training with Meta-RL, Base Model=Qwen3-4B2025.12 | 42.4 | 52 | 55.9 | |
| GiGPOTraining Strategy=Training with RL, Base Model=Qwen3-4B2025.12 | 41.6 | 43.6 | 44.1 | |
| GRPOTraining Strategy=Training with RL, Base Model=Qwen3-4B2025.12 | 22.9 | 26.4 | 27 | |
| RLOOTraining Strategy=Training with RL, Base Model=Qwen3-4B2025.12 | 13.5 | 16.6 | 18.8 | |
| PPOTraining Strategy=Training with RL, Base Model=Qwen3-4B2025.12 | 12.5 | 15.4 | 16.8 | |
| ReActTraining Strategy=Prompting, Base Model=Qwen3-4B2025.12 | 7.2 | 9.6 | 12.5 | |
| Zero-shotTraining Strategy=Prompting, Base Model=Qwen3-4B2025.12 | 6.8 | 9.8 | 12.9 | |
| ReflexionTraining Strategy=Prompting, Base Model=Qwen3-4B2025.12 | 6.4 | 9.8 | 12.1 |