Logical deduction on MineSweeper
52p@1GiGPO
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GiGPOTraining Strategy=Training with RL, Base Model=Qwen3-4B2025.12 | 52 | 54.9 | 55.1 | |
| RLOOTraining Strategy=Training with RL, Base Model=Qwen3-4B2025.12 | 48.8 | 51.2 | 51.6 | |
| LAMERTraining Strategy=Training with Meta-RL, Base Model=Qwen3-4B2025.12 | 44.1 | 66.4 | 74.4 | |
| GRPOTraining Strategy=Training with RL, Base Model=Qwen3-4B2025.12 | 36.3 | 40 | 40.4 | |
| PPOTraining Strategy=Training with RL, Base Model=Qwen3-4B2025.12 | 29.7 | 34.2 | 35.5 | |
| ReActTraining Strategy=Prompting, Base Model=Qwen3-4B2025.12 | 6.3 | 7 | 10.9 | |
| ReflexionTraining Strategy=Prompting, Base Model=Qwen3-4B2025.12 | 5.5 | 7.2 | 9.8 | |
| Zero-shotTraining Strategy=Prompting, Base Model=Qwen3-4B2025.12 | 4.5 | 6.6 | 8.6 |