Maze navigation on Maze 100 held-out mazes
52.6Best Success Rate @ 3Max-at-K
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Max-at-Kbackbone=Qwen3-4B, prompting strategy=single-answer prompt, training algorithm=Max-at-K2026.05 | 52.6 | 55.2 | 56.8 | 57.7 | 0.671 | |
| VPObackbone=Qwen3-4B, prompting strategy=multi-answer prompt, training algorithm=Vector Policy Optimization2026.05 | 51.2 | 56.4 | 59.1 | 59.3 | 1.006 | |
| GRPObackbone=Qwen3-4B, prompting strategy=single-answer prompt, training algorithm=GRPO2026.05 | 43.2 | 43.2 | 43.2 | 43.2 | 0.003 | |
| Multi-RLVRbackbone=Qwen3-4B, prompting strategy=multi-answer prompt, training algorithm=Multi-RLVR2026.05 | 42 | 43 | 43.5 | 43.6 | 0.187 | |
| Random-wbackbone=Qwen3-4B, prompting strategy=single-answer prompt, training algorithm=Random-w2026.05 | 41.4 | 42 | 42.5 | 43.2 | 0.09 | |
| MaxRLbackbone=Qwen3-4B, prompting strategy=single-answer prompt, training algorithm=MaxRL2026.05 | 41.4 | 42.6 | 44.1 | 46.4 | 0.206 | |
| Qwen3-4B (single-answer prompt)backbone=Qwen3-4B, prompting strategy=single-answer prompt2026.05 | 34.1 | 45.7 | 59.6 | 71.4 | 0.905 | |
| Qwen3-4B (multi-answer prompt)backbone=Qwen3-4B, prompting strategy=multi-answer prompt2026.05 | 6.6 | 10.3 | 17.8 | 33.4 | 0.176 |