Reinforcement Learning on Grid-World (Episode 500)
27.7True ReturnMulti-Head
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Multi-HeadAblation=true2026.04 | 27.7 | 42.1 | 14.5 | — | |
| Baseline Q-Learning2026.04 | 19.8 | 57.8 | 16.2 | — | |
| UARD Framework2026.04 | 4 | 9.1 | 0 | 3.2 |
| Method | Links | ||||
|---|---|---|---|---|---|
| Multi-HeadAblation=true2026.04 | 27.7 | 42.1 | 14.5 | — | |
| Baseline Q-Learning2026.04 | 19.8 | 57.8 | 16.2 | — | |
| UARD Framework2026.04 | 4 | 9.1 | 0 | 3.2 |