Multi-Objective Reinforcement Learning on Snake
1.168Hypervolume (HV)PPO
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| PPO2026.05 | 1.168 | 0.5 | 1.99 | 0.783 | — | — | — | |
| MOPPOConditioned=Yes2026.05 | 1.0214 | 8.8 | 1.971 | 0.783 | -0.16 | -0.655 | 0.959 | |
| MOPPO (no cond.)Conditioned=No2026.05 | 0.3098 | 0.4 | 1.762 | 0.76 | — | — | — |