Multi-objective Reinforcement Learning on MuJoCo 8 continuous-action tasks MO-Gymnasium (aggregated)
3.25Hypervolume (HV)PDMORL
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| PDMORLMethod Type=Model-Free, Environment Interaction Steps=2M2026.05 | 3.25 | 3.25 | 3.62 | |
| GPI-PDMethod Type=Model-Based, Environment Interaction Steps=500K2026.05 | 3 | 3 | 2.25 | |
| CAPQL-MFMethod Type=Model-Free, Environment Interaction Steps=2M2026.05 | 2.75 | 2.62 | 1.75 | |
| COLAMethod Type=Model-Free, Environment Interaction Steps=2M2026.05 | 2 | 2 | 2.12 | |
| PCSAC-MFMethod Type=Model-Free, Environment Interaction Steps=2M2026.05 | 2 | 2.12 | 2.5 | |
| CAPQL-MBMethod Type=Model-Based, Environment Interaction Steps=500K2026.05 | 1.75 | 1.75 | 2 | |
| PCSAC-MBMethod Type=Model-Based, Environment Interaction Steps=500K2026.05 | 1.25 | 1.25 | 1.75 |