Multi-objective Reinforcement Learning on Average-reward MDP
1Sample Complexity BoundActor-Critic
Evaluation Results
| Method | Links | |
|---|---|---|
| Actor-CriticParametrized Policy=true, Constraints=false, Concave Scalarization=false, Mixing-time Knowledge=No2026.06 | 1 | |
| Primal-Dual ACParametrized Policy=true, Constraints=true, Concave Scalarization=false, Mixing-time Knowledge=Yes2026.06 | 1 | |
| Model-basedParametrized Policy=false, Constraints=false, Concave Scalarization=true, Mixing-time Knowledge=No2026.06 | 1 | |
| Model-based PDParametrized Policy=false, Constraints=true, Concave Scalarization=true, Mixing-time Knowledge=No2026.06 | 1 | |
| MO-PDNACParametrized Policy=true, Constraints=true, Concave Scalarization=true, Mixing-time Knowledge=No2026.06 | 1 |