Reinforcement Learning on Atari Seaquest Expected Environment (train)
377.21Training RewardPPO + Naive Shield
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| PPO + Naive ShieldShield type=Naive2025.11 | 377.21 | 100 | |
| PPO + Repaired ShieldShield type=Repaired2025.11 | 261.77 | 100 | |
| PPO + Adaptive ShieldShield type=Adaptive2025.11 | 224.58 | 100 | |
| PPO + Static ShieldShield type=Static2025.11 | 224.03 | 100 | |
| PPOShield type=None2025.11 | 142.48 | 19 |