Reinforcement Learning on Craftax symbolic version
11.87Max Reward (%)RePPO (1.2)
Evaluation Results
| Method | Links | |
|---|---|---|
| RePPO (1.2)m=1.2, Entropy bonus=false, Training timesteps=1B, Number of seeds=52026.05 | 11.87 | |
| PPO-Q + EntEntropy bonus=0.01, Training timesteps=1B, Number of seeds=52026.05 | 11.85 | |
| RND + EntEntropy bonus=0.01, Training timesteps=1B, Number of seeds=52026.05 | 11.68 | |
| PPO-V + EntEntropy bonus=0.01, Training timesteps=1B, Number of seeds=52026.05 | 11.66 | |
| RePPO (1.4)m=1.4, Entropy bonus=false, Training timesteps=1B, Number of seeds=52026.05 | 10.79 | |
| RNDEntropy bonus=false, Training timesteps=1B, Number of seeds=52026.05 | 9.62 | |
| PPO-VEntropy bonus=false, Training timesteps=1B, Number of seeds=52026.05 | 9.31 | |
| PPO-QEntropy bonus=false, Training timesteps=1B, Number of seeds=52026.05 | 9.31 |