Instruction Following on BabyAI Synthseq
0.361Average Episodic RewardREINFORCE w/ UCB
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| REINFORCE w/ UCBRL Algorithm=REINFORCE, Exploration Strategy (UCB)=UCB2025.09 | 0.361 | 47.8 | |
| Poly-PPORL Algorithm=Poly-PPO, Exploration Strategy (UCB)=None2025.09 | 0.341 | 47 | |
| REINFORCERL Algorithm=REINFORCE, Exploration Strategy (UCB)=None2025.09 | 0.325 | 45.4 | |
| Poly-PPO w/ UCBRL Algorithm=Poly-PPO, Exploration Strategy (UCB)=UCB2025.09 | 0.317 | 43.2 | |
| PPORL Algorithm=PPO, Exploration Strategy (UCB)=None2025.09 | 0.277 | 32.2 | |
| PPO w/ UCBRL Algorithm=PPO, Exploration Strategy (UCB)=UCB2025.09 | 0.224 | 26.2 | |
| Pretrained policyRL Algorithm=Pretrained policy, Exploration Strategy (UCB)=None2025.09 | 0.157 | 20.2 |