Instruction Following on BabyAI Goto
0.575Average Episodic RewardPoly-PPO
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Poly-PPORL Algorithm=Poly-PPO, Exploration Strategy (UCB)=None2025.09 | 0.575 | 80.2 | |
| Poly-PPO w/ UCBRL Algorithm=Poly-PPO, Exploration Strategy (UCB)=UCB2025.09 | 0.561 | 76.2 | |
| REINFORCE w/ UCBRL Algorithm=REINFORCE, Exploration Strategy (UCB)=UCB2025.09 | 0.538 | 73.4 | |
| REINFORCERL Algorithm=REINFORCE, Exploration Strategy (UCB)=None2025.09 | 0.533 | 73 | |
| PPO w/ UCBRL Algorithm=PPO, Exploration Strategy (UCB)=UCB2025.09 | 0.428 | 47.4 | |
| PPORL Algorithm=PPO, Exploration Strategy (UCB)=None2025.09 | 0.406 | 46.2 | |
| Pretrained policyRL Algorithm=Pretrained policy, Exploration Strategy (UCB)=None2025.09 | 0.246 | 34.2 |