Instruction Following on BabyAI BossLevel
96.2Success RateThought Cloning
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Thought CloningNum Parameters=82.5M, Data=1M Episodes Actions + Language2023.06 | 96.2 | — | |
| Pure BC architecture (and matched num parameters), 2x dataNum Parameters=83.9M, Data=2M Episodes Actions2023.06 | 92.7 | — | |
| Pure BC architecture (and matched num parameters)Num Parameters=83.9M, Data=1M Episodes Actions2023.06 | 91.9 | — | |
| BC w/ 2x DataNum Parameters=20.6M, Data=2M Episodes Actions2023.06 | 91.4 | — | |
| Behavioral CloningNum Parameters=20.6M, Data=1M Episodes Actions2023.06 | 91.2 | — | |
| Think Before You ActNum Parameters=124M (estimated), Data=1M Episodes Actions + Language2023.06 | 85.2 | — | |
| Thought Cloning w/o Imitating ThoughtNum Parameters=82.5M, Data=1M Episodes Actions2023.06 | 65.5 | — | |
| Poly-PPO w/ UCBRL Algorithm=Poly-PPO, Exploration Strategy (UCB)=UCB2025.09 | 46.8 | 0.379 | |
| Poly-PPORL Algorithm=Poly-PPO, Exploration Strategy (UCB)=None2025.09 | 45.2 | 0.378 | |
| PPORL Algorithm=PPO, Exploration Strategy (UCB)=None2025.09 | 38.8 | 0.336 | |
| REINFORCE w/ UCBRL Algorithm=REINFORCE, Exploration Strategy (UCB)=UCB2025.09 | 36.4 | 0.286 | |
| PPO w/ UCBRL Algorithm=PPO, Exploration Strategy (UCB)=UCB2025.09 | 35.8 | 0.31 | |
| REINFORCERL Algorithm=REINFORCE, Exploration Strategy (UCB)=None2025.09 | 33.4 | 0.266 | |
| Pretrained policyRL Algorithm=Pretrained policy, Exploration Strategy (UCB)=None2025.09 | 20.6 | 0.212 |