Vision-and-Language Navigation on Touchdown Seen (test)
36.9TCORAR mixed model
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| ORAR mixed modelImage features=4th-to-last + pre-final2022.03 | 36.9 | 51.3 | |
| best merged2022.03 | 30.1 | 46 | |
| ORARImage features=ResNet 4th-to-last, Heading delta=false2022.03 | 29.3 | 45.3 | |
| best non-merged2022.03 | 29.1 | 44.9 | |
| ORARImage features=ResNet 4th-to-last2022.03 | 29.1 | 44.9 | |
| ORARImage features=ResNet 4th-to-last, Junction type=false2022.03 | 25.5 | 40.9 | |
| ORARImage features=ResNet pre-final2022.03 | 25.3 | 38.4 | |
| ORARImage features=ResNet 4th-to-last, Heading delta=false, Junction type=false2022.03 | 24.2 | 39.4 | |
| ARC+l2s2022.03 | 16.7 | — | |
| VLN Transformer2022.03 | 14.9 | 25.3 | |
| ARC2022.03 | 14.1 | — | |
| GA2022.03 | 11.9 | 24.9 | |
| RConcat2022.03 | 11.8 | 22.9 |