Aerial Vision-and-Language Navigation on ANDH Full Unseen 1.0 (test)
410SPLHAA-Transformer
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| HAA-Transformerhuman attention prediction training=true2022.05 | 410 | 630 | 63.2 | |
| E.T.Input Modality=Vision and Language2022.05 | 190 | 280 | 60.7 | |
| HAA-LSTMhuman attention prediction training=true2022.05 | 190 | 260 | 66.5 | |
| E.T.Input Modality=Language-only2022.05 | 180 | 220 | 58.2 | |
| LSTM2022.05 | 180 | 190 | 56.4 | |
| Random2022.05 | 20 | 10 | -158.4 | |
| E.T.Input Modality=Vision-only2022.05 | 20 | 20 | -1.6 |