Aerial Vision-and-Language Navigation on ANDH Unseen 1.0 (test)
12.9SPLHAA-Transformer
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| HAA-Transformerhuman attention prediction training=true2022.05 | 12.9 | 15.7 | 53.7 | |
| HAA-LSTMhuman attention prediction training=true2022.05 | 12.6 | 14.1 | 54.6 | |
| E.T.Input Modality=Vision and Language2022.05 | 11.3 | 13.3 | 51.7 | |
| E.T.Input Modality=Language-only2022.05 | 9.7 | 12.7 | 49.1 | |
| LSTM2022.05 | 9.7 | 10.8 | 40.4 | |
| E.T.Input Modality=Vision-only2022.05 | 3.2 | 3.9 | 0.2 | |
| Random2022.05 | 0.5 | 1.1 | -86.6 |