Vision-and-Language Navigation on Room-to-Room (R2R) (val unseen)
2.29NEHAMT
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| HAMTVLN specialist category=Pretraining-based VLN specialists2026.03 | 2.29 | 66 | — | 61 | |
| DUETVLN specialist category=Pretraining-based VLN specialists2026.03 | 3.31 | 72 | 81 | 60 | |
| NaviLLMVLN specialist category=Pretraining-based VLN specialists2026.03 | 3.51 | 67 | — | 59 | |
| PREVALENTVLN specialist category=Pretraining-based VLN specialists2026.03 | 4.71 | 58 | — | 53 | |
| NavGPTVLN specialist category=Zero-shot VLN with Foundation Models, Foundation Model=Q3, Re-implementation status=true2026.03 | 4.82 | 47 | 57.5 | 38.4 | |
| ProFocusVLN specialist category=Zero-shot VLN with Foundation Models, Foundation Models=Q3+Q3VL, Re-implementation status=true2026.03 | 4.92 | 52.5 | 65 | 39.8 | |
| MapGPTVLN specialist category=Zero-shot VLN with Foundation Models, Foundation Model=GLM, Re-implementation status=true2026.03 | 5 | 41.4 | 70.7 | 30.8 | |
| ProFocusVLN specialist category=Zero-shot VLN with Foundation Models, Foundation Models=DS3+GLM, Re-implementation status=true2026.03 | 5.21 | 50 | 63 | 41.2 | |
| EnvDropVLN specialist category=Training-based VLN specialists2026.03 | 5.22 | 52 | — | 48 | |
| MSNavVLN specialist category=Zero-shot VLN with Foundation Models, Foundation Model=GPT-4o2026.03 | 5.24 | 46 | 65 | 40 | |
| MapGPTVLN specialist category=Zero-shot VLN with Foundation Models, Foundation Model=Q3VL, Re-implementation status=true2026.03 | 5.27 | 32 | 47.5 | 28.7 | |
| RegretfulDecoding=greedy, Data Augmentation=true2019.03 | 5.32 | 50 | 59 | 0.41 | |
| DiscussNavVLN specialist category=Zero-shot VLN with Foundation Models, Foundation Model=GPT-42026.03 | 5.32 | 43 | 61 | 40 | |
| RegretfulDecoding=greedy, Data Augmentation=false2019.03 | 5.36 | 48 | 61 | 0.37 | |
| Self-MonitoringDecoding=greedy, Data Augmentation=true2019.03 | 5.52 | 45 | 56 | 0.32 | |
| MapGPTVLN specialist category=Zero-shot VLN with Foundation Models, Foundation Model=GPT-4V2026.03 | 5.63 | 43.7 | 57.6 | 34.8 | |
| RCMDecoding=greedy, Data Augmentation=true2019.03 | 5.88 | 43 | 52 | — | |
| NavCoTVLN specialist category=Pretraining-based VLN specialists2026.03 | 6.26 | 40 | 48 | 37 | |
| MapGPTVLN specialist category=Zero-shot VLN with Foundation Models, Foundation Model=GPT-42026.03 | 6.29 | 38.8 | 57.6 | 25.8 | |
| NavGPTVLN specialist category=Zero-shot VLN with Foundation Models, Foundation Model=GPT-42026.03 | 6.46 | 34 | 42 | 29 | |
| NavGPTVLN specialist category=Zero-shot VLN with Foundation Models, Foundation Model=DS3, Re-implementation status=true2026.03 | 6.52 | 36 | 50 | 28.1 | |
| Speaker-FollowerDecoding=greedy, Data Augmentation=true2019.03 | 6.62 | 36 | 45 | — | |
| Speaker-FollowerVLN specialist category=Training-based VLN specialists2026.03 | 6.62 | 36 | 45 | — | |
| RPADecoding=greedy2019.03 | 7.65 | 25 | 32 | — | |
| Student-forcingDecoding=greedy2019.03 | 7.81 | 22 | 28 | — | |
| Seq2SeqVLN specialist category=Training-based VLN specialists2026.03 | 7.81 | 21 | 28 | — | |
| RandomDecoding=greedy2019.03 | 9.23 | 16 | 22 | — |