Vision-Language Navigation on R2R-CE (val-unseen)
81.1Success Rate (SR)StereoNav
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| StereoNavSize=3B, Observation: Pano.=false, Observation: Depth=false, Observation: RGB=true, Prediction: Extra=true, Prediction: Action=true, Agent Type=Egocentric RGB Agent, Training Data=Additional external data2026.05 | 81.1 | 2.1 | 82.4 | 68.3 | — | — | — | — | |
| StereoNavSize=3B, Observation: Pano.=false, Observation: Depth=false, Observation: RGB=true, Prediction: Extra=true, Prediction: Action=true, Agent Type=Egocentric RGB Agent, Training Data=Standard navigation data only2026.05 | 72.8 | 3 | 76.6 | 56.4 | — | — | — | — | |
| Image2NavRGB=✓, FOV=180°2026.07 | 70.3 | 3.71 | 76.1 | 65.6 | — | — | — | — | |
| SpatialNavSettings=Zero-Shot, Ground truth spatial annotations=true2026.01 | 68 | 4.21 | 73 | 53.4 | — | 69.3 | — | — | |
| SpatialNavLearning paradigm=Pre-exploration2026.06 | 68 | 4.2 | 73 | 53.4 | — | — | — | — | |
| AstraNav-World w/ Diffusion PolicyObservation: Pano.=true, Action Decoder=Diffusion Policy2025.12 | 67.9 | 3.86 | 73.9 | 65.4 | — | — | — | — | |
| AstraNav-World w/ Action FormerObservation: Pano.=true, Action Decoder=Action Former2025.12 | 67.2 | 3.93 | 73.1 | 64.2 | — | — | — | — | |
| AgentVLN-3BS.RGB=true, Depth=true2026.03 | 67.2 | 3.88 | 73.5 | 64.7 | — | — | — | — | |
| AgentVLNSize=3B, Observation: Pano.=false, Observation: Depth=true, Observation: RGB=true, Prediction: Extra=false, Prediction: Action=true, Agent Type=Egocentric RGB-D Agent2026.05 | 67.2 | 3.9 | 73.5 | 64.7 | — | — | — | — | |
| ABot-N0Pano.=true2026.02 | 66.4 | 3.78 | 70.8 | 63.9 | — | — | — | — | |
| ABot-N0Size=4B, Observation: Pano.=true, Observation: Depth=false, Observation: RGB=false, Prediction: Extra=false, Prediction: Action=true, Agent Type=Panoramic RGB Agent2026.05 | 66.4 | 3.8 | 70.8 | 63.9 | — | — | — | — | |
| SPAN-NavRGB=true2026.03 | 66.3 | 4.07 | 75.3 | 59.3 | — | — | — | — | |
| SPAN-NavSize=-, Observation: Pano.=true, Observation: Depth=false, Observation: RGB=false, Prediction: Extra=true, Prediction: Action=true, Agent Type=Panoramic RGB Agent2026.05 | 66.3 | 4.1 | 75.3 | 59.3 | — | — | — | — | |
| Image2NavRGB=✓, FOV=90°2026.07 | 66.3 | 3.96 | 72.9 | 61.5 | — | — | — | — | |
| NavForeseePano=true2025.12 | 66.2 | 3.94 | 78.4 | 59.7 | — | — | — | — | |
| NavForeseeSize=3B, Observation: Pano.=true, Observation: Depth=false, Observation: RGB=false, Prediction: Extra=true, Prediction: Action=true, Agent Type=Panoramic RGB Agent2026.05 | 66.2 | 3.9 | 78.4 | 59.7 | — | — | — | — | |
| Dual-Anchoring Framework (Ours)S-RGB=true2026.04 | 65.6 | 4.15 | 69.2 | 62.1 | — | — | — | — | |
| AwareVLNS.RGB=true, Pano.=false, Depth=false, Odo.=false, Waypoint Predictor=false2026.05 | 65.4 | 4.02 | 73.5 | 55.1 | — | — | — | — | |
| FutureNav-4BVLM Params=4B, Single RGB=true, External Training Data=10457K2026.06 | 65.4 | 4.24 | 70.2 | 61.3 | — | — | — | — | |
| CorrectNavObservation: S.RGB=true2025.12 | 65.1 | 4.24 | 67.5 | 62.3 | — | — | — | — | |
| CorrectNavS.RGB=true2025.12 | 65.1 | 4.24 | 67.5 | 62.3 | — | — | — | — | |
| ETP-R1 (Ours-GRPO)training=GRPO-based online RFT2025.12 | 65 | 3.94 | 72 | 56 | — | — | — | — | |
| LCG-ETP-R1Method category=Map based2026.05 | 65 | 3.93 | 71 | 57 | 11.84 | — | — | — | |
| CLASHSize=-, Observation: Pano.=true, Observation: Depth=true, Observation: RGB=false, Prediction: Extra=false, Prediction: Action=true, Agent Type=Panoramic RGB-D Agent2026.05 | 65 | 4.1 | 73 | 55 | — | — | — | — | |
| ETP-R1Size=-, Observation: Pano.=true, Observation: Depth=true, Observation: RGB=false, Prediction: Extra=false, Prediction: Action=true, Agent Type=Panoramic RGB-D Agent2026.05 | 65 | 3.9 | 72 | 56 | — | — | — | — | |
| SACAPano.=false, Odo.=false, Depth=false, S.RGB=true, Waypoint Predictor=false, Additional training data=true2026.03 | 64.7 | 4.19 | 69.3 | 56.9 | — | — | — | — | |
| SACASize=8B, Observation: Pano.=false, Observation: Depth=false, Observation: RGB=true, Prediction: Extra=false, Prediction: Action=true, Agent Type=Egocentric RGB Agent2026.05 | 64.7 | 4.2 | 69.3 | 56.9 | — | — | — | — | |
| DualVLNPano.=false, Odo.=false, Depth=false, S.RGB=true, Waypoint predictor=false2025.12 | 64.3 | 4.05 | 70.7 | 58.5 | — | — | — | — | |
| DualVLNBackbone=QwenVL-2.5 7B2026.03 | 64.3 | 4.05 | 70.7 | 58.5 | — | — | — | — | |
| DualVLN-7.1BS.RGB=true2026.03 | 64.3 | 4.05 | 70.7 | 58.5 | — | — | — | — | |
| DualVLNS-RGB=true2026.04 | 64.3 | 4.05 | 70.7 | 58.5 | — | — | — | — | |
| DualVLNSize=8B, Observation: Pano.=false, Observation: Depth=false, Observation: RGB=true, Prediction: Extra=false, Prediction: Action=true, Agent Type=Egocentric RGB Agent2026.05 | 64.3 | 4.1 | 70.7 | 58.5 | — | — | — | — | |
| FutureNav-8BVLM Params=8B, Single RGB=true, External Training Data=10435K2026.06 | 64.3 | 4.29 | 69.2 | 60.3 | — | — | — | — | |
| DualVLNRGB=✓, Depth=-2026.07 | 64.3 | 4.05 | 70.7 | 58.5 | — | — | — | — | |
| Efficient-VLNSettings=Supervised Learning, Ground truth spatial annotations=false2026.01 | 64.2 | 4.18 | 73.7 | 55.9 | — | — | — | — | |
| Efficient-VLNObservation Encoder=Single RGB, Additional Training Data=true, Hardware=H8002025.12 | 64.2 | 4.18 | 73.7 | 55.9 | — | — | — | — | |
| Efficient-VLNLearning Paradigm=Supervised Learning2026.02 | 64.2 | 4.18 | 73.7 | 55.9 | — | — | — | — | |
| EfficientVLN-4BS.RGB=true2026.03 | 64.2 | 4.18 | 73.7 | 55.9 | — | — | — | — | |
| Efficient-VLNSize=4B, Observation: Pano.=false, Observation: Depth=false, Observation: RGB=true, Prediction: Extra=false, Prediction: Action=true, Agent Type=Egocentric RGB Agent2026.05 | 64.2 | 4.2 | 73.7 | 55.9 | — | — | — | — | |
| EfficientVLNRGB=✓, Depth=-2026.07 | 64.2 | 4.18 | 73.7 | 55.9 | — | — | — | — | |
| SpatialNavSettings=Zero-Shot, Ground truth spatial annotations=false2026.01 | 64 | 5.15 | 66 | 51.1 | — | 65.4 | — | — | |
| HSAN2026.06 | 64 | 3.28 | 71 | 59 | — | — | — | — | |
| SPAN-Nav GeneralizeRGB=true2026.03 | 63.3 | 4.02 | 73.1 | 57.3 | — | — | — | — | |
| VLN-CacheBackbone=QwenVL-2.5 7B2026.03 | 63.1 | 3.93 | 71.4 | 57.6 | — | — | — | — | |
| ETP-R1 (Ours-DAgger)training=online SFT2025.12 | 63 | 4.11 | 69 | 54 | — | — | — | — | |
| ETPNav w/ Φ-NavTraining paradigm=Scheduled Sampling, Observation view=panoramic-view, Φ-Nav application=true2026.07 | 62.58 | 4.29 | 68.84 | 52.74 | 12.76 | — | — | — | |
| SpaAct-stage2Pano.=false, Odo.=false, Depth=false, S.RGB=true, External Data=10476K2026.04 | 62.2 | 4.7 | 69.9 | 57.2 | — | — | — | — | |
| P3Nav2026.03 | 62 | 4.39 | 69 | 52 | — | — | — | — | |
| P³NavSize=-, Observation: Pano.=true, Observation: Depth=true, Observation: RGB=false, Prediction: Extra=false, Prediction: Action=true, Agent Type=Panoramic RGB-D Agent2026.05 | 62 | 4.4 | 69 | 52 | — | — | — | — | |
| IDEALearning Paradigm=Self-supervised, Base Model=BEVBert, Method Adjustment=IDEA2026.05 | 62 | 4.26 | 69 | 52 | 12.67 | — | — | — | |
| IDEAPolicy=BEVBert2026.05 | 62 | 4.26 | 69 | 52 | 12.67 | — | — | — | |
| NavFoMSettings=Supervised Learning, Ground truth spatial annotations=false2026.01 | 61.7 | 4.61 | 72.1 | 55.3 | — | — | — | — | |
| NavFoM (Four views)Pano.=true2026.02 | 61.7 | 4.61 | 72.1 | 55.3 | — | — | — | — | |
| NavFoMLearning Paradigm=Supervised Learning2026.02 | 61.7 | 4.61 | 72.1 | 55.3 | — | — | — | — | |
| NavFomRGB=true2026.03 | 61.7 | 4.61 | 72.1 | 55.3 | — | — | — | — | |
| NavFoMSize=7B, Observation: Pano.=true, Observation: Depth=false, Observation: RGB=false, Prediction: Extra=false, Prediction: Action=true, Agent Type=Panoramic RGB Agent2026.05 | 61.7 | 4.6 | 72.1 | 55.3 | — | — | — | — | |
| NAVIDAVLM # Params=3B, Observation: S.RGB=true2026.01 | 61.4 | 4.32 | 69.5 | 54.7 | — | — | — | — | |
| NaVIDASize=3B, Observation: Pano.=false, Observation: Depth=false, Observation: RGB=true, Prediction: Extra=false, Prediction: Action=true, Agent Type=Egocentric RGB Agent2026.05 | 61.4 | 4.3 | 69.5 | 54.7 | — | — | — | — | |
| D3D-VLPSize=2B, Observation: Pano.=true, Observation: Depth=true, Observation: RGB=false, Prediction: Extra=false, Prediction: Action=true, Agent Type=Panoramic RGB-D Agent2026.05 | 61.3 | 4.7 | 67.2 | 56.1 | — | — | — | — | |
| D3D-VLPRGB=✓, Depth=✓2026.07 | 61.3 | 4.73 | 67.2 | 56.1 | — | — | — | — | |
| HNRWaypoint Predictor=true, Cur. RGB=true, Panoramic View=true, Depth=true, Odometry=true2025.02 | 61 | 4.42 | 67 | 51 | — | — | — | — | |
| HNR2024.10 | 61 | 4.42 | 67 | 51 | 12.64 | — | — | — | |
| HNRSettings=Supervised Learning, Ground truth spatial annotations=false2026.01 | 61 | 4.42 | 67 | 51 | — | — | — | — | |
| HNR*Observation: Pano.=true, Observation: Depth=true, Observation: Odo.=true, Waypoint Predictor=true2025.12 | 61 | 4.42 | 67 | 51 | — | — | — | — | |
| HNRPano.=true, Depth=true, Odo.=true, waypoint predictor=true2026.02 | 61 | 4.42 | 67 | 51 | — | — | — | — | |
| HNR2025.12 | 61 | 4.42 | 67 | 51 | — | — | — | — | |
| G3D-LF2025.12 | 61 | 4.53 | 68 | 52 | — | — | — | — | |
| HNRPano=true, Depth=true, Odo=true2025.12 | 61 | 4.42 | 67 | 51 | — | — | — | — | |
| HNRLearning Paradigm=Supervised Learning2026.02 | 61 | 4.42 | 67 | 51 | — | — | — | — | |
| HNR*RGB=true, Depth=true, Odo.=true2026.03 | 61 | 4.42 | 67 | 51 | — | — | — | — | |
| HNR-VLN2026.03 | 61 | 4.42 | 67 | 51 | — | — | — | — | |
| G3D-LF2026.03 | 61 | 4.53 | 68 | 52 | — | — | — | — | |
| g3D-LFObservation Data=Complex Data, Waypoint Predictor=true2026.03 | 61 | 4.53 | 68 | 52 | — | — | — | — | |
| HNRPano.=true, Odo.=true, Depth=true, Waypoint Predictor=true2026.04 | 61 | 4.42 | 67 | 51 | — | — | — | — | |
| JanusVLNMethod category=End to End2026.05 | 61 | 4.78 | 65 | 57 | — | — | — | — | |
| GA-VLNSystem=GA-VLN, DAgger=false2026.05 | 61 | 4.8 | 67.6 | 55.2 | — | — | — | — | |
| HNRS.RGB=true, Pano.=true, Depth=true, Odo.=false, Waypoint Predictor=true2026.05 | 61 | 4.42 | 67 | 51 | — | — | — | — | |
| BEVBert + FEEDTTALearning Paradigm=With Feedback, Base Model=BEVBert, Method Adjustment=FEEDTTA2026.05 | 61 | 4.33 | 69 | 50 | 16.15 | — | — | — | |
| HNRTraining paradigm=Scheduled Sampling, Observation view=panoramic-view, Φ-Nav application=false2026.07 | 61 | 4.42 | 67 | 51 | 12.64 | — | — | — | |
| g3D-LFTraining paradigm=Scheduled Sampling, Observation view=panoramic-view, Φ-Nav application=false2026.07 | 61 | 4.53 | 68 | 52 | — | — | — | — | |
| Efficient-VLNObservation Encoder=Single RGB, Hardware=H800, Training cost=282 hours2025.12 | 60.8 | 4.36 | 69 | 53.7 | — | — | — | — | |
| DyGeoVLNObservation Data=Simple Data, Waypoint Predictor=false2026.03 | 60.8 | 4.41 | 70.1 | 55.8 | — | — | — | — | |
| DyGeoVLNSize=9B, Observation: Pano.=false, Observation: Depth=false, Observation: RGB=true, Prediction: Extra=true, Prediction: Action=true, Agent Type=Egocentric RGB Agent2026.05 | 60.8 | 4.4 | 70.1 | 55.8 | — | — | — | — | |
| JanusVLNVLM # Params=8B, Observation: S.RGB=true2026.01 | 60.5 | 4.78 | 65.2 | 56.8 | — | — | — | — | |
| JanusVLN-8.2BS.RGB=true2026.03 | 60.5 | 4.78 | 65.2 | 56.8 | — | — | — | — | |
| JanusVLNPano.=false, Odo.=false, Depth=false, S.RGB=true, External Data=10692K2026.04 | 60.5 | 4.78 | 65.2 | 56.8 | — | — | — | — | |
| JanusVLNSize=8B, Observation: Pano.=false, Observation: Depth=false, Observation: RGB=true, Prediction: Extra=false, Prediction: Action=true, Agent Type=Egocentric RGB Agent2026.05 | 60.5 | 4.8 | 65.2 | 56.8 | — | — | — | — | |
| JanusVLNVLM Params=7B, Single RGB=true, External Training Data=10692K2026.06 | 60.5 | 4.78 | 65.2 | 56.8 | — | — | — | — | |
| JanusVLNRGB=✓, Depth=-2026.07 | 60.5 | 4.78 | 65.2 | 56.8 | — | — | — | — | |
| SACAPano.=false, Odo.=false, Depth=false, S.RGB=true, Waypoint Predictor=false, Additional training data=false2026.03 | 60.3 | 4.57 | 64.9 | 55.1 | — | — | — | — | |
| Progress-ThinkParams=2B+2B, Time / Episode=36.60s2025.11 | 60.1 | — | — | — | — | — | — | — | |
| Progress-ThinkS.RGB=true, Depth=false, Pano.=false, External Data=0K, uses LLM=true2025.11 | 60.1 | 4.68 | 63.6 | 53.6 | — | — | — | — | |
| Safe-VLNParadigm=Explicit Map-based2026.01 | 60 | 4.48 | 68 | 47 | 15 | — | — | — | |
| BEVBertEvaluation Protocol=Supervised Learning2026.04 | 60 | 5.13 | 64 | 53.41 | 13.63 | 61.4 | — | — | |
| OVL-MAPMethod category=Map based2026.05 | 60 | 4.69 | 65 | 50 | 11.45 | — | — | — | |
| BEVBert + RLCFLearning Paradigm=With Feedback, Base Model=BEVBert, Method Adjustment=RLCF2026.05 | 60 | 4.53 | 66 | 50 | 13.15 | — | — | — | |
| BEVBert + ATENALearning Paradigm=With Feedback, Base Model=BEVBert, Method Adjustment=ATENA2026.05 | 60 | 4.5 | 67 | 51 | 13.48 | — | — | — | |
| BEVBert + FSTTALearning Paradigm=Self-supervised, Base Model=BEVBert, Method Adjustment=FSTTA2026.05 | 60 | 4.39 | 65 | 51 | 13.11 | — | — | — | |
| BEVBert + ReCAPLearning Paradigm=Self-supervised, Base Model=BEVBert, Method Adjustment=ReCAP2026.05 | 60 | 4.57 | 66 | 50 | 13.01 | — | — | — | |
| FSTTAPolicy=BEVBert2026.05 | 60 | 4.39 | 65 | 51 | 13.11 | — | — | — |