Ego-motion Estimation on OctoSense night (test)
0.05Translation RMSE (m)Late-fusion MAE (w/ full attention)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Late-fusion MAE (w/ full attention)Attention=full2026.06 | 0.05 | 0.17 | 2.25 | |
| Late-fusion MAE (w/o EV)Removed Modality=EV2026.06 | 0.05 | 0.2 | 2.42 | |
| Late-fusion MAEFusion=late2026.06 | 0.06 | 0.23 | 2.27 | |
| Late-fusion MAE (w/o RGB)Removed Modality=RGB2026.06 | 0.06 | 0.23 | 2.34 | |
| Late-fusion MAE (w/o IMU)Removed Modality=IMU2026.06 | 0.06 | 0.22 | 2.61 | |
| Early-fusion MAEFusion=early2026.06 | 0.15 | 0.53 | 4.48 | |
| Late-fusion MAE (w/o LiDAR)Removed Modality=LiDAR2026.06 | 0.57 | 0.45 | 4 | |
| V-JEPA 2.1 (stereo)*Encoder=V-JEPA 2.1, Input=stereo2026.06 | 0.71 | 0.47 | 6.47 | |
| RGB-only video MAEModality=RGB, Type=video2026.06 | 0.72 | 0.5 | 12.45 | |
| V-JEPA 2.1Encoder=V-JEPA 2.12026.06 | 0.76 | 0.47 | 6.83 | |
| Late-fusion MAE (w/ only RGB)Modality=only RGB2026.06 | 0.77 | 0.47 | 7.75 | |
| SigLIP 2Encoder=SigLIP 22026.06 | 0.78 | 0.94 | 36 | |
| DINOv3Encoder=DINOv32026.06 | 0.79 | 0.76 | 21.93 | |
| DINOv2Encoder=DINOv22026.06 | 0.84 | 0.75 | 23.11 | |
| Perception Enc.Encoder=Perception Enc.2026.06 | 0.91 | 0.93 | 34.54 | |
| RGB-only image MAEModality=RGB, Type=image2026.06 | 1.05 | 0.77 | 21.39 |