Ego-motion Estimation on OctoSense day (test)
0.05Translation RMSE (m)Late-fusion MAE (w/ full attention)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Late-fusion MAE (w/ full attention)Attention=full2026.06 | 0.05 | 0.17 | 2.46 | |
| Late-fusion MAE (w/o EV)Removed Modality=EV2026.06 | 0.05 | 0.21 | 2.59 | |
| Late-fusion MAEFusion=late2026.06 | 0.06 | 0.24 | 2.51 | |
| Late-fusion MAE (w/o RGB)Removed Modality=RGB2026.06 | 0.06 | 0.24 | 2.65 | |
| Late-fusion MAE (w/o IMU)Removed Modality=IMU2026.06 | 0.06 | 0.24 | 2.78 | |
| Early-fusion MAEFusion=early2026.06 | 0.14 | 0.56 | 4.58 | |
| Late-fusion MAE (w/o LiDAR)Removed Modality=LiDAR2026.06 | 0.73 | 0.46 | 3.97 | |
| RGB-only video MAEModality=RGB, Type=video2026.06 | 0.76 | 0.5 | 12.96 | |
| V-JEPA 2.1Encoder=V-JEPA 2.12026.06 | 0.77 | 0.47 | 5.98 | |
| V-JEPA 2.1 (stereo)*Encoder=V-JEPA 2.1, Input=stereo2026.06 | 0.77 | 0.48 | 5.69 | |
| Late-fusion MAE (w/ only RGB)Modality=only RGB2026.06 | 0.83 | 0.47 | 8.29 | |
| RGB-only image MAEModality=RGB, Type=image2026.06 | 0.92 | 0.88 | 31.57 | |
| DINOv3Encoder=DINOv32026.06 | 0.93 | 0.79 | 25.24 | |
| DINOv2Encoder=DINOv22026.06 | 0.96 | 0.8 | 22.62 | |
| Perception Enc.Encoder=Perception Enc.2026.06 | 0.96 | 0.94 | 35.62 | |
| SigLIP 2Encoder=SigLIP 22026.06 | 0.98 | 0.95 | 36.12 |