Semantic Segmentation on MADOS
76.407mIoUGFM-Swin
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GFM-SwinEvaluation Protocol=frozen-encoder, Data Modality=optical-only2026.05 | 76.407 | — | |
| TerraMind-LEvaluation Protocol=frozen-encoder, Data Modality=optical-only2026.05 | 75.492 | — | |
| BaselineModel Scale=600M, Params (M)=631.21, Train (M)=631.21, Time (min)=29.20±8.11, Mem (GB)=25.06±1.09, FLOPs (G)=646.85, Thr. (img/s)=16.28±0.31, Inf (s)=6.15±0.122026.03 | 69.6 | 94.1 | |
| CROMAEvaluation Protocol=frozen-encoder, Data Modality=optical-only2026.05 | 68.842 | — | |
| TerraMind-BParams=87.7M2026.03 | 67.44 | — | |
| SatlasNetEvaluation Protocol=frozen-encoder, Data Modality=optical-only2026.05 | 67.236 | — | |
| BaselineModel Scale=300M, Params (M)=303.90, Train (M)=303.90, Time (min)=15.90±4.80, Mem (GB)=11.70±0.47, FLOPs (G)=238.50, Thr. (img/s)=33.02±1.94, Inf (s)=3.04±0.182026.03 | 66.9 | 95.3 | |
| U-Net BaselineEvaluation Protocol=trained from scratch, Data Modality=optical-only2026.05 | 66.094 | — | |
| I-JEPAPre-training Dataset=IN-1k, CLS=Yes, PosEnc=default, Params=85.8M2026.03 | 65.59 | — | |
| Baseline with LoRAModel Scale=600M, Params (M)=635.20, Train (M)=3.99, Time (min)=29.55±4.85, Mem (GB)=17.99±1.05, FLOPs (G)=646.85, Thr. (img/s)=15.22±0.63, Inf (s)=6.58±0.282026.03 | 63.5 | 95.1 | |
| FLOROEvaluation Protocol=frozen-encoder, Data Modality=optical-only2026.05 | 63.113 | — | |
| SIMPLERModel Scale=300M, Params (M)=64.57, Train (M)=64.57, Time (min)=7.46±1.62, Mem (GB)=2.83±0.08, FLOPs (G)=50.67, Thr. (img/s)=88.72±15.04, Inf (s)=1.16±0.212026.03 | 62.8 | 94.2 | |
| SIMPLERModel Scale=600M, Params (M)=80.24, Train (M)=80.24, Time (min)=7.70±1.90, Mem (GB)=3.97±0.16, FLOPs (G)=82.22, Thr. (img/s)=77.26±11.27, Inf (s)=1.33±0.242026.03 | 62.2 | 90.5 | |
| RemoteCLIPEvaluation Protocol=frozen-encoder, Data Modality=optical-only2026.05 | 61.417 | — | |
| DOFAEvaluation Protocol=frozen-encoder, Data Modality=optical-only2026.05 | 61.366 | — | |
| SpectralGPTEvaluation Protocol=frozen-encoder, Data Modality=optical-only2026.05 | 61.127 | — | |
| Baseline + Prune 20%Model Scale=600M, Params (M)=513.14, Train (M)=513.14, Time (min)=48.75±10.08, Mem (GB)=25.06±1.09, FLOPs (G)=525.86, Thr. (img/s)=18.62±1.15, Inf (s)=5.39±0.342026.03 | 60.8 | 92.8 | |
| SSL4EO-S12-MAEEvaluation Protocol=frozen-encoder, Data Modality=optical-only2026.05 | 60.773 | — | |
| SIMPLER with LoRAModel Scale=300M, Params (M)=65.12, Train (M)=0.55, Time (min)=4.31±0.28, Mem (GB)=2.46±0.10, FLOPs (G)=50.67, Thr. (img/s)=79.51±14.44, Inf (s)=1.30±0.212026.03 | 60.4 | 91.8 | |
| SSL4EO-S12-DINOEvaluation Protocol=frozen-encoder, Data Modality=optical-only2026.05 | 60.341 | — | |
| Baseline with LoRAModel Scale=300M, Params (M)=306.32, Train (M)=2.42, Time (min)=11.77±2.17, Mem (GB)=8.76±0.50, FLOPs (G)=238.50, Thr. (img/s)=31.62±1.13, Inf (s)=3.17±0.112026.03 | 59.6 | 94.1 | |
| SSL4EO-S12-MoCoEvaluation Protocol=frozen-encoder, Data Modality=optical-only2026.05 | 59.575 | — | |
| I-JEPAPre-training Dataset=IN-1k, CLS=No, PosEnc=default, Params=85.8M2026.03 | 59.28 | — | |
| Baseline + Prune 20%Model Scale=300M, Params (M)=240.92, Train (M)=240.92, Time (min)=24.34±5.09, Mem (GB)=11.70±0.47, FLOPs (G)=189.07, Thr. (img/s)=35.33±4.02, Inf (s)=2.87±0.362026.03 | 58.4 | 93.2 | |
| SIMPLER with LoRAModel Scale=600M, Params (M)=80.79, Train (M)=0.55, Time (min)=6.49±1.18, Mem (GB)=3.27±0.12, FLOPs (G)=82.22, Thr. (img/s)=64.92±9.56, Inf (s)=1.57±0.192026.03 | 58.1 | 88.8 | |
| Baseline + Prune 40%Model Scale=600M, Params (M)=375.40, Train (M)=375.40, Time (min)=42.34±8.84, Mem (GB)=25.06±1.09, FLOPs (G)=384.70, Thr. (img/s)=25.08±0.90, Inf (s)=3.99±0.142026.03 | 55.6 | 90.2 | |
| SSL4EO-S12-Data2VecEvaluation Protocol=frozen-encoder, Data Modality=optical-only2026.05 | 55.062 | — | |
| CROMAParams=201.5M2026.03 | 54.97 | — | |
| Prithvi 1.0 100MEvaluation Protocol=frozen-encoder, Data Modality=optical-only2026.05 | 54.088 | — | |
| RemoteCLIPParams=87.5M2026.03 | 53.11 | — | |
| Scale-MAEEvaluation Protocol=frozen-encoder, Data Modality=optical-only2026.05 | 52.371 | — | |
| LEPAPre-training Dataset=HLS, CLS=No, PosEnc=CondPos, Params=86.4M2026.03 | 51.4 | — | |
| ViT BaselineEvaluation Protocol=trained from scratch, Data Modality=optical-only2026.05 | 50.974 | — | |
| DOFAParams=111.3M2026.03 | 50.32 | — | |
| Baseline + Prune 40%Model Scale=300M, Params (M)=177.94, Train (M)=177.94, Time (min)=22.51±4.83, Mem (GB)=11.70±0.47, FLOPs (G)=139.64, Thr. (img/s)=47.03±6.03, Inf (s)=2.16±0.252026.03 | 47.9 | 87.2 | |
| I-JEPAPre-training Dataset=HLS, CLS=Yes, PosEnc=CondPos, Params=86.4M2026.03 | 47.59 | — | |
| I-JEPAPre-training Dataset=HLS, CLS=Yes, PosEnc=default, Params=86.4M2026.03 | 46.4 | — | |
| I-JEPAPre-training Dataset=HLS, CLS=No, PosEnc=default, Params=86.4M2026.03 | 45.74 | — | |
| LEPAPre-training Dataset=HLS, CLS=Yes, PosEnc=CondPos, Params=86.4M2026.03 | 45.14 | — | |
| LEPAPre-training Dataset=HLS, CLS=No, PosEnc=default, Params=86.4M2026.03 | 43.36 | — | |
| Prithvi-EO-2.0-100MParams=86.4M2026.03 | 41.46 | — | |
| NoMaskPre-training Dataset=HLS, CLS=No, PosEnc=CondPos, Params=86.4M2026.03 | 33.51 | — |