Semantic Segmentation on ADE20K (val)
62.9mIoUM3I Pre-training
Evaluation Results
| Method | Links | |||||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| M3I Pre-trainingModel=InternImage-H (1B), Pipeline=Single Stage: M3I Pre-training, Public Data=427M image-text, 15M image-category2022.11 | 62.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| BEIT-32022.08 | 62.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| BEIT-3Model=BEIT-3 (2B), Pipeline=Stage 1: CLIP, Stage 2: Dense Distillation, Stage 3: Masked Data Modeling, Public Data=21M image-text, 15M image-category, Private Data=400M image-text2022.11 | 62.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| BEiT3(w/ ViT-Adapter)Backbone=BEiT32022.05 | 62.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| BEIT-3Crop Size=896^22022.08 | 62 | — | 62.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| BEiT3Params=1.0 G, Pre-training Images=35 M, Pre-training Annotation=labeled & image-text, Segmenter=Mask2Former2022.12 | 62 | — | 62.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| BEIT-3 (w/ ViT-Adapter)Framework=Mask2Former, Backbone Pre-train=MM, BEIT-3, Extra Pre-train=COCO-Stuff, sup, Crop Size=896, Iters=80k, #Param=1.3B2022.05 | 62 | — | 62.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| ViT-Adapter-LBackbone=ViT-L2022.05 | 61.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| FD-SwinV2-GBackbone=SwinV2-G, Protocol=Fine-tuning2022.05 |