Semantic Segmentation on ADE20k (mIoU, GFLOPs, FPS)
48.3mIoUSegViT
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| SegViTBackbone=ViT-Base, Input Resolution=512 x 5122025.09 | 48.3 | 113 | 53 | |
| SegViT + dCTS τ-6899Backbone=ViT-Base, Input Resolution=512 x 5122025.09 | 48.2 | 73 | 98 | |
| SegViT + CTSBackbone=ViT-Base, Input Resolution=512 x 512, Configuration=Default from original paper2025.09 | 47.8 | 75 | 40 | |
| SegViT + STEP@[8]τ-6899Backbone=ViT-Base, Input Resolution=512 x 5122025.09 | 47.1 | 68 | 32 | |
| SegViT + STEP@[6,8]τ-6899Backbone=ViT-Base, Input Resolution=512 x 5122025.09 | 46.9 | 64 | 24 | |
| SegViT + CTS & DToPBackbone=ViT-Base, Input Resolution=512 x 512, Configuration=Default from original paper2025.09 | 46.3 | 62 | 25 | |
| SegViT + DToPBackbone=ViT-Base, Input Resolution=512 x 512, Configuration=Default from original paper2025.09 | 45.8 | 91 | 25 | |
| SegViT + STEP@[8]τ-4999Backbone=ViT-Base, Input Resolution=512 x 5122025.09 | 45.8 | 53 | 43 | |
| SegViT + STEP@[6,8]τ-4999Backbone=ViT-Base, Input Resolution=512 x 5122025.09 | 45.3 | 50 | 34 |