Object Detection on MS COCO 2014 (test)
58.7AP (Box)FasterViT-4
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| FasterViT-4Model=DINO, Epochs=36, FLOPs (G)=1364, Throughput=84, Pre-training=ImageNet-21K2023.06 | 58.7 | — | — | |
| Swin-LModel=DINO, Epochs=36, FLOPs (G)=1285, Throughput=71, Pre-training=ImageNet-21K2023.06 | 58.5 | — | — | |
| Swin-LModel=HTC++, Epochs=72, FLOPs (G)=1470, Pre-training=ImageNet-21K2023.06 | 57.1 | — | — | |
| FasterViT-4Backbone=FasterViT-4, Throughput (im/sec)=117, Schedule=3x, Detection Framework=Cascade Mask R-CNN, Input Resolution=1280 x 8002023.06 | 52.9 | 71.6 | 57.7 | |
| ConvNeXt-BBackbone=ConvNeXt-B, Throughput (im/sec)=101, Schedule=3x, Detection Framework=Cascade Mask R-CNN, Input Resolution=1280 x 8002023.06 | 52.7 | — | 57.2 | |
| FasterViT-3Backbone=FasterViT-3, Throughput (im/sec)=159, Schedule=3x, Detection Framework=Cascade Mask R-CNN, Input Resolution=1280 x 8002023.06 | 52.4 | 71.1 | 56.7 | |
| FasterViT-2Backbone=FasterViT-2, Throughput (im/sec)=287, Schedule=3x, Detection Framework=Cascade Mask R-CNN, Input Resolution=1280 x 8002023.06 | 52.1 | 71 | 56.6 | |
| Swin-SBackbone=Swin-S, Throughput (im/sec)=119, Schedule=3x, Detection Framework=Cascade Mask R-CNN, Input Resolution=1280 x 8002023.06 | 51.9 | 70.7 | 56.3 | |
| ConvNeXt-SBackbone=ConvNeXt-S, Throughput (im/sec)=128, Schedule=3x, Detection Framework=Cascade Mask R-CNN, Input Resolution=1280 x 8002023.06 | 51.9 | 70.8 | 56.5 | |
| Swin-BBackbone=Swin-B, Throughput (im/sec)=90, Schedule=3x, Detection Framework=Cascade Mask R-CNN, Input Resolution=1280 x 8002023.06 | 51.9 | 70.5 | 56.4 | |
| Swin-TBackbone=Swin-T, Throughput (im/sec)=161, Schedule=3x, Detection Framework=Cascade Mask R-CNN, Input Resolution=1280 x 8002023.06 | 50.4 | 69.2 | 54.7 | |
| ConvNeXt-TBackbone=ConvNeXt-T, Throughput (im/sec)=166, Schedule=3x, Detection Framework=Cascade Mask R-CNN, Input Resolution=1280 x 8002023.06 | 50.4 | 69.1 | 54.8 | |
| X101-64Backbone=X101-64, Throughput (im/sec)=86, Schedule=3x, Detection Framework=Cascade Mask R-CNN, Input Resolution=1280 x 8002023.06 | 48.3 | 66.4 | 52.3 | |
| X101-32Backbone=X101-32, Throughput (im/sec)=124, Schedule=3x, Detection Framework=Cascade Mask R-CNN, Input Resolution=1280 x 8002023.06 | 48.1 | 66.5 | 52.4 | |
| DeiT-Small/16Backbone=DeiT-Small/16, Throughput (im/sec)=269, Schedule=3x, Detection Framework=Cascade Mask R-CNN, Input Resolution=1280 x 8002023.06 | 48 | 67.2 | 51.7 |