Scene Text Recognition on SVT 647 images
99.6AccuracyCLIP4STR-L
Evaluation Results
| Method | Links | |
|---|---|---|
| CLIP4STR-LData=RBU, Params=446 M2025.03 | 99.6 | |
| CLIP4STR-BData=RBU, Params=158 M2025.03 | 99.5 | |
| CLIP4STR-HData=RBU, Params=1 B2025.03 | 99.5 | |
| CSD-DData=RBU, Params=110.2 M2025.03 | 99.2 | |
| CSD-DData=Real, Params=110.2 M2025.03 | 99.1 | |
| CSD-SData=RBU, Params=40.8 M2025.03 | 98.8 | |
| CSD-BData=RBU, Params=104.9 M2025.03 | 98.8 | |
| CLIP4STR-LData=Real, Params=446 M2025.03 | 98.5 | |
| CSD-SData=Real, Params=40.8 M2025.03 | 98.5 | |
| CLIP4STR-BData=Real, Params=158 M2025.03 | 98.3 | |
| CSD-BData=Real, Params=104.9 M2025.03 | 98 | |
| PARSeqData=Real, Params=22.5 M2025.03 | 97.9 | |
| ABINETData=Real, Params=23.5 M2025.03 | 97.8 | |
| TRBAData=Real, Params=49.6 M2025.03 | 97 | |
| VITSTR-SData=Real, Params=21.7 M2025.03 | 95.8 | |
| SIGATType=V, Backbone=ViT-B, Structure=Transformer, Size=32×1282022.03 | 95.1 | |
| ABINet+ConCLRType=VL, Backbone=ResNet45-Trns, Structure=Transformer, Size=32×1282022.03 | 94.3 | |
| S-GTRType=VL, Backbone=ResNet50Dilated-PPM, Structure=ResNet, Size=64×2562022.03 | 94.1 | |
| PARSeqType=VL, Backbone=DeiT, Structure=Transformer, Size=32×1282022.03 | 93.6 | |
| PARSeqTraining Datasets=MJ and ST2023.06 | 93.6 | |
| DiffusionSTRTraining Datasets=MJ and ST2023.06 | 93.6 | |
| ABINetType=VL, Backbone=ResNet45-Trns, Structure=Transformer, Size=32×1282022.03 | 93.5 | |
| ABINetTraining Datasets=MJ and ST2023.06 | 93.5 | |
| LevOCRType=VL, Backbone=ResNet45, Structure=ResNet, Size=32×1282022.03 | 92.9 | |
| SIGARType=V, Backbone=ResNet45, Structure=ResNet, Size=32×1282022.03 | 92.7 | |
| Bhunia et al.Type=VL, Backbone=ResNet50-FPN, Structure=ResNet, Size=32×1002022.03 | 92.2 | |
| LevOCRType=VL, Backbone=ViT, Structure=Transformer, Size=32×1282022.03 | 91.8 | |
| VisionLANType=VL, Backbone=ResNet45, Structure=ResNet, Size=64×2562022.03 | 91.7 | |
| SRNType=VL, Backbone=ResNet50-FPN, Structure=ResNet, Size=64×2562022.03 | 91.5 | |
| CRNNData=Real, Params=8.5 M2025.03 | 90.7 | |
| TRBATraining Datasets=MJ and ST2023.06 | 88.9 | |
| ViTSTRTraining Datasets=MJ and ST2023.06 | 87.7 | |
| CRNNTraining Datasets=MJ and ST2023.06 | 80.1 |