Scene Text Recognition on CUTE80
99.7AccuracyCSD-D
Evaluation Results
| Method | Links | |
|---|---|---|
| CSD-DData=RBU, Params=110.2 M2025.03 | 99.7 | |
| CLIP4STR-BData=Real, Params=158 M2025.03 | 99.3 | |
| CSD-DData=Real, Params=110.2 M2025.03 | 99.3 | |
| CSD-BData=RBU, Params=104.9 M2025.03 | 99.3 | |
| CLIP4STR-HData=RBU, Params=1 B2025.03 | 99.1 | |
| CLIP4STR-LData=Real, Params=446 M2025.03 | 99 | |
| CSD-SData=Real, Params=40.8 M2025.03 | 99 | |
| CSD-BData=Real, Params=104.9 M2025.03 | 99 | |
| CSD-SData=RBU, Params=40.8 M2025.03 | 99 | |
| CLIP4STR-LData=RBU, Params=446 M2025.03 | 98.6 | |
| PARSeqData=Real, Params=22.5 M2025.03 | 98.3 | |
| CLIP4STR-BData=RBU, Params=158 M2025.03 | 98.3 | |
| TRBAData=Real, Params=49.6 M2025.03 | 97.7 | |
| ABINETData=Real, Params=23.5 M2025.03 | 97.7 | |
| VITSTR-SData=Real, Params=21.7 M2025.03 | 96.1 | |
| SIGATType=V, Backbone=ViT-B, Structure=Transformer, Size=32×1282022.03 | 93.1 | |
| S-GTRType=VL, Backbone=ResNet50Dilated-PPM, Structure=ResNet, Size=64×2562022.03 | 92.3 | |
| PARSeqType=VL, Backbone=DeiT, Structure=Transformer, Size=32×1282022.03 | 92.2 | |
| LevOCRType=VL, Backbone=ResNet45, Structure=ResNet, Size=32×1282022.03 | 91.7 | |
| SIGARType=V, Backbone=ResNet45, Structure=ResNet, Size=32×1282022.03 | 91.7 | |
| ABINet+ConCLRType=VL, Backbone=ResNet45-Trns, Structure=Transformer, Size=32×1282022.03 | 91.3 | |
| Bhunia et al.Type=VL, Backbone=ResNet50-FPN, Structure=ResNet, Size=32×1002022.03 | 89.7 | |
| ABINetType=VL, Backbone=ResNet45-Trns, Structure=Transformer, Size=32×1282022.03 | 89.2 | |
| CRNNData=Real, Params=8.5 M2025.03 | 89.1 | |
| VisionLANType=VL, Backbone=ResNet45, Structure=ResNet, Size=64×2562022.03 | 88.5 | |
| SRNType=VL, Backbone=ResNet50-FPN, Structure=ResNet, Size=64×2562022.03 | 87.8 | |
| LevOCRType=VL, Backbone=ViT, Structure=Transformer, Size=32×1282022.03 | 86.8 | |
| DAN-2DRect=false, 2D=true, character-level annotation=false2019.12 | 84.4 | |
| Li et al. 2019Rect=false, 2D=true, character-level annotation=false2019.12 | 83.3 | |
| Zhan and Lu 2019Rect=true, 2D=false, character-level annotation=false2019.12 | 83.3 | |
| Xie et al. 2019Rect=false, 2D=true, character-level annotation=false2019.12 | 82.6 | |
| DAN-1DRect=false, 2D=false, character-level annotation=false2019.12 | 80.6 | |
| CA-FCNLexicon Size=0, Extra synthetic data=true2018.09 | 79.9 | |
| Liao et al. 2019Rect=false, 2D=true, character-level annotation=true2019.12 | 79.9 | |
| Shi et al. 2018Rect=true, 2D=false, character-level annotation=false2019.12 | 79.5 | |
| CA-FCNLexicon Size=0, Extra synthetic data=false2018.09 | 78.1 | |
| MORANLexicon=None2019.01 | 77.4 | |
| Luo, Jin, and Sun 2019Rect=true, 2D=true, character-level annotation=false2019.12 | 77.4 | |
| AONLexicon Size=02018.09 | 76.8 | |
| Cheng et al. [6]Lexicon=None2019.01 | 76.8 | |
| Cheng et al. 2018Rect=true, 2D=false, character-level annotation=false2019.12 | 76.8 | |
| Yang et al. 2017Lexicon Size=02018.09 | 69.3 | |
| Yang et al.Lexicon=None2019.01 | 69.3 | |
| DOTA2026.03 | 66.67 | |
| DOTA + CRFRefinement=CRF2026.03 | 66.67 | |
| RES50-ViT-PEConfig=Positional Encoding2026.03 | 65.97 | |
| RES50-DEF(L4)-ViT-AdaptiveDeformable Layer=L4, Config=Adaptive2026.03 | 64.24 | |
| Cheng et al. [5]Lexicon=None2019.01 | 63.9 | |
| Liu et al. 2018Rect=true, 2D=false, character-level annotation=false2019.12 | 62.5 | |
| RES50-ATT-ViT2026.03 | 62.5 | |
| RES50-ViTIteration=22026.03 | 61.81 | |
| RARELexicon Size=02018.09 | 59.2 | |
| Shi et al.Lexicon=None2019.01 | 59.2 | |
| RES50-ViT2026.03 | 56.94 | |
| RES50-ATT-AdaptiveConfig=Adaptive2026.03 | 55.9 | |
| ResNext2026.03 | 53.47 | |
| RES50-ATTIteration=22026.03 | 52.78 | |
| RES50-DEF-(L3–L4)Deformable Layers=L3-L42026.03 | 51.39 | |
| RES50-ATT2026.03 | 44.44 |