Scene Text Recognition on SVT
95.5AccuracyABINet-LVest
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| ABINet-LVestLabeled Datasets=MJ+ST, Unlabeled Datasets=Uber-Text2021.03 | 95.5 | — | — | — | |
| ABINet-LVstLabeled Datasets=MJ+ST, Unlabeled Datasets=Uber-Text2021.03 | 94.9 | — | — | — | |
| Our model2021.07 | 94.7 | — | — | — | |
| PREN2DLanguage Mode=Lan-aware2022.04 | 94 | — | — | — | |
| VSTLanguage Mode=Lan-aware2022.04 | 93.8 | — | — | — | |
| CSTR2021.07 | 93.7 | — | — | — | |
| ASTER2021.07 | 93.7 | — | — | — | |
| ABINet-LVLabeled Datasets=MJ+ST2021.03 | 93.5 | — | — | — | |
| ABINetLanguage Mode=Lan-aware2022.04 | 93.5 | — | — | — | |
| ABINet-SVLabeled Datasets=MJ+ST2021.03 | 93.2 | — | — | — | |
| SVTR-SBackbone=Small2022.04 | 93 | — | — | — | |
| SCATTERVenue=CVPR’20, Training Set=MJ+ST+Extra, RNN=Y, FPS=-2022.04 | 92.7 | — | — | — | |
| SRN-LV (Reproduced)Labeled Datasets=MJ+ST2021.03 | 92.3 | — | — | — | |
| PREN*Language Mode=Lan-free2022.04 | 92 | — | — | — | |
| VST*Language Mode=Lan-free2022.04 | 91.9 | — | — | — | |
| RCEED2021.07 | 91.8 | — | — | — | |
| VisionLANLanguage Mode=Lan-aware2022.04 | 91.7 | — | — | — | |
| SVTR-LBackbone=Large2022.04 | 91.7 | — | — | — | |
| SVTR-TBackbone=Tiny2022.04 | 91.6 | — | — | — | |
| SRNLabeled Datasets=MJ+ST2021.03 | 91.5 | — | — | — | |
| SRNLanguage Mode=Lan-aware2022.04 | 91.5 | — | — | — | |
| SVTR-BBackbone=Base2022.04 | 91.5 | — | — | — | |
| SATRN2021.07 | 91.3 | — | — | — | |
| SRN-SV (Reproduced)Labeled Datasets=MJ+ST2021.03 | 90.9 | — | — | — | |
| AutoSTRLanguage Mode=Lan-aware2022.04 | 90.9 | — | — | — | |
| SAMLabeled Datasets=MJ+ST2021.03 | 90.6 | — | — | — | |
| ABINet*Language Mode=Lan-free2022.04 | 90.4 | — | — | — | |
| Zhan et al. [57]lexicon=lexicon-free, training_annotations=word-level2020.05 | 90.2 | — | — | — | |
| ESIRVenue=CVPR’19, Training Set=MJ+ST, RNN=Y, FPS=-2022.04 | 90.2 | — | — | — | |
| ParallelLabeled Datasets=MJ+ST2021.03 | 90.1 | — | — | — | |
| TextscannerLabeled Datasets=MJ+ST2021.03 | 90.1 | — | — | — | |
| TextScanner*Venue=AAAI’20, Training Set=MJ+ST+Extra, RNN=N, FPS=-2022.04 | 90.1 | — | — | — | |
| SE-ASTER (Ours)lexicon=lexicon-free, training_annotations=word-level2020.05 | 89.6 | — | — | — | |
| SE-ASTERLabeled Datasets=MJ+ST2021.03 | 89.6 | — | — | — | |
| SEED2021.07 | 89.6 | — | — | — | |
| SEEDVenue=CVPR’20, Training Set=MJ+ST, RNN=Y, FPS=-2022.04 | 89.6 | — | — | — | |
| ASTER [45]lexicon=lexicon-free, training_annotations=word-level2020.05 | 89.5 | — | — | — | |
| ASTERLanguage Mode=Lan-aware2022.04 | 89.5 | — | — | — | |
| DANLabeled Datasets=MJ+ST2021.03 | 89.2 | — | — | — | |
| DANVenue=AAAI’20, Training Set=MJ+ST, RNN=Y, FPS=-2022.04 | 89.2 | — | — | — | |
| Yang et al. [53]*lexicon=lexicon-free, training_annotations=word-level and character-level2020.05 | 88.9 | — | — | — | |
| Luo et al. [32]lexicon=lexicon-free, training_annotations=word-level2020.05 | 88.3 | — | — | — | |
| MORANLanguage Mode=Lan-aware2022.04 | 88.3 | — | — | — | |
| NRTRLanguage Mode=Lan-aware2022.04 | 88.3 | — | — | — | |
| RobustScannerLabeled Datasets=MJ+ST2021.03 | 88.1 | — | — | — | |
| SRN*Language Mode=Lan-free2022.04 | 88.1 | — | — | — | |
| DOTA2026.03 | 88.1 | — | — | — | |
| DOTA + CRFRefinement=CRF2026.03 | 88.1 | — | — | — | |
| Baek et al.2021.07 | 87.9 | — | — | — | |
| RES50-DEF(L4)-ViT-AdaptiveDeformable Layer=L4, Config=Adaptive2026.03 | 87.79 | — | — | — | |
| ViTSTRLanguage Mode=Lan-free2022.04 | 87.7 | — | — | — | |
| Bai et al. [2]*lexicon=lexicon-free, training_annotations=word-level and character-level2020.05 | 87.5 | — | — | — | |
| Comb.BestVenue=ICCV’19, Training Set=MJ+ST, RNN=Y, FPS=36.232022.04 | 87.5 | — | — | — | |
| RES50-ViT-PEConfig=Positional Encoding2026.03 | 87.33 | — | — | — | |
| ASTER baseline reproducedlexicon=lexicon-free, training_annotations=word-level2020.05 | 87.2 | — | — | — | |
| Liu et al. [29]lexicon=lexicon-free, training_annotations=word-level2020.05 | 87.1 | — | — | — | |
| Liao et al. [25]*lexicon=lexicon-free, training_annotations=word-level and character-level2020.05 | 86.4 | — | — | — | |
| Ours-LargeVenue=-, Training Set=MJ+ST, RNN=N, FPS=66.91/255+2022.04 | 85.93 | — | — | — | |
| Cheng et al. [7]*lexicon=lexicon-free, training_annotations=word-level and character-level2020.05 | 85.9 | — | — | — | |
| RES50-ATT-ViT2026.03 | 85.78 | — | — | — | |
| Liu et al. [28]*lexicon=lexicon-free, training_annotations=word-level and character-level2020.05 | 85.5 | — | — | — | |
| RES50-ViTIteration=22026.03 | 85.16 | — | — | — | |
| RES50-ViT2026.03 | 84.85 | — | — | — | |
| RosettaLanguage Mode=Lan-free2022.04 | 84.7 | — | — | — | |
| RosettaVenue=KDD’18, Training Set=MJ+ST, RNN=N, FPS=212.762022.04 | 84.7 | — | — | — | |
| Li et al. [23]lexicon=lexicon-free, training_annotations=word-level2020.05 | 84.5 | — | — | — | |
| SARLanguage Mode=Lan-aware2022.04 | 84.5 | — | — | — | |
| SARVenue=AAAI’19, Training Set=MJ+ST, RNN=Y, FPS=-2022.04 | 84.5 | — | — | — | |
| Cheng et al. [8]lexicon=lexicon-free, training_annotations=word-level2020.05 | 82.8 | — | — | — | |
| Shi et al. [43]lexicon=lexicon-free, training_annotations=word-level2020.05 | 82.7 | — | — | — | |
| RES50-DEF-(L3–L4)Deformable Layers=L3-L42026.03 | 82.23 | — | — | — | |
| CA-FCN*Venue=AAAI’19, Training Set=ST, RNN=N, FPS=452022.04 | 82.1 | — | — | — | |
| Shi et al. [44]lexicon=lexicon-free, training_annotations=word-level2020.05 | 81.9 | — | — | — | |
| CRNNLanguage Mode=Lan-free2022.04 | 81.6 | — | — | — | |
| RES50-ATT-AdaptiveConfig=Adaptive2026.03 | 80.99 | — | — | — | |
| Lee et al. [22]lexicon=lexicon-free, training_annotations=word-level2020.05 | 80.7 | — | — | — | |
| RES50-ATTIteration=22026.03 | 79.13 | — | — | — | |
| ResNext2026.03 | 78.52 | — | — | — | |
| RES50-ATT2026.03 | 71.87 | — | — | — | |
| ABBYY2015.07 | — | 35 | — | — | |
| Almazán et al.2015.07 | — | 89.2 | — | — | |
| Almazán et al.Training Data=-2019.12 | — | 89.2 | — | — | |
| Alsharif and Pineau2015.07 | — | 74.3 | — | — | |
| AONTraining Data=ST + 90k2019.12 | — | 96 | 82.8 | — | |
| AONVenue=CVPR’18, Lexicon-based evaluation=Close-set2022.04 | — | 96 | — | — | |
| ASTERTraining Data=ST + 90k2019.12 | — | 97.4 | 89.5 | — | |
| Bissacco et al.2015.07 | — | 90.4 | 78 | — | |
| Bleeker and de RijkeRect.=No, Encoder/Decoder=Transformer, Train Dataset=MJ + ST2021.04 | — | — | — | 89 | |
| CA-FCNTraining Data=ST2019.12 | — | 98.8 | 86.4 | — | |
| CA-FCNVenue=AAAI’19, Lexicon-based evaluation=Close-set, Training Data=Datasets other than MJ and ST2022.04 | — | 98.5 | — | — | |
| CRNN2015.07 | — | 96.4 | 80.8 | — | |
| CRNNTraining Data=90k2019.12 | — | 96.4 | 80.8 | — | |
| EPTraining Data=ST + 90k2019.12 | — | 96.6 | 87.5 | — | |
| ESIRVenue=CVPR’19, Lexicon-based evaluation=Close-set2022.04 | — | 97.4 | — | — | |
| FANTraining Data=ST + 90k2019.12 | — | 97.1 | 85.9 | — | |
| Goel et al.2015.07 | — | 77.3 | — | — | |
| Gordo2015.07 | — | 91.8 | — | — | |
| GordoTraining Data=-2019.12 | — | 91.8 | — | — | |
| Jaderberg et al. (2014a)Training Data=-2019.12 | — | — | 86.1 | — | |
| Jaderberg et al. (2014a)Training Data=90k2019.12 | — | 93.2 | 71.7 | — |