Scene Text Recognition on SVTP
90.1AccuracyABINet-LVst
Evaluation Results
| Method | Links | |
|---|---|---|
| ABINet-LVstLabeled Datasets=MJ+ST, Unlabeled Datasets=Uber-Text2021.03 | 90.1 | |
| ABINet-LVestLabeled Datasets=MJ+ST, Unlabeled Datasets=Uber-Text2021.03 | 89.9 | |
| SVTR-BBackbone=Base2022.04 | 89.9 | |
| ABINet-LVLabeled Datasets=MJ+ST2021.03 | 89.3 | |
| ABINetLanguage Mode=Lan-aware2022.04 | 89.3 | |
| ABINetTraining Datasets=MJ and ST2023.06 | 89.3 | |
| DiffusionSTRTraining Datasets=MJ and ST2023.06 | 89.2 | |
| PARSeqTraining Datasets=MJ and ST2023.06 | 88.9 | |
| VSTLanguage Mode=Lan-aware2022.04 | 88.7 | |
| SVTR-LBackbone=Large2022.04 | 88.4 | |
| SRN-LV (Reproduced)Labeled Datasets=MJ+ST2021.03 | 87.9 | |
| SVTR-SBackbone=Small2022.04 | 87.9 | |
| PREN2DLanguage Mode=Lan-aware2022.04 | 87.6 | |
| ABINet-SVLabeled Datasets=MJ+ST2021.03 | 87 | |
| VST*Language Mode=Lan-free2022.04 | 87 | |
| SRN-SV (Reproduced)Labeled Datasets=MJ+ST2021.03 | 86.4 | |
| VisionLANLanguage Mode=Lan-aware2022.04 | 86 | |
| SVTR-TBackbone=Tiny2022.04 | 85.4 | |
| SRNLabeled Datasets=MJ+ST2021.03 | 85.1 | |
| SRNLanguage Mode=Lan-aware2022.04 | 85.1 | |
| TextscannerLabeled Datasets=MJ+ST2021.03 | 84.3 | |
| ABINet*Language Mode=Lan-free2022.04 | 84.2 | |
| PREN*Language Mode=Lan-free2022.04 | 83.9 | |
| ParallelLabeled Datasets=MJ+ST2021.03 | 82.3 | |
| SAMLabeled Datasets=MJ+ST2021.03 | 82.2 | |
| DOTA + CRFRefinement=CRF2026.03 | 82.17 | |
| DOTA2026.03 | 82.02 | |
| RES50-DEF(L4)-ViT-AdaptiveDeformable Layer=L4, Config=Adaptive2026.03 | 81.86 | |
| ViTSTRLanguage Mode=Lan-free2022.04 | 81.8 | |
| ViTSTRTraining Datasets=MJ and ST2023.06 | 81.8 | |
| AutoSTRLanguage Mode=Lan-aware2022.04 | 81.7 | |
| SE-ASTER (Ours)lexicon=lexicon-free, training_annotations=word-level2020.05 | 81.4 | |
| SE-ASTERLabeled Datasets=MJ+ST2021.03 | 81.4 | |
| RES50-ViT-PEConfig=Positional Encoding2026.03 | 81.09 | |
| Yang et al. [53]*lexicon=lexicon-free, training_annotations=word-level and character-level2020.05 | 80.8 | |
| DANLabeled Datasets=MJ+ST2021.03 | 80 | |
| Zhan et al. [57]lexicon=lexicon-free, training_annotations=word-level2020.05 | 79.6 | |
| RobustScannerLabeled Datasets=MJ+ST2021.03 | 79.5 | |
| TRBATraining Datasets=MJ and ST2023.06 | 79.5 | |
| SRN*Language Mode=Lan-free2022.04 | 79.4 | |
| ASTER baseline reproducedlexicon=lexicon-free, training_annotations=word-level2020.05 | 79.1 | |
| Liu et al. [28]*lexicon=lexicon-free, training_annotations=word-level and character-level2020.05 | 78.9 | |
| ASTER [45]lexicon=lexicon-free, training_annotations=word-level2020.05 | 78.5 | |
| ASTERLanguage Mode=Lan-aware2022.04 | 78.5 | |
| RES50-ATT-ViT2026.03 | 78.14 | |
| RES50-ViTIteration=22026.03 | 77.05 | |
| Li et al. [23]lexicon=lexicon-free, training_annotations=word-level2020.05 | 76.4 | |
| SARLanguage Mode=Lan-aware2022.04 | 76.4 | |
| RES50-ViT2026.03 | 76.28 | |
| Luo et al. [32]lexicon=lexicon-free, training_annotations=word-level2020.05 | 76.1 | |
| MORANLanguage Mode=Lan-aware2022.04 | 76.1 | |
| Yang et al. [54]*lexicon=lexicon-free, training_annotations=word-level and character-level2020.05 | 75.8 | |
| Liu et al. [29]lexicon=lexicon-free, training_annotations=word-level2020.05 | 73.9 | |
| RosettaLanguage Mode=Lan-free2022.04 | 73.8 | |
| Cheng et al. [8]lexicon=lexicon-free, training_annotations=word-level2020.05 | 73 | |
| RES50-DEF-(L3–L4)Deformable Layers=L3-L42026.03 | 72.25 | |
| RES50-ATT-AdaptiveConfig=Adaptive2026.03 | 72.25 | |
| Shi et al. [44]lexicon=lexicon-free, training_annotations=word-level2020.05 | 71.8 | |
| Xie et al. [52]lexicon=lexicon-free, training_annotations=word-level2020.05 | 70.1 | |
| CRNNLanguage Mode=Lan-free2022.04 | 70 | |
| RES50-ATTIteration=22026.03 | 69.61 | |
| ResNext2026.03 | 67.6 | |
| CRNNTraining Datasets=MJ and ST2023.06 | 65.9 | |
| RES50-ATT2026.03 | 59.84 |