Scene Text Recognition on IC13
97.7AccuracyABINet-LVest
Evaluation Results
| Method | Links | |
|---|---|---|
| ABINet-LVestLabeled Datasets=MJ+ST, Unlabeled Datasets=Uber-Text2021.03 | 97.7 | |
| ABINet-LVLabeled Datasets=MJ+ST2021.03 | 97.4 | |
| ABINetTraining Datasets=MJ and ST2023.06 | 97.4 | |
| ABINet-LVstLabeled Datasets=MJ+ST, Unlabeled Datasets=Uber-Text2021.03 | 97.3 | |
| DiffusionSTRTraining Datasets=MJ and ST2023.06 | 97.1 | |
| PARSeqTraining Datasets=MJ and ST2023.06 | 97 | |
| ABINet-SVLabeled Datasets=MJ+ST2021.03 | 96.8 | |
| SRN-LV (Reproduced)Labeled Datasets=MJ+ST2021.03 | 96.8 | |
| Our model2021.07 | 96.8 | |
| SRN-SV (Reproduced)Labeled Datasets=MJ+ST2021.03 | 96.3 | |
| S-Attn/CTC + LMRect.=No, Encoder/Decoder=S-Attn/CTC + LM, Train Dataset=Internal2021.04 | 95.98 | |
| SRNLabeled Datasets=MJ+ST2021.03 | 95.5 | |
| Yu et al.Rect.=No, Encoder/Decoder=Attn / Semantic Attn., Train Dataset=MJ + ST2021.04 | 95.5 | |
| SAMLabeled Datasets=MJ+ST2021.03 | 95.3 | |
| Lu et al.Rect.=No, Encoder/Decoder=Global Context Attn / Tfmr Dec., Train Dataset=MJ + ST + SA2021.04 | 95.3 | |
| S-Attn/CTC + LMRect.=No, Encoder/Decoder=S-Attn/CTC + LM, Train Dataset=Internal + Public2021.04 | 95.25 | |
| RobustScannerLabeled Datasets=MJ+ST2021.03 | 94.8 | |
| SCATTERRect.=Yes, Encoder/Decoder=CNN / BiLSTM, Train Dataset=MJ + ST + SA2021.04 | 94.7 | |
| RCEED2021.07 | 94.7 | |
| S-Attn/CTCRect.=No, Encoder/Decoder=S-Attn/CTC, Train Dataset=Internal2021.04 | 94.43 | |
| Bai et al. [2]*lexicon=lexicon-free, training_annotations=word-level and character-level2020.05 | 94.4 | |
| TransformerRect.=No, Encoder/Decoder=Transformer, Train Dataset=Internal2021.04 | 94.34 | |
| SATRN2021.07 | 94.1 | |
| Liu et al. [29]lexicon=lexicon-free, training_annotations=word-level2020.05 | 94 | |
| TransformerRect.=No, Encoder/Decoder=Transformer, Train Dataset=Internal + Public2021.04 | 93.97 | |
| Yang et al. [53]*lexicon=lexicon-free, training_annotations=word-level and character-level2020.05 | 93.9 | |
| DANLabeled Datasets=MJ+ST2021.03 | 93.9 | |
| Wang et al.Rect.=No, Encoder/Decoder=FCN/GRU, Train Dataset=MJ + ST2021.04 | 93.9 | |
| SCATTERVenue=CVPR’20, Training Set=MJ+ST+Extra, RNN=Y, FPS=-2022.04 | 93.9 | |
| DANVenue=AAAI’20, Training Set=MJ+ST, RNN=Y, FPS=-2022.04 | 93.9 | |
| TransformerRect.=No, Encoder/Decoder=Transformer, Train Dataset=MJ + ST + Public2021.04 | 93.88 | |
| S-Attn/CTCRect.=No, Encoder/Decoder=S-Attn/CTC, Train Dataset=Internal + Public2021.04 | 93.7 | |
| S-Attn/CTC + LMRect.=No, Encoder/Decoder=S-Attn/CTC + LM, Train Dataset=MJ + ST + Public2021.04 | 93.61 | |
| Bleeker and de RijkeRect.=No, Encoder/Decoder=Transformer, Train Dataset=MJ + ST2021.04 | 93.4 | |
| Cheng et al. [7]*lexicon=lexicon-free, training_annotations=word-level and character-level2020.05 | 93.3 | |
| Qiao et al.Rect.=No, Encoder/Decoder=LSTM/LSTM + Attn, Train Dataset=MJ + ST2021.04 | 93.3 | |
| CSTR2021.07 | 93.2 | |
| ASTER2021.07 | 93.2 | |
| ViTSTRTraining Datasets=MJ and ST2023.06 | 93.2 | |
| Liu et al. [30]*lexicon=lexicon-free, training_annotations=word-level and character-level2020.05 | 92.9 | |
| TextscannerLabeled Datasets=MJ+ST2021.03 | 92.9 | |
| TextScanner*Venue=AAAI’20, Training Set=MJ+ST+Extra, RNN=N, FPS=-2022.04 | 92.9 | |
| SE-ASTER (Ours)lexicon=lexicon-free, training_annotations=word-level2020.05 | 92.8 | |
| SE-ASTERLabeled Datasets=MJ+ST2021.03 | 92.8 | |
| SEED2021.07 | 92.8 | |
| SEEDVenue=CVPR’20, Training Set=MJ+ST, RNN=Y, FPS=-2022.04 | 92.8 | |
| ParallelLabeled Datasets=MJ+ST2021.03 | 92.7 | |
| Luo et al. [32]lexicon=lexicon-free, training_annotations=word-level2020.05 | 92.4 | |
| Baek et al.2021.07 | 92.3 | |
| Comb.BestVenue=ICCV’19, Training Set=MJ+ST, RNN=Y, FPS=36.232022.04 | 92.3 | |
| Ours-LargeVenue=-, Training Set=MJ+ST, RNN=N, FPS=66.91/255+2022.04 | 92.21 | |
| S-Attn/CTCRect.=No, Encoder/Decoder=S-Attn/CTC, Train Dataset=MJ + ST + Public2021.04 | 92.15 | |
| ASTER [45]lexicon=lexicon-free, training_annotations=word-level2020.05 | 91.8 | |
| Sheng et al.Rect.=Yes, Encoder/Decoder=S-Attn/ Attn, Train Dataset=MJ + ST2021.04 | 91.8 | |
| Liao et al. [25]*lexicon=lexicon-free, training_annotations=word-level and character-level2020.05 | 91.5 | |
| CA-FCN*Venue=AAAI’19, Training Set=ST, RNN=N, FPS=452022.04 | 91.4 | |
| Zhan et al. [57]lexicon=lexicon-free, training_annotations=word-level2020.05 | 91.3 | |
| Liu et al. [28]*lexicon=lexicon-free, training_annotations=word-level and character-level2020.05 | 91.1 | |
| Li et al. [23]lexicon=lexicon-free, training_annotations=word-level2020.05 | 91 | |
| ASTER baseline reproducedlexicon=lexicon-free, training_annotations=word-level2020.05 | 90.9 | |
| Lee et al. [22]lexicon=lexicon-free, training_annotations=word-level2020.05 | 90 | |
| Shi et al.Rect.=Yes, Encoder/Decoder=LSTM/LSTM + Attn, Train Dataset=MJ + ST2021.04 | 89.75 | |
| Shi et al. [43]lexicon=lexicon-free, training_annotations=word-level2020.05 | 89.6 | |
| CRNNTraining Datasets=MJ and ST2023.06 | 89.4 | |
| RosettaVenue=KDD’18, Training Set=MJ+ST, RNN=N, FPS=212.762022.04 | 89 | |
| Shi et al. [44]lexicon=lexicon-free, training_annotations=word-level2020.05 | 88.6 |