Scene Text Recognition on CUTE 288 samples (test)
99.65Word AccuracyPARSeq-S
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| PARSeq-SParameters=22M, Train Data=REBU-Syn2023.12 | 99.65 | — | |
| PARSeq-HParameters=0.6B, Train Data=REBU-Syn2023.12 | 99.65 | — | |
| PARSeq-BTrain Data=REBU-Syn2023.12 | 99.31 | — | |
| PARSeq-LTrain Data=REBU-Syn2023.12 | 99.31 | — | |
| DTrOCRTraining data=Real2023.08 | 99.1 | — | |
| PARSeq_ATrain data=R2022.07 | 98.3 | — | |
| PARSeqTraining data=Real2023.08 | 98.3 | — | |
| TRBATrain data=R2022.07 | 97.7 | — | |
| ABINetTrain data=R2022.07 | 97.7 | — | |
| PARSeq_NTrain data=R2022.07 | 97.7 | — | |
| TRBATraining data=Real2023.08 | 97.7 | — | |
| ABINetTraining data=Real2023.08 | 97.7 | — | |
| DTrOCRTraining data=Synth2023.08 | 97.6 | — | |
| VITSTR-STrain data=R2022.07 | 96.1 | — | |
| TrOCRBackbone=TrOCR-Large, Training Data=Synthetic + Benchmark2021.09 | 95.1 | — | |
| SVTR_LARGETraining data=Synth2023.08 | 95.1 | — | |
| ABINet++ (LV)Labeled Datasets=MJ+ST+Real, Vision Model Scale=LV2022.11 | 94.4 | — | |
| ABINet++‡ (LV)Labeled Datasets=MJ+ST, Unlabeled Datasets=Uber-Text, Vision Model Scale=LV, Training Strategy=ensemble self-training2022.11 | 94.1 | — | |
| ABINet++† (LV)Labeled Datasets=MJ+ST, Unlabeled Datasets=Uber-Text, Vision Model Scale=LV, Training Strategy=self-training2022.11 | 93.4 | — | |
| MaskOCRBackbone=ViT-L2021.09 | 92.7 | — | |
| PIMNet Qiao et al.Year=2021, Labeled Datasets=MJ+ST+Real2022.11 | 92.7 | — | |
| MaskOCR_LARGETraining data=Synth2023.08 | 92.7 | — | |
| DiffusionSTRTraining Datasets=MJ and ST2023.06 | 92.5 | — | |
| DiffusionSTRTraining data=Synth2023.08 | 92.5 | — | |
| Robust ScannerTrain data=S,B2022.07 | 92.4 | — | |
| RobustScanner2021.09 | 92.4 | — | |
| PARSeq_ATrain data=S2022.07 | 92.2 | — | |
| PARSeq2021.09 | 92.2 | — | |
| GTC Hu et al.Year=2020, Labeled Datasets=MJ+ST+Real2022.11 | 92.2 | — | |
| PARSeqTraining Datasets=MJ and ST2023.06 | 92.2 | — | |
| PARSeqTraining data=Synth2023.08 | 92.2 | — | |
| RCEEDTrain data=S,B2022.07 | 91.7 | — | |
| PREN2DTrain data=S2022.07 | 91.7 | — | |
| RCEED2021.09 | 91.7 | — | |
| PREN2D2021.09 | 91.7 | — | |
| PREN2DTraining data=Synth2023.08 | 91.7 | — | |
| SVTR_BASETraining data=Synth2023.08 | 91.7 | — | |
| TextScannerTrain data=S*2022.07 | 91.6 | — | |
| TextScanner2021.09 | 91.6 | — | |
| Textscanner Wan et al.Year=2020, Labeled Datasets=MJ+ST+Real2022.11 | 91.6 | — | |
| TextScannerTraining data=Synth2023.08 | 91.6 | — | |
| PARSeq_NTrain data=S2022.07 | 91.4 | — | |
| TRBATrain data=S2022.07 | 91.3 | — | |
| TrOCRBackbone=TrOCR-Base, Training Data=Synthetic + Benchmark2021.09 | 90.6 | — | |
| RobustScanner Yue et al.Year=2020, Labeled Datasets=MJ+ST2022.11 | 90.3 | 36 | |
| Bhunia et al.Train data=S2022.07 | 89.7 | — | |
| CVAE-FeedTrain data=S2022.07 | 89.7 | — | |
| ABINetTrain data=S2022.07 | 89.7 | — | |
| Bhunia2021.09 | 89.7 | — | |
| CVAE-Feed.2021.09 | 89.7 | — | |
| Bhunia et al.Year=2021, Labeled Datasets=MJ+ST2022.11 | 89.7 | — | |
| JVSRTraining data=Synth2023.08 | 89.7 | — | |
| CVAE-FeedTraining data=Synth2023.08 | 89.7 | — | |
| TrOCRBackbone=TrOCR-Large, Training Data=Synthetic2021.09 | 89.6 | — | |
| TrOCR_LARGETraining data=Synth2023.08 | 89.6 | — | |
| ABINetTrain data=S32022.07 | 89.2 | — | |
| ABINet2021.09 | 89.2 | — | |
| MaskOCRBackbone=ViT-B2021.09 | 89.2 | — | |
| ABINet++ (LV)Labeled Datasets=MJ+ST, Vision Model Scale=LV2022.11 | 89.2 | 33.9 | |
| ABINetTraining Datasets=MJ and ST2023.06 | 89.2 | — | |
| ABINetTraining data=Synth2023.08 | 89.2 | — | |
| MaskOCR_BASETraining data=Synth2023.08 | 89.2 | — | |
| CRNNTrain data=R2022.07 | 89.1 | — | |
| CRNNTraining data=Real2023.08 | 89.1 | — | |
| ABINet++ (SV)Labeled Datasets=MJ+ST, Vision Model Scale=SV2022.11 | 88.9 | 31.6 | |
| VisionLANTrain data=S2022.07 | 88.5 | — | |
| VisionLAN2021.09 | 88.5 | — | |
| VisionLan Wang et al.Year=2021, Labeled Datasets=MJ+ST2022.11 | 88.5 | 11.5 | |
| VisionLANTraining data=Synth2023.08 | 88.5 | — | |
| VITSTR-STrain data=S2022.07 | 88.2 | — | |
| SRN* (LV)Labeled Datasets=MJ+ST, Reproduced=true, Vision Model Scale=LV2022.11 | 88.2 | 26.9 | |
| SRNTrain data=S2022.07 | 87.8 | — | |
| SRN2021.09 | 87.8 | — | |
| SRN Yu et al.Year=2020, Labeled Datasets=MJ+ST2022.11 | 87.8 | 46.2 | |
| SRNTraining data=Synth2023.08 | 87.8 | — | |
| SRN* (SV)Labeled Datasets=MJ+ST, Reproduced=true, Vision Model Scale=SV2022.11 | 87.5 | 24.2 | |
| TrOCRBackbone=TrOCR-Base, Training Data=Synthetic2021.09 | 86.8 | — | |
| TrOCR_BASETraining data=Synth2023.08 | 86.8 | — | |
| PlugNetTrain data=S2022.07 | 85 | — | |
| PlugNet2021.09 | 85 | — | |
| PlugNetTraining data=Synth2023.08 | 85 | — | |
| DAN Wang et al.Year=2020, Labeled Datasets=MJ+ST2022.11 | 84.4 | — | |
| PIMNet Qiao et al.Year=2021, Labeled Datasets=MJ+ST2022.11 | 84.4 | 19.1 | |
| SEED Qiao et al.Year=2020, Labeled Datasets=MJ+ST2022.11 | 83.6 | 53.3 | |
| Textscanner Wan et al.Year=2020, Labeled Datasets=MJ+ST2022.11 | 83.3 | — | |
| VITSTR-BTrain data=S22022.07 | 81.3 | — | |
| VITSTR-B2021.09 | 81.3 | — | |
| ViTSTRTraining Datasets=MJ and ST2023.06 | 81.3 | — | |
| ViTSTR_BASETraining data=Synth2023.08 | 81.3 | — | |
| CRNNTrain data=S2022.07 | 78.7 | — | |
| TRBATrain data=S2022.07 | 78.2 | — | |
| TRBA2021.09 | 78.2 | — | |
| TRBATraining Datasets=MJ and ST2023.06 | 78.2 | — | |
| TRBATraining data=Synth2023.08 | 78.2 | — | |
| CRNNTraining Datasets=MJ and ST2023.06 | 61.5 | — | |
| CRNNTraining data=Synth2023.08 | 61.5 | — | |
| CRNNTrain data=S2022.07 | 61.3 | — | |
| CRNN2021.09 | 61.3 | — |