Scene Text Recognition on IC13 (test)
99.42Word AccuracyCLIP4STR-L*
Evaluation Results
| Method | Links | |
|---|---|---|
| CLIP4STR-L*Training data=REBU-Syn2023.12 | 99.42 | |
| DTrOCRTraining data=Real2023.08 | 99.4 | |
| CLIP4STR-B*Training data=REBU-Syn2023.12 | 99.29 | |
| DTrOCRTraining data=Real2023.08 | 99.1 | |
| DTrOCRTraining data=Synth2023.08 | 98.8 | |
| CLIP4STR-LTraining data=Real2023.12 | 98.48 | |
| PARSeq_ATrain data=R2022.07 | 98.4 | |
| PARSeqTraining data=Real2023.08 | 98.4 | |
| CLIP4STR-BTraining data=Real2023.12 | 98.36 | |
| PARSeq_ATrain data=R2022.07 | 98.3 | |
| TrOCR_LARGETraining data=Synth2023.08 | 98.3 | |
| PARSeqTraining data=Real2023.08 | 98.3 | |
| MaskOCR(ViT-L)Training data=MJ+ST2023.12 | 98.2 | |
| PARSeq_NTrain data=R2022.07 | 98.1 | |
| MaskOCR_BASETraining data=Synth2023.08 | 98.1 | |
| MaskOCR(ViT-B)Training data=MJ+ST2023.12 | 98.1 | |
| MAERec-BTraining data=Union14M-L2023.12 | 98.1 | |
| ABINetTrain data=R2022.07 | 98 | |
| PARSeq_NTrain data=R2022.07 | 98 | |
| ABINetTraining data=Real2023.08 | 98 | |
| ABINetTraining data=Real2023.08 | 98 | |
| ABINetTraining data=Real2023.12 | 98 | |
| ABINetTrain data=R2022.07 | 97.8 | |
| ABINet-LV+TPS++Training Data=90K+ST, Params=37.2M, Time=41.5ms2023.05 | 97.8 | |
| MaskOCR_LARGETraining data=Synth2023.08 | 97.8 | |
| DTrOCRTraining data=Synth2023.08 | 97.8 | |
| ViTSTR-STraining data=Real2023.12 | 97.8 | |
| VITSTR-STrain data=R2022.07 | 97.7 | |
| VITSTR-STrain data=R2022.07 | 97.6 | |
| TRBATrain data=R2022.07 | 97.6 | |
| TRBATrain data=R2022.07 | 97.6 | |
| TRBATraining data=Real2023.08 | 97.6 | |
| TRBATraining data=Real2023.08 | 97.6 | |
| DiG-ViT-BTraining data=Real2023.12 | 97.6 | |
| ABINetTrain data=S32022.07 | 97.4 | |
| TrOCRBackbone=TrOCR-Base, Training Data=Synthetic + Benchmark2021.09 | 97.4 | |
| ABINet-LVTraining Data=90K+ST, Params=36.7M, Time=37.2ms2023.05 | 97.4 | |
| ABINetTraining data=Synth2023.08 | 97.4 | |
| ABINetYear=2021, Training Data=90K+ST2021.11 | 97.4 | |
| CDistNet w/o TPSTraining Data=90K+ST2021.11 | 97.4 | |
| CDistNetTraining Data=90K+ST2021.11 | 97.4 | |
| ABINetTraining data=MJ+ST2023.12 | 97.4 | |
| PARSeqATraining data=Real, Reproduced=true2023.12 | 97.32 | |
| TrOCRBackbone=TrOCR-Large, Training Data=Synthetic + Benchmark2021.09 | 97.3 | |
| TrOCR_BASETraining data=Synth2023.08 | 97.3 | |
| TrOCRLargeTraining data=MJ+ST+B2023.12 | 97.3 | |
| SVTR_LARGETraining data=Synth2023.08 | 97.2 | |
| ABINetTrain data=S2022.07 | 97.1 | |
| DiffusionSTRTraining data=Synth2023.08 | 97.1 | |
| SVTR_BASETraining data=Synth2023.08 | 97.1 | |
| PARSeq_ATrain data=S2022.07 | 97 | |
| TrOCRBackbone=TrOCR-Large, Training Data=Synthetic2021.09 | 97 | |
| PARSeqTraining Data=90K+ST, Params=23.8M, Time=11.8ms2023.05 | 97 | |
| PARSeqTraining data=Synth2023.08 | 97 | |
| TrOCR_LARGETraining data=Synth2023.08 | 97 | |
| PETRTraining data=MJ+ST2023.12 | 97 | |
| DiG-ViT-BTraining data=MJ+ST2023.12 | 96.9 | |
| GTRTraining Data=90K+ST, Params=42.1M, Time=18.8ms2023.05 | 96.8 | |
| S-GTRYear=2022, Training Data=90K+ST2021.11 | 96.8 | |
| NRTR+TPS++Training Data=90K+ST, Params=35.5M, Time=218ms2023.05 | 96.6 | |
| PREN2DTrain data=S2022.07 | 96.4 | |
| Pren2DTraining Data=90K+ST+Real, Time=67.4ms2023.05 | 96.4 | |
| PREN2DTraining data=Synth2023.08 | 96.4 | |
| DiffusionSTRTraining data=Synth2023.08 | 96.4 | |
| Pren2DYear=2021, Training Data=90K+ST+Real2021.11 | 96.4 | |
| STN-CSTRTrain data=S2022.07 | 96.3 | |
| TRBATrain data=S2022.07 | 96.3 | |
| PARSeq_NTrain data=S2022.07 | 96.3 | |
| TrOCRBackbone=TrOCR-Base, Training Data=Synthetic2021.09 | 96.3 | |
| TrOCR_BASETraining data=Synth2023.08 | 96.3 | |
| PARSeq_ATrain data=S2022.07 | 96.2 | |
| PARSeq2021.09 | 96.2 | |
| PARSeqTraining data=Synth2023.08 | 96.2 | |
| PARSeqATraining data=MJ+ST2023.12 | 96.2 | |
| NRTRTraining Data=90K+ST, Params=31.7M, Time=212ms2023.05 | 95.8 | |
| NRTRYear=2019, Training Data=90K+ST2021.11 | 95.8 | |
| MATRNTraining data=MJ+ST2023.12 | 95.8 | |
| VisionLANTrain data=S2022.07 | 95.7 | |
| CVAE-FeedTrain data=S2022.07 | 95.7 | |
| CVAE-Feed.2021.09 | 95.7 | |
| VisionLANTraining Data=90K+ST, Params=32.8M, Time=28.0ms2023.05 | 95.7 | |
| VisionLANTraining data=Synth2023.08 | 95.7 | |
| CVAE-FeedTraining data=Synth2023.08 | 95.7 | |
| VisionLANYear=2021, Training Data=90K+ST2021.11 | 95.7 | |
| VisionLANTraining data=MJ+ST2023.12 | 95.7 | |
| SRNBackbone=ResNet, Training Data=Synth90K+SynthText, Annotations=word-level, Lexicon=None2020.03 | 95.5 | |
| SRNTrain data=S2022.07 | 95.5 | |
| Bhunia et al.Train data=S2022.07 | 95.5 | |
| PARSeq_NTrain data=S2022.07 | 95.5 | |
| Bhunia2021.09 | 95.5 | |
| SRNTraining Data=90K+ST, Params=49.3M, Time=26.9ms2023.05 | 95.5 | |
| SRNTraining data=Synth2023.08 | 95.5 | |
| JVSRTraining data=Synth2023.08 | 95.5 | |
| SRNYear=2020, Training Data=90K+ST2021.11 | 95.5 | |
| JVSRYear=2021, Training Data=90K+ST2021.11 | 95.5 | |
| SRNTraining data=MJ+ST2023.12 | 95.5 | |
| STAR-NetTrain Dataset=S90k+ST+TextOCR2021.05 | 95.33 | |
| Liao et al. (SAM)Backbone=ResNet, Training Data=Synth90K+SynthText, Annotations=word-level, Lexicon=None2020.03 | 95.3 | |
| MaskTextSpotterYear=2019, Training Data=90K+ST2021.11 | 95.3 | |
| VITSTR-STrain data=S2022.07 | 95.1 |