Scene Text Recognition on SVT (test)
99.2Word AccuracyShi et al. 2018
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Shi et al. 2018Lexicon=50, Training Level=word-level2018.11 | 99.2 | — | |
| DTrOCRTraining data=Real2023.08 | 98.9 | — | |
| CLIP4STR-B*Training data=REBU-Syn2023.12 | 98.76 | — | |
| CLIP4STR-L*Training data=REBU-Syn2023.12 | 98.61 | — | |
| SARLexicon=50, Training Level=word-level2018.11 | 98.5 | — | |
| ABINetTraining data=Real2023.12 | 98.2 | — | |
| CLIP4STR-LTraining data=Real2023.12 | 98.15 | — | |
| PARSeqTraining data=Real2023.08 | 97.9 | — | |
| ABINetTraining data=Real2023.08 | 97.8 | — | |
| MAERec-BTraining data=Union14M-L2023.12 | 97.8 | — | |
| CLIP4STR-BTraining data=Real2023.12 | 97.68 | — | |
| Shi, Bai, and YaoLexicon Size=502019.01 | 97.5 | — | |
| Shi, Bai, and Yao 2017Lexicon=50, Training Level=word-level2018.11 | 97.5 | — | |
| PARSeqATraining data=Real, Reproduced=true2023.12 | 97.37 | — | |
| Cheng et al. 2017Lexicon=50, Training Level=word-level and character-level2018.11 | 97.1 | — | |
| TRBATraining data=Real2023.08 | 97 | — | |
| DTrOCRTraining data=Synth2023.08 | 96.9 | — | |
| MaskOCR(ViT-L)Training data=MJ+ST2023.12 | 96.9 | — | |
| MORANLexicon Size=502019.01 | 96.6 | — | |
| Bai et al. 2018Lexicon=50, Training Level=word-level and character-level2018.11 | 96.6 | — | |
| DiG-ViT-BTraining data=Real2023.12 | 96.5 | — | |
| Shi et al.Lexicon Size=502016.03 | 96.4 | — | |
| Lee and OsinderoLexicon Size=502019.01 | 96.3 | — | |
| Lee and Osindero 2016Lexicon=50, Training Level=word-level2018.11 | 96.3 | — | |
| Wang and Hu 2017Lexicon=50, Training Level=word-level2018.11 | 96.3 | — | |
| RARE (SRN only)Lexicon Size=502016.03 | 96.1 | — | |
| TrOCRLargeTraining data=MJ+ST+B2023.12 | 96.1 | — | |
| Cheng et al.Lexicon Size=50, Citation ID=[6]2019.01 | 96 | — | |
| Cheng et al. 2018Lexicon=50, Training Level=word-level2018.11 | 96 | — | |
| ViTSTR-STraining data=Real2023.12 | 96 | — | |
| Cheng et al.Lexicon Size=50, Citation ID=[5]2019.01 | 95.7 | — | |
| RARELexicon Size=502016.03 | 95.5 | — | |
| Shi et al.Lexicon Size=50, Citation ID=[42]2019.01 | 95.5 | — | |
| Liu et al.Lexicon Size=502019.01 | 95.5 | — | |
| Shi et al. 2016Lexicon=50, Training Level=word-level2018.11 | 95.5 | — | |
| Liu et al. 2016Lexicon=50, Training Level=word-level2018.11 | 95.5 | — | |
| Jaderberg et al. [17]Lexicon Size=502016.03 | 95.4 | — | |
| Jaderberg et al.Lexicon Size=50, Citation ID=[22]2019.01 | 95.4 | — | |
| Jaderberg et al. 2015aLexicon=50, Training Level=word-level2018.11 | 95.4 | — | |
| Yang et al.Lexicon Size=502019.01 | 95.2 | — | |
| Yang et al. 2017Lexicon=50, Training Level=word-level and character-level2018.11 | 95.2 | — | |
| Liu et al. 2018Lexicon=50, Training Level=word-level and character-level2018.11 | 95.2 | — | |
| Yin et al.Lexicon Size=502019.01 | 95.1 | — | |
| MATRN2021.11 | 95 | — | |
| MATRNTraining data=MJ+ST2023.12 | 95 | — | |
| MGP-STR FuseTraining Data=MJ and ST, Lexicon=None2022.09 | 94.74 | — | |
| MaskOCR_BASETraining data=Synth2023.08 | 94.7 | — | |
| MaskOCR(ViT-B)Training data=MJ+ST2023.12 | 94.7 | — | |
| NRTR+TPS++Training Data=90K+ST, Params=35.5M, Time=218ms2023.05 | 94.6 | — | |
| DiG-ViT-BTraining data=MJ+ST2023.12 | 94.6 | — | |
| ABINet-LV+TPS++Training Data=90K+ST, Params=37.2M, Time=41.5ms2023.05 | 94.3 | — | |
| GTRTraining Data=90K+ST, Params=42.1M, Time=18.8ms2023.05 | 94.1 | — | |
| MaskOCR_LARGETraining data=Synth2023.08 | 94.1 | — | |
| S-GTRYear=2022, Training Data=90K+ST2021.11 | 94.1 | — | |
| PREN2DYear=20212021.11 | 94 | — | |
| Pren2DTraining Data=90K+ST+Real, Time=67.4ms2023.05 | 94 | — | |
| PREN2DTraining data=Synth2023.08 | 94 | — | |
| Pren2DYear=2021, Training Data=90K+ST+Real2021.11 | 94 | — | |
| ABINet (reproduced)Implementation=reproduced2021.11 | 93.7 | — | |
| Shi et al. 2018Lexicon=None, Training Level=word-level2018.11 | 93.6 | — | |
| ASTERFeature map=STN-1D, Training data=MJ+ST, Dictionary matching=False2019.10 | 93.6 | — | |
| PARSeqTraining Data=90K+ST, Params=23.8M, Time=11.8ms2023.05 | 93.6 | — | |
| DiffusionSTRTraining data=Synth2023.08 | 93.6 | — | |
| PARSeqTraining data=Synth2023.08 | 93.6 | — | |
| CDistNet w/o TPSTraining Data=90K+ST2021.11 | 93.6 | — | |
| PARSeqATraining data=MJ+ST2023.12 | 93.6 | — | |
| ABINet + InfoBatchPruned %=26.6%2025.03 | 93.6 | — | |
| He et al. 2016bLexicon=50, Training Level=word-level2018.11 | 93.5 | — | |
| ABINetYear=20212021.11 | 93.5 | — | |
| ABINetTraining Data=MJ and ST, Lexicon=None2022.09 | 93.5 | — | |
| ABINet-LVTraining Data=90K+ST, Params=36.7M, Time=37.2ms2023.05 | 93.5 | — | |
| ABINetTraining data=Synth2023.08 | 93.5 | — | |
| ABINetYear=2021, Training Data=90K+ST2021.11 | 93.5 | — | |
| CDistNetTraining Data=90K+ST2021.11 | 93.5 | — | |
| ABINetTraining data=MJ+ST2023.12 | 93.5 | — | |
| ABINetPruned %=0%2025.03 | 93.4 | — | |
| ABINet + SeTaPruned %=28.1%2025.03 | 93.4 | — | |
| Jaderberg et al. [16]Lexicon Size=502016.03 | 93.2 | — | |
| Jaderberg et al.Lexicon Size=50, Citation ID=[21]2019.01 | 93.2 | — | |
| MGP-STR VisionTraining Data=MJ and ST, Lexicon=None2022.09 | 93.2 | — | |
| TrOCR_LARGETraining data=Synth2023.08 | 93.2 | — | |
| ABINet + InfoBatchPruned %=38.1%2025.03 | 93.2 | — | |
| ABINet + InfoBatchPruned %=50.3%2025.03 | 93.2 | — | |
| ABINet + SeTaPruned %=40.4%2025.03 | 93.2 | — | |
| ABINet + SeTaPruned %=71.0%2025.03 | 92.9 | — | |
| TextScannerTraining data=Synth2023.08 | 92.7 | — | |
| TextScannerTraining data=MJ+ST2023.12 | 92.7 | — | |
| PETRTraining data=MJ+ST2023.12 | 92.4 | — | |
| PlugNetTraining data=Synth2023.08 | 92.3 | — | |
| JVSRYear=20212021.11 | 92.2 | — | |
| JVSRTraining data=Synth2023.08 | 92.2 | — | |
| JVSRYear=2021, Training Data=90K+ST2021.11 | 92.2 | — | |
| STAR-NetTrain Dataset=S90k+ST+TextOCR2021.05 | 92.12 | — | |
| GordoLexicon Size=502016.03 | 91.8 | — | |
| GordoLexicon Size=502019.01 | 91.8 | — | |
| Mask TextSpotterLexicon-free=true, Character-level annotations=true2021.06 | 91.8 | — | |
| RCEEDLexicon-free=true, Character-level annotations=false2021.06 | 91.8 | — | |
| VisionLANYear=20212021.11 | 91.7 | — | |
| VisionLANTraining Data=MJ and ST, Lexicon=None2022.09 | 91.7 | — | |
| VisionLANTraining Data=90K+ST, Params=32.8M, Time=28.0ms2023.05 | 91.7 | — |