Scene Text Recognition on IIIT5K (test)
99.6Word AccuracyCheng et al.
Evaluation Results
| Method | Links | |
|---|---|---|
| Cheng et al.Lexicon Size=50, Citation ID=[6]2019.01 | 99.6 | |
| Cheng et al. 2018Lexicon=50, Training Level=word-level2018.11 | 99.6 | |
| Shi et al. 2018Lexicon=50, Training Level=word-level2018.11 | 99.6 | |
| Bai et al. 2018Lexicon=50, Training Level=word-level and character-level2018.11 | 99.5 | |
| CLIP4STR-LTraining data=Real2023.12 | 99.43 | |
| SARLexicon=50, Training Level=word-level2018.11 | 99.4 | |
| CLIP4STR-LTraining Data=3.3M real samples, Backbone=ViT-L2023.05 | 99.4 | |
| Cheng et al. 2017Lexicon=50, Training Level=word-level and character-level2018.11 | 99.3 | |
| MPSTRTraining Data=3.3M real samples2023.05 | 99.2 | |
| CLIP4STR-BTraining Data=3.3M real samples, Backbone=ViT-B2023.05 | 99.2 | |
| CLIP4STR-L*Training data=REBU-Syn2023.12 | 99.13 | |
| CLIP4STR-B*Training data=REBU-Syn2023.12 | 98.96 | |
| Cheng et al.Lexicon Size=50, Citation ID=[5]2019.01 | 98.9 | |
| PARSeqTraining Data=3.3M real samples2023.05 | 98.9 | |
| Shi et al. 2018Lexicon=1k, Training Level=word-level2018.11 | 98.8 | |
| CLIP4STR-BTraining data=Real2023.12 | 98.73 | |
| Yin et al.Lexicon Size=502019.01 | 98.7 | |
| ABINetTraining data=Real2023.12 | 98.6 | |
| ABINetTraining Data=3.3M real samples, Backbone=ResNet-452023.05 | 98.6 | |
| MAERec-BTraining data=Union14M-L2023.12 | 98.5 | |
| SARLexicon=1k, Training Level=word-level2018.11 | 98.2 | |
| Cheng et al.Lexicon Size=1k, Citation ID=[6]2019.01 | 98.1 | |
| Cheng et al. 2018Lexicon=1k, Training Level=word-level2018.11 | 98.1 | |
| Wang and Hu 2017Lexicon=50, Training Level=word-level2018.11 | 98 | |
| MaskOCR(ViT-L)Training data=MJ+ST2023.12 | 98 | |
| MORANLexicon Size=502019.01 | 97.9 | |
| Bai et al. 2018Lexicon=1k, Training Level=word-level and character-level2018.11 | 97.9 | |
| ViTSTR-STraining data=Real2023.12 | 97.9 | |
| PARSeqATraining data=Real, Reproduced=true2023.12 | 97.87 | |
| Shi, Bai, and YaoLexicon Size=502019.01 | 97.8 | |
| Yang et al.Lexicon Size=502019.01 | 97.8 | |
| Shi, Bai, and Yao 2017Lexicon=50, Training Level=word-level2018.11 | 97.8 | |
| Yang et al. 2017Lexicon=50, Training Level=word-level and character-level2018.11 | 97.8 | |
| Liu et al.Lexicon Size=502019.01 | 97.7 | |
| Liu et al. 2016Lexicon=50, Training Level=word-level2018.11 | 97.7 | |
| Shi et al.Lexicon Size=502016.03 | 97.6 | |
| DiG-ViT-BTraining data=Real2023.12 | 97.6 | |
| Cheng et al. 2017Lexicon=1k, Training Level=word-level and character-level2018.11 | 97.5 | |
| Jaderberg et al. [17]Lexicon Size=502016.03 | 97.1 | |
| Jaderberg et al.Lexicon Size=50, Citation ID=[22]2019.01 | 97.1 | |
| Jaderberg et al. 2015aLexicon=50, Training Level=word-level2018.11 | 97.1 | |
| Liu et al. 2018Lexicon=50, Training Level=word-level and character-level2018.11 | 97 | |
| PARSeqTraining Data=90K+ST, Params=23.8M, Time=11.8ms2023.05 | 97 | |
| PARSeqATraining data=MJ+ST2023.12 | 97 | |
| Lee and OsinderoLexicon Size=502019.01 | 96.8 | |
| Cheng et al.Lexicon Size=1k, Citation ID=[5]2019.01 | 96.8 | |
| Lee and Osindero 2016Lexicon=50, Training Level=word-level2018.11 | 96.8 | |
| CDistNet w/o TPSTraining Data=90K+ST2021.11 | 96.7 | |
| DiG-ViT-BTraining data=MJ+ST2023.12 | 96.7 | |
| MATRNTraining data=MJ+ST2023.12 | 96.6 | |
| RARE (SRN only)Lexicon Size=502016.03 | 96.5 | |
| CDistNetTraining Data=90K+ST2021.11 | 96.4 | |
| NRTR+TPS++Training Data=90K+ST, Params=35.5M, Time=218ms2023.05 | 96.3 | |
| ABINet-LV+TPS++Training Data=90K+ST, Params=37.2M, Time=41.5ms2023.05 | 96.3 | |
| ABINet + InfoBatchPruned %=26.6%2025.03 | 96.3 | |
| RARELexicon Size=502016.03 | 96.2 | |
| Shi et al.Lexicon Size=50, Citation ID=[42]2019.01 | 96.2 | |
| MORANLexicon Size=1k2019.01 | 96.2 | |
| Shi et al. 2016Lexicon=50, Training Level=word-level2018.11 | 96.2 | |
| ABINet-LVTraining Data=90K+ST, Params=36.7M, Time=37.2ms2023.05 | 96.2 | |
| ABINetYear=2021, Training Data=90K+ST2021.11 | 96.2 | |
| ABINetTraining data=MJ+ST2023.12 | 96.2 | |
| ABINet + SeTaPruned %=28.1%2025.03 | 96.2 | |
| Yang et al.Lexicon Size=1k2019.01 | 96.1 | |
| Yin et al.Lexicon Size=1k2019.01 | 96.1 | |
| Yang et al. 2017Lexicon=1k, Training Level=word-level and character-level2018.11 | 96.1 | |
| ABINetPruned %=0%2025.03 | 96.1 | |
| ABINet + InfoBatchPruned %=38.1%2025.03 | 95.9 | |
| ABINet + SeTaPruned %=40.4%2025.03 | 95.9 | |
| VisionLANTraining Data=90K+ST, Params=32.8M, Time=28.0ms2023.05 | 95.8 | |
| GTRTraining Data=90K+ST, Params=42.1M, Time=18.8ms2023.05 | 95.8 | |
| VisionLANYear=2021, Training Data=90K+ST2021.11 | 95.8 | |
| S-GTRYear=2022, Training Data=90K+ST2021.11 | 95.8 | |
| VisionLANTraining data=MJ+ST2023.12 | 95.8 | |
| PETRTraining data=MJ+ST2023.12 | 95.8 | |
| MaskOCR(ViT-B)Training data=MJ+ST2023.12 | 95.8 | |
| ABINet + InfoBatchPruned %=50.3%2025.03 | 95.8 | |
| ABINet + SeTaPruned %=71.0%2025.03 | 95.8 | |
| TextScannerTraining data=MJ+ST2023.12 | 95.7 | |
| Wang and Hu 2017Lexicon=1k, Training Level=word-level2018.11 | 95.6 | |
| Pren2DTraining Data=90K+ST+Real, Time=67.4ms2023.05 | 95.6 | |
| Pren2DYear=2021, Training Data=90K+ST+Real2021.11 | 95.6 | |
| Jaderberg et al. [16]Lexicon Size=502016.03 | 95.5 | |
| Jaderberg et al.Lexicon Size=50, Citation ID=[21]2019.01 | 95.5 | |
| RobustScannerLexicon-free=true, Character-level annotations=false2021.06 | 95.4 | |
| SGBANetTraining Data=90K+ST2023.05 | 95.4 | |
| Mask TextSpotterLexicon-free=true, Character-level annotations=true2021.06 | 95.3 | |
| RobustScannerTraining Data=90K+ST2023.05 | 95.3 | |
| RobustScannerYear=2020, Training Data=90K+ST2021.11 | 95.3 | |
| SPINTraining Data=90K+ST2023.05 | 95.2 | |
| JVSRYear=2021, Training Data=90K+ST2021.11 | 95.2 | |
| Shi, Bai, and YaoLexicon Size=1k2019.01 | 95 | |
| Shi, Bai, and Yao 2017Lexicon=1k, Training Level=word-level2018.11 | 95 | |
| SARLexicon=None, Training Level=word-level2018.11 | 95 | |
| MASTERLexicon-free=true, Character-level annotations=false2021.06 | 95 | |
| SARLexicon-free=true, Character-level annotations=false2021.06 | 95 | |
| RCEEDLexicon-free=true, Character-level annotations=false2021.06 | 94.9 | |
| SRNBackbone=ResNet, Training Data=Synth90K+SynthText, Annotations=word-level, Lexicon=None2020.03 | 94.8 | |
| SRNTraining Data=90K+ST, Params=49.3M, Time=26.9ms2023.05 | 94.8 | |
| SRNYear=2020, Training Data=90K+ST2021.11 | 94.8 |