Named Entity Recognition on NER (test)
95.25F1 ScoreZEN (P)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| ZEN (P)Initialization=Pre-trained, Backbone Scale=Base2019.11 | 95.25 | — | |
| BERT-WWMWhole Word Masking=true, Backbone Scale=Base2019.11 | 95.1 | — | |
| ERNIE 1.0Backbone Scale=Base2019.11 | 95.1 | — | |
| BERT (P)Initialization=Pre-trained, Backbone Scale=Base2019.11 | 94.78 | — | |
| ZEN (R)Initialization=Random, Backbone Scale=Base2019.11 | 93.24 | — | |
| BERT (R)Initialization=Random, Backbone Scale=Base2019.11 | 93.12 | — | |
| CVT + Multi-Task + LargeSupervision Type=Semi-supervised, Learning Paradigm=Multi-task, Model Scale=Large2018.09 | 92.61 | — | |
| CVT + Multi-task + LargeSemi-supervised=true, Multi-task=true, Model Size=Large2018.09 | 92.6 | — | |
| CVT + Multi-TaskSupervision Type=Semi-supervised, Learning Paradigm=Multi-task, Model Scale=Standard2018.09 | 92.42 | — | |
| CVT + Multi-taskSemi-supervised=true, Multi-task=true2018.09 | 92.4 | — | |
| CVTSupervision Type=Semi-supervised, Learning Paradigm=Single-task, Model Scale=Standard2018.09 | 92.34 | — | |
| ELMo + Multi-taskSupervision Type=Semi-supervised, Learning Paradigm=Multi-task, Model Scale=Standard2018.09 | 92.32 | — | |
| ELMo + Multi-taskSemi-supervised=true, Multi-task=true2018.09 | 92.3 | — | |
| CVTSemi-supervised=true, Multi-task=false2018.09 | 92.3 | — | |
| ELMoSupervision Type=Semi-supervised, Learning Paradigm=Single-task, Model Scale=Standard2018.09 | 92.24 | — | |
| ELMoSource=Peters et al. (2018)2018.09 | 92.22 | — | |
| ELMoSemi-supervised=true, Multi-task=false2018.09 | 92.2 | — | |
| ELMo (our implementation)Semi-supervised=true, Multi-task=false2018.09 | 92.2 | — | |
| Word DropoutSupervision Type=Semi-supervised, Learning Paradigm=Single-task, Model Scale=Standard2018.09 | 92.14 | — | |
| Word DropoutSemi-supervised=true, Multi-task=false2018.09 | 92.1 | — | |
| TagLMSupervision Type=Supervised2018.09 | 91.93 | — | |
| TagLMSemi-supervised=true, Multi-task=false2018.09 | 91.9 | — | |
| Virtual Adversarial TrainingSemi-supervised=true, Multi-task=false2018.09 | 91.8 | — | |
| LM-LSTM-CNN-CRFSupervision Type=Supervised2018.09 | 91.71 | — | |
| LSTM-CNNSupervision Type=Supervised2018.09 | 91.62 | — | |
| LSTM-CNN-CRFSupervision Type=Supervised2018.09 | 91.21 | — | |
| SupervisedSemi-supervised=false, Multi-task=false2018.09 | 91.2 | — | |
| SupervisedSupervision Type=Supervised, Learning Paradigm=Single-task, Model Scale=Standard2018.09 | 91.16 | — | |
| Virtual Adversarial TrainingSupervision Type=Semi-supervised, Learning Paradigm=Single-task, Model Scale=Standard2018.09 | 91.15 | — | |
| ID-CNN-CRFSemi-supervised=false, Multi-task=false2018.09 | 90.7 | — | |
| ID-CNN-CRFSupervision Type=Supervised2018.09 | 90.65 | — | |
| RA-CAShots=250, Selection Strategy=Checkpoint averaging, Run Type=Ensemble2023.05 | 82.2 | 0.2 | |
| Ref.2023.05 | 78.25 | — | |
| TAPIR-Trfdelay=22023.05 | 78.04 | — | |
| TAPIR-Trfdelay=12023.05 | 76.85 | — | |
| TAPIR-LTdelay=22023.05 | 75.75 | — | |
| TAPIR-Trfdelay=02023.05 | 74.13 | — | |
| TAPIR-LTdelay=12023.05 | 73.79 | — | |
| TAPIR-LTdelay=02023.05 | 73.12 | — | |
| Our methodM (few-shot samples)=202023.05 | 72.9 | — | |
| NoisyTuneM (few-shot samples)=202023.05 | 72.7 | — | |
| Fine-tuning slow algorithm (FS)M (few-shot samples)=202023.05 | 72.7 | — | |
| Direct Fine-tuning (DF)M (few-shot samples)=202023.05 | 72.5 | — | |
| Fine-tuning fast algorithm (FF)M (few-shot samples)=202023.05 | 72.5 | — | |
| Our methodM (few-shot samples)=102023.05 | 71.7 | — | |
| Fine-tuning slow algorithm (FS)M (few-shot samples)=102023.05 | 71.2 | — | |
| NoisyTuneM (few-shot samples)=102023.05 | 70.7 | — | |
| Fine-tuning fast algorithm (FF)M (few-shot samples)=102023.05 | 70.7 | — | |
| Direct Fine-tuning (DF)M (few-shot samples)=102023.05 | 70.6 | — | |
| RA-CAShots=5, Selection Strategy=Checkpoint averaging, Run Type=Ensemble2023.05 | 70.3 | 1 | |
| RA-LASTShots=5, Selection Strategy=Last checkpoint, Run Type=Ensemble2023.05 | 69.7 | 1 | |
| Our methodM (few-shot samples)=52023.05 | 69.1 | — | |
| CAShots=5, Selection Strategy=Checkpoint averaging, Run Type=Single Run2023.05 | 69.1 | 1 | |
| Fine-tuning slow algorithm (FS)M (few-shot samples)=52023.05 | 68.5 | — | |
| Fine-tuning fast algorithm (FF)M (few-shot samples)=52023.05 | 68.3 | — | |
| Direct Fine-tuning (DF)M (few-shot samples)=52023.05 | 68.1 | — | |
| NoisyTuneM (few-shot samples)=52023.05 | 67.8 | — | |
| FILTERProtocol=Translate-train, Learning approach=Self-Teaching2020.09 | 67.7 | — | |
| FILTERProtocol=Translate-train2020.09 | 66.7 | — | |
| XLM-RProtocol=Cross-lingual zero-shot transfer2020.09 | 65.4 | — | |
| X-STILTSProtocol=Cross-lingual zero-shot transfer2020.09 | 64 | — | |
| Our methodM (few-shot samples)=02023.05 | 62.5 | — | |
| Fine-tuning slow algorithm (FS)M (few-shot samples)=02023.05 | 62.3 | — | |
| mBERTProtocol=Cross-lingual zero-shot transfer2020.09 | 62.2 | — | |
| Fine-tuning fast algorithm (FF)M (few-shot samples)=02023.05 | 62.1 | — | |
| Direct Fine-tuning (DF)M (few-shot samples)=02023.05 | 61.3 | — | |
| NoisyTuneM (few-shot samples)=02023.05 | 61.3 | — | |
| XLMProtocol=Cross-lingual zero-shot transfer2020.09 | 61.2 | — |