Named Entity Recognition on CoNLL NER 2002/2003 (test)
86.2German F1 ScoreERNIE-M
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| ERNIE-MTraining protocol=Fine-tune on all dataset, Model size=Large2020.12 | 86.2 | 94.01 | 89.23 | 93.81 | 90.81 | |
| XLM-RTraining protocol=Fine-tune on all dataset, Model size=Large2020.12 | 84.6 | 92 | 89.52 | 91.6 | 89.43 | |
| ERNIE-MTraining protocol=Fine-tune on all dataset, Model size=Base2020.12 | 84.2 | 93.04 | 88.33 | 91.73 | 89.32 | |
| XLM-RTraining protocol=Fine-tune on all dataset, Model size=Base2020.12 | 83.17 | 91.08 | 87.28 | 89.09 | 87.66 | |
| MELMsamples_per_language=4002021.08 | 80.33 | 86.14 | 86.6 | 85.99 | 84.76 | |
| Code-Mix-esssamples_per_language=4002021.08 | 80.03 | 85.74 | 85.18 | 85.36 | 84.08 | |
| MELM-goldsamples_per_language=4002021.08 | 79.09 | 86.04 | 85.76 | 84.83 | 83.93 | |
| MulDAsamples_per_language=4002021.08 | 78.41 | 84.37 | 84.54 | 83.09 | 82.6 | |
| mLUKE-EModel scale=large, Input features=entity-based2021.10 | 78.3 | 94 | 81.4 | 83.5 | 84.3 | |
| MELMsamples_per_language=2002021.08 | 78.24 | 83.56 | 84.98 | 82.79 | 82.29 | |
| MELM-goldsamples_per_language=2002021.08 | 78.05 | 82.9 | 85.93 | 81 | 81.97 | |
| Code-Mix-randomsamples_per_language=4002021.08 | 77.91 | 85.04 | 84.44 | 83.56 | 82.74 | |
| Gold-Onlysamples_per_language=4002021.08 | 77.4 | 83.92 | 83.22 | 84.04 | 82.14 | |
| mLUKE-EModel scale=base, Input features=entity-based2021.10 | 77.2 | 93.6 | 77.7 | 81.8 | 82.6 | |
| Code-Mix-esssamples_per_language=2002021.08 | 76.64 | 83.34 | 83.02 | 82.27 | 81.07 | |
| mLUKE-WModel scale=large, Input features=word-based2021.10 | 76.5 | 92.3 | 80.7 | 82.6 | 83 | |
| Gold-Onlysamples_per_language=2002021.08 | 76.39 | 83.06 | 82.71 | 79.19 | 80.34 | |
| Code-Mix-randomsamples_per_language=2002021.08 | 75.7 | 82.86 | 83.13 | 79.08 | 80.19 | |
| XLM-RModel scale=base, Extra training=true2021.10 | 75.7 | 91.8 | 79.8 | 80.3 | 81.9 | |
| MELMsamples_per_language=1002021.08 | 75.61 | 80.96 | 81.47 | 80.14 | 79.54 | |
| mLUKE-WModel scale=base, Input features=word-based2021.10 | 75.1 | 91.6 | 79.2 | 80.2 | 81.5 | |
| XLM-RModel scale=large2021.10 | 75.1 | 92.5 | 80.5 | 82.9 | 82.8 | |
| AdvPickerBackbone=multilingual BERT (cased), Number of parameters=177M, zero-shot protocol=cross-lingual, Use of additional translation data=false, Training time=≈21min2021.06 | 75.01 | — | 79 | 82.9 | 78.97 | |
| MELM-goldsamples_per_language=1002021.08 | 74.79 | 78.71 | 81.25 | 78.85 | 78.4 | |
| MulDAsamples_per_language=2002021.08 | 74.57 | 82.32 | 82.73 | 79.06 | 79.67 | |
| XLM-RModel scale=base2021.10 | 74.3 | 91.5 | 79.8 | 80.7 | 81.6 | |
| mBERT-TLADVBackbone=multilingual BERT (cased), Number of parameters=178M, zero-shot protocol=cross-lingual, Training time=≈130min2021.06 | 73.89 | — | 76.92 | 80.62 | 77.14 | |
| UniTrans (Wu et al. 2020a)zero-shot protocol=cross-lingual, Use of additional translation data=false2021.06 | 73.61 | — | 77.3 | 81.2 | 77.37 | |
| XLM-K2021.10 | 73.3 | 90.7 | 76.6 | 80 | 80.1 | |
| Wu et al.zero-shot protocol=cross-lingual2021.06 | 73.16 | — | 76.75 | 80.44 | 76.78 | |
| ERNIE-MTraining protocol=Fine-tune on English dataset, Model size=Large2020.12 | 72.99 | 93.28 | 78.83 | 81.45 | 81.64 | |
| mBERT-ftBackbone=multilingual BERT (cased), zero-shot protocol=cross-lingual2021.06 | 72.59 | — | 75.12 | 80.34 | 76.02 | |
| Keung et al.zero-shot protocol=cross-lingual2021.06 | 71.9 | — | 74.3 | 77.6 | 74.6 | |
| Code-Mix-esssamples_per_language=1002021.08 | 71.56 | 79.55 | 79.58 | 76.49 | 76.8 | |
| XLM-RTraining protocol=Fine-tune on English dataset, Model size=Large2020.12 | 71.4 | 92.92 | 78.64 | 80.8 | 80.94 | |
| LIMode=Zero-shot (except English)2019.11 | 71.28 | 73.87 | 75.4 | 80.06 | — | |
| Code-Mix-randomsamples_per_language=1002021.08 | 70.58 | 77.38 | 78.61 | 76.45 | 75.75 | |
| MulDAsamples_per_language=1002021.08 | 70.47 | 73.67 | 75.53 | 72.4 | 73.02 | |
| BERT-ML^kMode=Zero-shot, Training=k languages2019.11 | 70.23 | 72.57 | 77.17 | 79.76 | — | |
| mBERT2021.10 | 70 | 89.7 | 77.1 | 75.2 | 78 | |
| Pires, Schlinger, and GarretteMode=Zero-shot (except English)2019.11 | 69.74 | 90.7 | 73.59 | 77.36 | — | |
| XLM-RTraining protocol=Fine-tune on English dataset, Model size=Base2020.12 | 69.6 | 92.25 | 76.53 | 78.08 | 79.11 | |
| Wu and Dredzezero-shot protocol=cross-lingual2021.06 | 69.56 | — | 74.96 | 77.57 | 73.57 | |
| mBERTTraining protocol=Fine-tune on English dataset, Model size=Base2020.12 | 69.56 | 91.97 | 74.96 | 77.57 | 78.52 | |
| BERT-SLMode=Zero-shot, Source Language=English2019.11 | 69.42 | — | 73.62 | 78.61 | — | |
| Gold-Onlysamples_per_language=1002021.08 | 69.35 | 75.62 | 75.85 | 74.33 | 73.79 | |
| CL+LIMode=Zero-shot (except English)2019.11 | 68.94 | 74.28 | 73.68 | 80.78 | — | |
| ERNIE-MTraining protocol=Fine-tune on English dataset, Model size=Base2020.12 | 68.08 | 92.78 | 79.37 | 78.01 | 79.56 | |
| PCMode=Zero-shot (except English)2019.11 | 67.91 | 73.72 | 73.68 | 80.28 | — | |
| CLMode=Zero-shot (except English)2019.11 | 67.04 | 73.3 | 75.29 | 81.76 | — | |
| PC+LIMode=Zero-shot (except English)2019.11 | 66.31 | 74.79 | 74.12 | 79.53 | — | |
| Bari et al.zero-shot protocol=cross-lingual2021.06 | 65.24 | — | 75.93 | 74.61 | 71.93 | |
| Jain et al.zero-shot protocol=cross-lingual2021.06 | 61.5 | — | 73.5 | 69.9 | 68.3 | |
| Ni et al.zero-shot protocol=cross-lingual2021.06 | 58.5 | — | 65.1 | 65.4 | 63 | |
| Xie et al.zero-shot protocol=cross-lingual2021.06 | 57.76 | — | 72.37 | 71.25 | 67.13 | |
| Mayhew et al.zero-shot protocol=cross-lingual2021.06 | 57.23 | — | 64.1 | 63.37 | 61.57 | |
| Xie et al.Mode=Zero-shot2019.11 | 56.9 | — | 72.4 | 71.3 | — | |
| Tsai et al.zero-shot protocol=cross-lingual2021.06 | 48.12 | — | 60.55 | 61.56 | 56.74 | |
| Täckström et al.zero-shot protocol=cross-lingual2021.06 | 40.4 | — | 59.3 | 58.4 | 52.7 |