Named Entity Recognition on CoNLL 2003 (test)
98.31F1 ScoreMAGNET
Evaluation Results
| Method | Links | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| MAGNETModel Category=Llama 2 models, Adaptation=MNTP, SSCL, and MSG2025.01 | 98.31 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| StructBERT-LargeModel Category=Encoder models, Evaluation Mode=Linear probing2025.01 | 97.31 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| LLM2Vec [MNTP]Model Category=Llama 2 models, Adaptation=MNTP only2025.01 | 97.16 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Llama-2-7BModel Category=Llama 2 models, Evaluation Mode=Linear probing2025.01 | 96.59 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| LLM2VecModel Category=Llama 2 models, Adaptation=MNTP and SimCSE2025.01 | 96.05 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DeBERTa-LargeModel Category=Encoder models, Evaluation Mode=Linear probing2025.01 | 94.97 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SOTA2024.09 | 94.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| VanillaBackbone=LUKE-LARGE2024.09 | 94.32 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| LUKE-CRBackbone=LUKE, Noise handling=CR2021.04 | 94.22 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SubRegWeigh (Random)Backbone=LUKE-LARGE, Processing time (hh:mm)=2:592024.09 | 94.22 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SubRegWeigh (K-means)Backbone=LUKE-LARGE, Processing time (hh:mm)=6:362024.09 | 94.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| CrossWeighBackbone=LUKE-LARGE, Processing time (hh:mm)=26:192024.09 | 94.12 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SubRegWeigh (Cos-Sim)Backbone=LUKE-LARGE, Processing time (hh:mm)=6:192024.09 | 94.11 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| LUKE-CrossWeighBackbone=LUKE, Noise handling=CrossWeigh2021.04 | 93.98 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| LUKEBackbone=LUKE, Noise handling=None2021.04 | 93.91 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SubRegWeigh (K-means)Backbone=RoBERTa-LARGE, Processing time (hh:mm)=5:212024.09 | 93.81 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| BSBackbone=RoBERTa-large2023.05 | 93.77 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DSpERTbest score=true2022.10 | 93.7 | 93.48 | 93.93 | — | — | — | — | — | — | — | — | — | — | — | |
| Weighted Sampling DistributionEncoder=BERT2021.08 | 93.68 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| XLNet-LargeModel Category=Encoder models, Evaluation Mode=Linear probing2025.01 | 93.67 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Baseline + BSboundary smoothing=true2022.04 | 93.65 | 93.61 | 93.68 | — | — | — | — | — | — | — | — | — | — | — | |
| Zhu and Liyear=20222022.10 | 93.65 | 93.61 | 93.68 | — | — | — | — | — | — | — | — | — | — | — | |
| DSpERTaverage of multiple independent runs=true2022.10 | 93.64 | 93.39 | 93.88 | — | — | — | — | — | — | — | — | — | — | — | |
| VanillaBackbone=RoBERTa-LARGE2024.09 | 93.54 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SubRegWeigh (Random)Backbone=RoBERTa-LARGE, Processing time (hh:mm)=3:262024.09 | 93.51 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| CNN Large + fine-tunebackbone=CNN Large, stacking_strategy=fine-tuning2019.03 | 93.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Baevski et al.POS tags=false, Contextual word embeddings=true, Designed for nested entities=false2019.09 | 93.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| BERT-Biaffine ModelEncoder=BERT2021.08 | 93.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Yu et al.trained with dev split=true2022.04 | 93.5 | 93.7 | 93.3 | — | — | — | — | — | — | — | — | — | — | — | |
| Yu et al.trained with train and dev splits=true, year=20202022.10 | 93.5 | 93.7 | 93.3 | — | — | — | — | — | — | — | — | — | — | — | |
| Biaffine2023.05 | 93.5 | 93.7 | 93.3 | — | — | — | — | — | — | — | — | — | — | — | |
| Baseline2022.04 | 93.48 | 92.93 | 94.03 | — | — | — | — | — | — | — | — | — | — | — | |
| GCDTBackbone=BERT_LARGE2019.06 | 93.47 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GCDT + BERT_LARGEadditional resources=BERT_LARGE2019.06 | 93.47 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Liu et al.POS tags=false, Contextual word embeddings=true2019.09 | 93.47 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Jiang et al.POS tags=true, Contextual word embeddings=true2019.09 | 93.47 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SubRegWeigh (Cos-Sim)Backbone=RoBERTa-LARGE, Processing time (hh:mm)=4:512024.09 | 93.44 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Pooled-FlairCrossWeigh=true2019.09 | 93.43 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Wang et al.trained with dev split=true2022.04 | 93.43 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Vanilla Negative SamplingEncoder=BERT2021.08 | 93.42 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| CrossWeighBackbone=RoBERTa-LARGE, Processing time (hh:mm)=30:552024.09 | 93.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Straková et al.trained with dev split=true2022.04 | 93.38 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| HCR w/ BERTEncoder=BERT2021.08 | 93.37 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| BINDER2022.08 | 93.33 | 93.08 | 93.57 | — | — | — | — | — | — | — | — | — | — | — | |
| BARTNER (Word)Backbone=BART-Large, Entity Representation=Word2021.06 | 93.24 | 92.61 | 93.87 | — | — | — | — | — | — | — | — | — | — | — | |
| Yan et al.trained with train and dev splits=true, year=20212022.10 | 93.24 | 92.61 | 93.87 | — | — | — | — | — | — | — | — | — | — | — | |
| BARTNER2023.05 | 93.24 | 92.61 | 93.87 | — | — | — | — | — | — | — | — | — | — | — | |
| CNN Large + ELMobackbone=CNN Large, stacking_strategy=ELMo-style2019.03 | 93.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| FlairCrossWeigh=true2019.09 | 93.19 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Akbik et al. (2019)Backbone=Flair2021.06 | 93.18 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Seq2Seqparadigm=generative, calibration=E-NER2023.05 | 93.15 | — | — | — | — | — | — | — | — | — | — | 0.0225 | — | — | |
| Pooled-FlairCrossWeigh=false2019.09 | 93.14 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Pooled Flairruns=Average of 52019.09 | 93.14 | — | — | 0.14 | — | — | — | — | — | — | — | — | — | — | |
| CSE2018.10 | 93.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Akbik et al.Training Data=training and development data2019.06 | 93.09 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Flair2019.07 | 93.09 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Akbik et al.POS tags=true, Contextual word embeddings=true, Backbone=FLAIR2019.09 | 93.09 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Flair EmbeddingEncoder=Flair2021.08 | 93.09 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| BSBackbone=BERT-large2023.05 | 93.08 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| PromptNERBackbone=RoBERTa-large2023.05 | 93.08 | 92.96 | 93.18 | — | — | — | — | — | — | — | — | — | — | — | |
| Straková et al. (2019)Backbone=BERT-Large2021.06 | 93.07 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| W2NERMethod category=Seq2Seq2021.12 | 93.07 | 92.71 | 93.44 | — | — | — | — | — | — | — | — | — | — | — | |
| Akbik et al.trained with dev split=true2022.04 | 93.07 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Li et al.average of multiple independent runs=true, year=20222022.10 | 93.07 | 92.71 | 93.44 | — | — | — | — | — | — | — | — | — | — | — | |
| W2NER2023.05 | 93.07 | 92.71 | 93.44 | — | — | — | — | — | — | — | — | — | — | — | |
| Yan et al.Method category=Seq2Seq, re-implementation=true2021.12 | 93.05 | 92.56 | 93.56 | — | — | — | — | — | — | — | — | — | — | — | |
| Seq2Seqparadigm=generative, calibration=softmax2023.05 | 93.05 | — | — | — | — | — | — | — | — | — | — | 0.0324 | — | — | |
| BERT-MRCEncoder=BERT2021.08 | 93.04 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Li et al.source=2020b2022.04 | 93.04 | 92.33 | 94.61 | — | — | — | — | — | — | — | — | — | — | — | |
| Li et al.year=2020b2022.10 | 93.04 | 92.33 | 94.61 | — | — | — | — | — | — | — | — | — | — | — | |
| MRC2023.05 | 93.04 | 92.33 | 94.61 | — | — | — | — | — | — | — | — | — | — | — | |
| UIE2023.05 | 92.99 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Straková et al.Method category=Seq2Seq2021.12 | 92.98 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| BARTNER (BPE)Backbone=BART-Large, Entity Representation=BPE2021.06 | 92.96 | 92.6 | 93.22 | — | — | — | — | — | — | — | — | — | — | — | |
| Shen et al.Method category=Span-based2021.12 | 92.94 | 92.13 | 93.73 | — | — | — | — | — | — | — | — | — | — | — | |
| BARTNER (Span)Backbone=BART-Large, Entity Representation=Span2021.06 | 92.88 | 92.31 | 93.45 | — | — | — | — | — | — | — | — | — | — | — | |
| FlairCrossWeigh=false2019.09 | 92.87 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Flairruns=Average of 52019.09 | 92.87 | — | — | 0.08 | — | — | — | — | — | — | — | — | — | — | |
| Li et al. (2020b)Backbone=BERT-Large, Note=Rerun of code2021.06 | 92.87 | 92.47 | 93.27 | — | — | — | — | — | — | — | — | — | — | — | |
| Shen et al.trained with train and dev splits=true, year=20222022.10 | 92.87 | 93.29 | 92.46 | — | — | — | — | — | — | — | — | — | — | — | |
| PIQN2023.05 | 92.87 | 93.29 | 92.46 | — | — | — | — | — | — | — | — | — | — | — | |
| Seq2Seqparadigm=generative, calibration=EDL2023.05 | 92.84 | — | — | — | — | — | — | — | — | — | — | 0.0322 | — | — | |
| BERTLARGE-CRBackbone=BERT-large, Noise handling=CR2021.04 | 92.82 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| BERT_LARGE2019.03 | 92.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| BERTLARGEEvaluation Approach=Fine-tuning2018.10 | 92.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| BERTBackbone=LARGE2019.06 | 92.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| BERT LargeModel size=Large2019.07 | 92.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Devlin et al.POS tags=true, Contextual word embeddings=true, Backbone=BERT2019.09 | 92.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Devlin et al.2022.04 | 92.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Devlin et al.year=20192022.10 | 92.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Clark et al.2019.06 | 92.61 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| CVT2018.10 | 92.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| CVT + MultiMulti-task=true2019.07 | 92.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Clark et al. (2018)Backbone=GloVe300d2021.06 | 92.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Template BART2022.08 | 92.55 | 91.72 | 93.4 | — | — | — | — | — | — | — | — | — | — | — | |
| Yu et al.Method category=Span-based, re-implementation=true2021.12 | 92.52 | 92.91 | 92.13 | — | — | — | — | — | — | — | — | — | — | — | |
| Yu et al. (2020)Backbone=BERT-Large, Note=Reproduction with sentence-level context2021.06 | 92.5 | 92.85 | 92.15 | — | — | — | — | — | — | — | — | — | — | — | |
| TablERTSentence=Multi2020.10 | 92.5 | 92 | 92.9 | — | — | — | — | — | — | — | — | — | — | — | |
| BERTLARGE-CrossWeighBackbone=BERT-large, Noise handling=CrossWeigh2021.04 | 92.49 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Joint Stack-LSTMjoint learning=true2019.07 | 92.43 | — | — | — | — | — | — | — | — | — | — | — | — | — |