Named Entity Recognition on OntoNotes 5.0
92.1F1 ScoreLS-unLLaMA-2-7B
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| LS-unLLaMA-2-7BModel Variant=LS-unLLaMA, Model Scale=7B2023.10 | 92.1 | — | — | — | |
| RoBERTa-LargeSetting=Discriminant baselines2023.10 | 91.72 | — | — | — | |
| SPANNERvariant=both (sp2)2021.06 | 91.59 | — | — | — | |
| RoBERTa-BaseSetting=Discriminant baselines2023.10 | 91.55 | — | — | — | |
| SpanDec (RoBERTa-L)Params=0.35B, Throughput=118.7 (1×)2026.04 | 91.3 | — | — | — | |
| LS-unLLaMA-2-13BModel Variant=LS-unLLaMA, Model Scale=13B2023.10 | 91.07 | — | — | — | |
| SplitNERParams=0.35B, Throughput=104.1 (0.88×)2026.04 | 90.9 | — | — | — | |
| SPANNERvariant=generic (sp1)2021.06 | 90.84 | — | — | — | |
| sq3Char/Sub.=BERT, Word=none, Sent.=LSTM2021.06 | 90.77 | — | — | — | |
| sq5Char/Sub.=ELMo, Word=none, Sent.=LSTM2021.06 | 90.44 | — | — | — | |
| AESINER2020.10 | 90.32 | — | — | — | |
| Luo et al. (2020)text_encoder=BERT2020.10 | 90.3 | — | — | — | |
| sq1Char/Sub.=Flair, Word=none, Sent.=LSTM2021.06 | 90.23 | — | — | — | |
| InstructUIESetting=Instruction-tuning2023.10 | 90.19 | — | — | — | |
| sq0Char/Sub.=Flair, Word=GloVe, Sent.=LSTM2021.06 | 90.11 | — | — | — | |
| sq4Char/Sub.=ELMo, Word=rand, Sent.=LSTM2021.06 | 90.1 | — | — | — | |
| sq2Char/Sub.=BERT, Word=GloVe, Sent.=LSTM2021.06 | 90.08 | — | — | — | |
| Liu et al. (2019b)2020.10 | 89.94 | — | — | — | |
| UniNERParams=7B, Throughput=< 6.4 (0.05×)2026.04 | 89.9 | — | — | — | |
| Jie and Lu (2019)2020.10 | 89.88 | — | — | — | |
| Dai et al. (2019)2020.10 | 89.83 | — | — | — | |
| Yan et al. (2019)2020.10 | 89.78 | — | — | — | |
| RANSetting=Discriminant baselines2023.10 | 89.38 | — | — | — | |
| Akbik et al. (2018)2020.10 | 89.3 | — | — | — | |
| BERT-LargeSetting=Discriminant baselines2023.10 | 89.27 | — | — | — | |
| Devlin et al. (2019)text_encoder=BERT2020.10 | 89.16 | — | — | — | |
| Dang et al. (2018)2020.10 | 88.91 | — | — | — | |
| BERT-BaseSetting=Discriminant baselines2023.10 | 88.88 | — | — | — | |
| Luo et al. (2018)result_source=Run by authors2020.10 | 88.79 | — | — | — | |
| InstructUIEParams=13B, Throughput=< 5.8 (0.05×)2026.04 | 88.6 | — | — | — | |
| sq6Char/Sub.=cnn, Word=glove, Sent.=LSTM2021.06 | 88.31 | — | — | — | |
| GLiNERParams=0.35B, Throughput=76.2 (0.64×)2026.04 | 88.1 | — | — | — | |
| GRN2019.07 | 87.67 | 87.79 | 87.56 | — | |
| CNN-BiLSTM-Att-CRFArchitecture=CNN-BiLSTM with Attention and CRF2019.07 | 87.25 | — | — | — | |
| sq7Char/Sub.=cnn, Word=none, Sent.=LSTM2021.06 | 86.87 | — | — | — | |
| ID-CNN2019.07 | 86.84 | — | — | — | |
| Shen et al. (2017)2019.07 | 86.63 | — | — | — | |
| Chiu and Nichols (2016)2019.07 | 86.28 | 86.04 | 86.53 | — | |
| Chiu and Nichols (2016)2020.10 | 86.12 | — | — | — | |
| sq8Char/Sub.=none, Word=glove, Sent.=LSTM2021.06 | 86.03 | — | — | — | |
| sq9Char/Sub.=none, Word=rand, Sent.=LSTM2021.06 | 85.73 | — | — | — | |
| Durrett and Klein (2014)2019.07 | 84.04 | 85.22 | 82.89 | — | |
| SPANNERvariant=both (sp2)2021.06 | 83.22 | — | — | — | |
| Passos, Kumar, and McCallum (2014)2019.07 | 82.3 | — | — | — | |
| SPANNERvariant=generic (sp1)2021.06 | 82.24 | — | — | — | |
| GPT-NERParams=175B, Throughput=< 5 (0.04×)2026.04 | 82.2 | — | — | — | |
| sq2Char/Sub.=BERT, Word=GloVe, Sent.=LSTM2021.06 | 81.55 | — | — | — | |
| sq3Char/Sub.=BERT, Word=none, Sent.=LSTM2021.06 | 80.11 | — | — | — | |
| sq1Char/Sub.=Flair, Word=none, Sent.=LSTM2021.06 | 79.55 | — | — | — | |
| sq5Char/Sub.=ELMo, Word=none, Sent.=LSTM2021.06 | 79.32 | — | — | — | |
| sq4Char/Sub.=ELMo, Word=rand, Sent.=LSTM2021.06 | 78.28 | — | — | — | |
| sq0Char/Sub.=Flair, Word=GloVe, Sent.=LSTM2021.06 | 78.17 | — | — | — | |
| LS-LLaMA-2-13BModel Variant=LS-LLaMA, Model Scale=13B2023.10 | 77.73 | — | — | — | |
| LS-LLaMA-2-7BModel Variant=LS-LLaMA, Model Scale=7B2023.10 | 77.41 | — | — | — | |
| sq6Char/Sub.=cnn, Word=glove, Sent.=LSTM2021.06 | 75.1 | — | — | — | |
| sq7Char/Sub.=cnn, Word=none, Sent.=LSTM2021.06 | 74.63 | — | — | — | |
| sq2Char/Sub.=BERT, Word=GloVe, Sent.=LSTM2021.06 | 71.07 | — | — | — | |
| MUCOShot=5-shot2021.06 | 71.06 | 73.27 | 69 | — | |
| sq3Char/Sub.=BERT, Word=none, Sent.=LSTM2021.06 | 71.01 | — | — | — | |
| SPANNERvariant=both (sp2)2021.06 | 69.91 | — | — | — | |
| sq8Char/Sub.=none, Word=glove, Sent.=LSTM2021.06 | 69.81 | — | — | — | |
| MAMLShot=5-shot2021.06 | 67.57 | 65.99 | 69.31 | — | |
| SPANNERvariant=generic (sp1)2021.06 | 66.67 | — | — | — | |
| sq0Char/Sub.=Flair, Word=GloVe, Sent.=LSTM2021.06 | 66.57 | — | — | — | |
| sq9Char/Sub.=none, Word=rand, Sent.=LSTM2021.06 | 66.42 | — | — | — | |
| WPNShot=5-shot2021.06 | 66.34 | 65.28 | 67.66 | — | |
| sq1Char/Sub.=Flair, Word=none, Sent.=LSTM2021.06 | 65.58 | — | — | — | |
| sq5Char/Sub.=ELMo, Word=none, Sent.=LSTM2021.06 | 65.57 | — | — | — | |
| SOLARParams=10.7B, Throughput=< 8 (0.07×)2026.04 | 64.84 | — | — | — | |
| sq4Char/Sub.=ELMo, Word=rand, Sent.=LSTM2021.06 | 64.62 | — | — | — | |
| sq6Char/Sub.=cnn, Word=glove, Sent.=LSTM2021.06 | 64.36 | — | — | — | |
| PNShot=5-shot2021.06 | 60.12 | 61.84 | 58.61 | — | |
| Mistral-12BParams=12B, Throughput=< 7 (0.06×)2026.04 | 59.3 | — | — | — | |
| BERTShot=5-shot2021.06 | 59.04 | 61.81 | 56.64 | — | |
| MUCOshot=1-shot2021.06 | 57.89 | 60.43 | 55.82 | — | |
| WPNshot=1-shot2021.06 | 56.2 | 58.29 | 54.39 | — | |
| sq7Char/Sub.=cnn, Word=none, Sent.=LSTM2021.06 | 56.16 | — | — | — | |
| MAMLshot=1-shot2021.06 | 56.15 | 56.63 | 55.84 | — | |
| LLaMA2-7BParams=7B, Throughput=< 9 (0.08×)2026.04 | 52.78 | — | — | — | |
| LTCShot=5-shot2021.06 | 52.35 | 62.06 | 46.08 | — | |
| sq8Char/Sub.=none, Word=glove, Sent.=LSTM2021.06 | 51.83 | — | — | — | |
| ChatGPTSetting=Zero- and few-shot2023.10 | 51.1 | — | — | — | |
| LTCshot=1-shot2021.06 | 50.04 | 60.83 | 43.25 | — | |
| sq9Char/Sub.=none, Word=rand, Sent.=LSTM2021.06 | 46.84 | — | — | — | |
| BERTshot=1-shot2021.06 | 39.92 | 54.92 | 32.09 | — | |
| PNshot=1-shot2021.06 | 38.67 | 55.77 | 30.56 | — | |
| LLaMA2-13BParams=13B, Throughput=< 7 (0.06×)2026.04 | 38.32 | — | — | — | |
| LLaMA3-8BParams=8B, Throughput=< 9 (0.08×)2026.04 | 22.34 | — | — | — | |
| GPT-3.5-TurboSetting=Zero- and few-shot2023.10 | 18.22 | — | — | — | |
| LLaMA-2-7B (zero-shot)Setting=Zero-shot, Model Scale=7B2023.10 | 1.2 | — | — | — | |
| Single(QA)Setup=Standard single-model QA-based2023.10 | — | — | — | 89.02 | |
| Single(SeqTag)Setup=Standard single-model sequence tagging2023.10 | — | — | — | 88.64 | |
| SplitNER(QA-QA)2023.10 | — | — | — | 90.86 | |
| SplitNER(QANoCharPattern-QA)Features=No additional character and pattern features2023.10 | — | — | — | 90.58 | |
| SplitNER(SeqTag-QA)Span Detection=Sequence tagging2023.10 | — | — | — | 90.3 |