Coreference Resolution on CoNLL English 2012 (test)
88MUC F1 ScoreWu et al. (2020)
Evaluation Results
| Method | Links | |||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Wu et al. (2020)2022.03 | 88 | 88.6 | 87.4 | 82.4 | 82 | 82.2 | — | — | — | — | 79.9 | 78.3 | 79.1 | 83.1 | — | |
| Wu et al. (2020)LM=SpanBERT, Decoder=QA, additional training data=true2022.11 | 88 | 88.6 | 87.4 | 82.4 | 82 | 82.2 | — | — | — | — | 79.9 | 78.3 | 79.1 | 83.1 | — | |
| Link-AppendLM=mT5, Decoder=transition2022.11 | 87.8 | 87.4 | 88.3 | 81.8 | 83.4 | 82.6 | — | — | — | — | 79.1 | 79.9 | 79.5 | 83.3 | — | |
| Dobrovolskii (2021)LM=ROBERTa, Decoder=c2f2022.11 | 86.3 | 84.9 | 87.9 | 77.4 | 82.6 | 79.9 | — | — | — | — | 76.1 | 77.1 | 76.6 | 81 | — | |
| CorefQA+SpanBERT2022.05 | 86.3 | — | — | — | — | 77.6 | — | — | — | — | — | — | 75.8 | 79.9 | — | |
| G2GTBackbone=SpanBERT-large, Handling Strategy=reduced2022.03 | 85.9 | 85.9 | 86 | 79.3 | 79.4 | 79.3 | — | — | — | — | 76.4 | 75.9 | 76.1 | 80.5 | — | |
| Kirstain et al. (2021)LM=LongFormer, Decoder=bilinear2022.11 | 85.8 | 86.5 | 85.1 | 80.3 | 77.9 | 79.1 | — | — | — | — | 76.8 | 75.4 | 76.1 | 80.3 | — | |
| SpanBERT + CMEncoder=SpanBERT-large, HOI approach=Cluster Merging2020.09 | 85.7 | 85.9 | 85.5 | 79 | 78.9 | 79 | — | — | — | — | 76.7 | 75.2 | 75.9 | 80.2 | 79.9 | |
| Xu and Choi (2020)2022.03 | 85.7 | 85.9 | 85.5 | 79 | 78.9 | 79 | — | — | — | — | 76.7 | 75.2 | 75.9 | 80.2 | — | |
| Xu and Choi (2020)LM=SpanBERT, Decoder=hoi2022.11 | 85.7 | 85.9 | 85.5 | 79 | 78.9 | 79 | — | — | — | — | 76.7 | 75.2 | 75.9 | 80.2 | — | |
| SpanBERTEncoder=SpanBERT-large2020.09 | 85.5 | 85.7 | 85.3 | 78.6 | 78.6 | 78.6 | — | — | — | — | 76.8 | 74.8 | 75.8 | 79.9 | 79.7 | |
| SpanBERT + AAEncoder=SpanBERT-large, HOI approach=Attend-to-Antecedent2020.09 | 85.4 | 86.1 | 84.8 | 79.3 | 77.3 | 78.3 | — | — | — | — | 76 | 74.7 | 75.4 | 79.7 | 79.4 | |
| SpanBERT + SCEncoder=SpanBERT-large, HOI approach=Span Clustering2020.09 | 85.4 | 85.5 | 85.2 | 78.4 | 78.5 | 78.4 | — | — | — | — | 76.5 | 74.1 | 75.2 | 79.7 | 79.2 | |
| SpanBERT2019.07 | 85.3 | 85.8 | 84.8 | 78.3 | 77.9 | 78.1 | — | — | — | — | 76.4 | 74.2 | 75.3 | 79.6 | — | |
| Joshi et al.Embedding Category=Fine-tuned on BERT, Reference=2019a2019.11 | 85.3 | 85.8 | 84.8 | 78.3 | 77.9 | 78.1 | — | — | — | — | 76.4 | 74.2 | 75.3 | 79.6 | — | |
| J-202020.09 | 85.3 | 85.8 | 84.8 | 78.3 | 77.9 | 78.1 | — | — | — | — | 76.4 | 74.2 | 75.3 | 79.6 | — | |
| Baseline + SpanBERT-largeBackbone=SpanBERT-large2022.03 | 85.3 | 85.8 | 84.8 | 78.3 | 77.9 | 78.1 | — | — | — | — | 76.4 | 74.2 | 75.3 | 79.6 | — | |
| G2GTBackbone=SpanBERT-large, Handling Strategy=overlap2022.03 | 85.3 | 85.8 | 84.9 | 78.7 | 78 | 78.3 | — | — | — | — | 76.4 | 74.5 | 75.4 | 79.7 | — | |
| Joshi et al. (2020)LM=SpanBERT, Decoder=c2f2022.11 | 85.3 | 85.8 | 84.8 | 78.3 | 77.9 | 78.1 | — | — | — | — | 76.4 | 74.2 | 75.3 | 79.6 | — | |
| Xia et al. (2020)LM=SpanBERT, Decoder=transitions2022.11 | 85.3 | 85.7 | 84.8 | 78.1 | 77.5 | 77.8 | — | — | — | — | 76.3 | 74.1 | 75.2 | 79.4 | — | |
| SpanBERT + EEEncoder=SpanBERT-large, HOI approach=Entity-Equalization2020.09 | 85.1 | 85.7 | 84.5 | 78.5 | 77.4 | 77.9 | — | — | — | — | 76.7 | 73.4 | 75 | 79.4 | 78.9 | |
| BERT-1seqImplementation=Authors' baseline replication, Objective=Single-sequence training2019.07 | 84.8 | 85.5 | 84.1 | 77.8 | 76.7 | 77.2 | — | — | — | — | 75.3 | 73.5 | 74.4 | 78.8 | — | |
| BERTImplementation=Authors' baseline replication2019.07 | 84.3 | 85.1 | 83.5 | 77.3 | 75.5 | 76.4 | — | — | — | — | 75 | 71.9 | 73.9 | 78.3 | — | |
| G2GTBackbone=BERT-large, Handling Strategy=reduced2022.03 | 83.9 | 84.7 | 83.1 | 76.8 | 74 | 75.4 | — | — | — | — | 75.3 | 70.1 | 72.6 | 77.3 | — | |
| BERTEncoder=BERT-base2020.09 | 83.8 | 85 | 82.5 | 77.3 | 74 | 75.6 | — | — | — | — | 74.9 | 70.7 | 72.8 | 77.4 | 77.3 | |
| Google BERT2019.07 | 83.7 | 84.9 | 82.5 | 76.7 | 74.2 | 75.4 | — | — | — | — | 74.6 | 70.1 | 72.3 | 77.1 | — | |
| BERT-large + c2f-corefvariant=independent2019.08 | 83.5 | 84.7 | 82.4 | 76.5 | 74 | 75.3 | — | — | — | — | 74.1 | 69.8 | 71.9 | 76.9 | — | |
| Joshi et al.Embedding Category=Fine-tuned on BERT, Reference=2019b2019.11 | 83.5 | 84.7 | 82.4 | 76.5 | 74 | 75.3 | — | — | — | — | 74.1 | 69.8 | 71.9 | 76.9 | — | |
| J-192020.09 | 83.5 | 84.7 | 82.4 | 76.5 | 74 | 75.3 | — | — | — | — | 74.1 | 69.8 | 71.9 | 76.9 | — | |
| Baseline + BERT-largeBackbone=BERT-large2022.03 | 83.5 | 84.7 | 82.4 | 76.5 | 74 | 75.3 | — | — | — | — | 74.1 | 69.8 | 71.9 | 76.9 | — | |
| Joshi et al. (2019)LM=BERT, Decoder=c2f2022.11 | 83.5 | 84.7 | 82.4 | 76.5 | 74 | 75.3 | — | — | — | — | 74.1 | 69.8 | 71.9 | 76.9 | — | |
| EEcitation=Kantor and Globerson (2019)2019.08 | 83.4 | 82.6 | 84.1 | 73.3 | 76.2 | 74.7 | — | — | — | — | 72.4 | 71.1 | 71.8 | 76.6 | — | |
| Kantor and GlobersonEmbedding Category=Pre-trained Contextual Embeddings2019.11 | 83.4 | 82.6 | 84.1 | 73.3 | 76.2 | 74.7 | — | — | — | — | 72.4 | 71.1 | 71.8 | 76.6 | — | |
| K-192020.09 | 83.4 | 82.6 | 84.1 | 73.3 | 76.2 | 74.7 | — | — | — | — | 72.4 | 71.1 | 71.8 | 76.6 | — | |
| G2GTBackbone=BERT-large, Handling Strategy=overlap2022.03 | 83.3 | 83.5 | 83.2 | 74.5 | 74.1 | 74.3 | — | — | — | — | 75.2 | 70.1 | 72.6 | 76.7 | — | |
| G2GTBackbone=BERT-base, Handling Strategy=reduced2022.03 | 83.2 | 83.4 | 83.1 | 70.1 | 73.7 | 71.9 | — | — | — | — | 72.1 | 70.1 | 71 | 75.4 | — | |
| Cluster Ranking ModelEmbedding Category=Pre-trained Contextual Embeddings2019.11 | 83 | 82.7 | 83.3 | 73.8 | 75.6 | 74.7 | — | — | — | — | 72.2 | 71 | 71.6 | 76.4 | — | |
| Yu et al. (2020)LM=BERT, Decoder=Ranking2022.11 | 83 | 82.7 | 83.3 | 73.8 | 75.6 | 74.7 | — | — | — | — | 72.2 | 71 | 71.6 | 76.4 | — | |
| BERT-large + c2f-corefvariant=overlap2019.08 | 82.8 | 85.1 | 80.5 | 77.5 | 70.9 | 74.1 | — | — | — | — | 73.8 | 69.3 | 71.5 | 76.1 | — | |
| G2GTBackbone=BERT-base, Handling Strategy=overlap2022.03 | 82 | 81.2 | 82.8 | 69.8 | 73.6 | 71.6 | — | — | — | — | 69.6 | 69.3 | 69.4 | 74.4 | — | |
| Fei et al.2019.08 | 81.4 | 85.4 | 77.9 | 77.9 | 66.4 | 71.7 | — | — | — | — | 70.6 | 66.3 | 68.4 | 73.8 | — | |
| BERT-base + c2f-corefvariant=overlap2019.08 | 81.4 | 80.4 | 82.3 | 69.6 | 73.8 | 71.7 | — | — | — | — | 69 | 68.5 | 68.8 | 73.9 | — | |
| F-192020.09 | 81.4 | 85.4 | 77.9 | 77.9 | 66.4 | 71.7 | — | — | — | — | 70.6 | 66.3 | 68.4 | 73.8 | — | |
| Fei et al. (2019)2022.03 | 81.4 | 85.4 | 77.9 | 77.9 | 66.4 | 71.7 | — | — | — | — | 70.6 | 66.3 | 68.4 | 73.8 | — | |
| Baseline + BERT-baseBackbone=BERT-base2022.03 | 81.4 | 80.4 | 82.3 | 69.6 | 73.8 | 71.7 | — | — | — | — | 69 | 68.5 | 68.8 | 73.9 | — | |
| BERT+c2f-coref2022.05 | 81.4 | — | — | — | — | 71.7 | — | — | — | — | — | — | 68.8 | 73.9 | — | |
| BERT-base + c2f-corefvariant=independent2019.08 | 81.3 | 80.2 | 82.4 | 69.6 | 73.8 | 71.6 | — | — | — | — | 69 | 68.6 | 68.8 | 73.9 | — | |
| TANL2022.05 | 81 | — | — | — | — | 69 | — | — | — | — | — | — | 68.4 | 72.8 | — | |
| Second-order inference (Full Approach)ELMo=true, hyperparameter tuning=true, coarse-to-fine=true, second-order inference=true2018.04 | 80.4 | 81.4 | 79.5 | 72.2 | 69.5 | 70.8 | — | — | — | — | 68.2 | 67.1 | 67.6 | 73 | — | |
| c2f-corefcitation=Lee et al. (2018)2019.08 | 80.4 | 81.4 | 79.5 | 72.2 | 69.5 | 70.8 | — | — | — | — | 68.2 | 67.1 | 67.6 | 73 | — | |
| Lee et al.Note=Previous SOTA2019.07 | 80.4 | 81.4 | 79.5 | 72.2 | 69.5 | 70.8 | — | — | — | — | 68.2 | 67.1 | 67.6 | 73 | — | |
| Lee et al.Embedding Category=Pre-trained Contextual Embeddings2019.11 | 80.4 | 81.4 | 79.5 | 72.2 | 69.5 | 70.8 | — | — | — | — | 68.2 | 67.1 | 67.6 | 73 | — | |
| L-182020.09 | 80.4 | 81.4 | 79.5 | 72.2 | 69.5 | 70.8 | — | — | — | — | 68.2 | 67.1 | 67.6 | 73 | — | |
| BaselineBackbone=ELMo2022.03 | 80.4 | 81.4 | 79.5 | 72.2 | 69.5 | 70.8 | — | — | — | — | 68.2 | 67.1 | 67.6 | 73 | — | |
| Lee et al. (2018)LM=Elmo, Decoder=c2f2022.11 | 80.4 | 81.4 | 79.5 | 72.2 | 69.5 | 70.8 | — | — | — | — | 68.2 | 67.1 | 67.6 | 73 | — | |
| Higher-order c2f-coref2022.05 | 80.4 | — | — | — | — | 70.8 | — | — | — | — | — | — | 67.6 | 73 | — | |
| Coarse-to-fine inferenceELMo=true, hyperparameter tuning=true, coarse-to-fine=true2018.04 | 80.1 | 80.4 | 79.9 | 71 | 70 | 70.5 | — | — | — | — | 67.5 | 67.2 | 67.3 | 72.6 | — | |
| Lee et al. (2017) + ELMo + Hyperparameter TuningELMo=true, hyperparameter tuning=true2018.04 | 79.8 | 80.7 | 78.8 | 71.7 | 68.7 | 70.2 | — | — | — | — | 67.2 | 66.8 | 67 | 72.3 | — | |
| G2GTBackbone=BERT-large, Handling Strategy=truncated2022.03 | 79.6 | 80.1 | 79.2 | 71.3 | 71 | 71.1 | — | — | — | — | 69.1 | 68.8 | 68.9 | 73.2 | — | |
| TANL (multitask)training=multi-task2022.05 | 78.7 | — | — | — | — | 65.7 | — | — | — | — | — | — | 63.8 | 69.4 | — | |
| Lee et al. (2017) + ELMoELMo=true2018.04 | 78.6 | 80.1 | 77.2 | 69.8 | 66.5 | 68.1 | — | — | — | — | 66.4 | 62.9 | 64.6 | 70.4 | — | |
| G2GTBackbone=BERT-base, Handling Strategy=truncated2022.03 | 78.1 | 78.4 | 77.9 | 69.6 | 71 | 70.3 | — | — | — | — | 66.8 | 67.3 | 67 | 71.8 | — | |
| End-to-end neural coreference resolution modelModel configuration=ensemble2017.07 | 77.2 | 81.2 | 73.6 | 72.3 | 61.7 | 66.6 | — | — | — | — | 65.2 | 60.2 | 62.6 | 68.8 | — | |
| Zhang et al.Embedding Category=Context Independent2019.11 | 76.5 | 79.4 | 73.8 | 69 | 62.3 | 65.5 | — | — | — | — | 64.9 | 58.3 | 61.4 | 67.8 | — | |
| Lee et al. (2017)2018.04 | 75.8 | 78.4 | 73.4 | 68.6 | 61.8 | 65 | — | — | — | — | 62.7 | 59 | 60.8 | 67.2 | — | |
| End-to-end neural coreference resolution modelModel configuration=single2017.07 | 75.8 | 78.4 | 73.4 | 68.6 | 61.8 | 65 | — | — | — | — | 62.7 | 59 | 60.8 | 67.2 | — | |
| e2e-corefcitation=Lee et al. (2017)2019.08 | 75.8 | 78.4 | 73.4 | 68.6 | 61.8 | 65 | — | — | — | — | 62.7 | 59 | 60.8 | 67.2 | — | |
| Lee et al.Embedding Category=Context Independent2019.11 | 75.8 | 78.4 | 73.4 | 68.6 | 61.8 | 65 | — | — | — | — | 62.7 | 59 | 60.8 | 67.2 | — | |
| L-172020.09 | 75.8 | 78.4 | 73.4 | 68.6 | 61.8 | 65 | — | — | — | — | 62.7 | 59 | 60.8 | 67.2 | — | |
| Lee et al. (2017)2022.03 | 75.8 | 78.4 | 73.4 | 68.6 | 61.8 | 65 | — | — | — | — | 62.7 | 59 | 60.8 | 67.2 | — | |
| Lee et al. (2017)Decoder=neural e2e2022.11 | 75.8 | 78.4 | 73.4 | 68.6 | 61.8 | 65 | — | — | — | — | 62.7 | 59 | 60.8 | 67.2 | — | |
| DEEPSTRUCTtraining=w/ finetune2022.05 | 74.9 | — | — | — | — | 71.3 | — | — | — | — | — | — | 73.1 | 73.1 | — | |
| Heuristic Loss2016.09 | 74.65 | 79.63 | 70.25 | 69.21 | 57.87 | 63.03 | — | — | — | — | 63.62 | 53.97 | 58.4 | 65.36 | — | |
| Clark and Manning (2016a)2018.04 | 74.6 | 79.2 | 70.4 | 69.9 | 58 | 63.4 | — | — | — | — | 63.5 | 55.5 | 59.2 | 65.7 | — | |
| Clark and ManningVersion=2016a2017.07 | 74.6 | 79.2 | 70.4 | 69.9 | 58 | 63.4 | — | — | — | — | 63.5 | 55.5 | 59.2 | 65.7 | — | |
| Clark and Manningyear=20162019.08 | 74.6 | 79.2 | 70.4 | 69.9 | 58 | 63.4 | — | — | — | — | 63.5 | 55.5 | 59.2 | 65.7 | — | |
| Clark and ManningEmbedding Category=Context Independent2019.11 | 74.6 | 79.2 | 70.4 | 69.9 | 58 | 63.4 | — | — | — | — | 63.5 | 55.5 | 59.2 | 65.7 | — | |
| Clark and Manning (2016)2022.03 | 74.6 | 79.2 | 70.4 | 69.9 | 58 | 63.4 | — | — | — | — | 63.5 | 55.5 | 59.2 | 65.7 | — | |
| Reward Rescaling2016.09 | 74.56 | 79.19 | 70.44 | 69.93 | 57.99 | 63.4 | — | — | — | — | 63.46 | 55.52 | 59.23 | 65.73 | — | |
| REINFORCE2016.09 | 74.48 | 80.08 | 69.61 | 70.7 | 56.96 | 63.09 | — | — | — | — | 63.59 | 54.46 | 58.67 | 65.41 | — | |
| Clark & Manning (2016)2016.09 | 74.23 | 79.91 | 69.3 | 71.01 | 56.53 | 62.95 | — | — | — | — | 63.84 | 54.33 | 58.7 | 65.29 | — | |
| Clark and Manning (2016b)2018.04 | 74.2 | 79.9 | 69.3 | 71 | 56.5 | 63 | — | — | — | — | 63.8 | 54.3 | 58.7 | 65.3 | — | |
| Clark and ManningVersion=2016b2017.07 | 74.2 | 79.9 | 69.3 | 71 | 56.5 | 63 | — | — | — | — | 63.8 | 54.3 | 58.7 | 65.3 | — | |
| NN Cluster Ranker2016.06 | 74.06 | 78.93 | 69.75 | 70.08 | 56.98 | 62.86 | 62.48 | 55.82 | 58.96 | — | — | — | — | 65.29 | — | |
| NN Mention Ranker2016.06 | 74.05 | 79.77 | 69.1 | 69.68 | 56.37 | 62.32 | 63.02 | 53.59 | 57.92 | — | — | — | — | 64.76 | — | |
| RNN-based Coreference Resolutionmodel=This work2016.04 | 73.42 | 77.49 | 69.75 | 66.83 | 56.95 | 61.5 | 62.14 | 53.85 | 57.7 | 64.21 | — | — | — | — | — | |
| Wiseman et al. (2016)2016.06 | 73.42 | 77.49 | 69.75 | 66.83 | 56.95 | 61.5 | 62.14 | 53.85 | 57.7 | — | — | — | — | 64.21 | — | |
| Wiseman et al. (2016)2016.09 | 73.42 | 77.49 | 69.75 | 66.83 | 56.95 | 61.5 | — | — | — | — | 62.14 | 53.85 | 57.7 | 64.21 | — | |
| Wiseman et al. (2016)2018.04 | 73.4 | 77.5 | 69.8 | 66.8 | 57 | 61.5 | — | — | — | — | 62.1 | 53.9 | 57.7 | 64.2 | — | |
| Wiseman et al.Version=20162017.07 | 73.4 | 77.5 | 69.8 | 66.8 | 57 | 61.5 | — | — | — | — | 62.1 | 53.9 | 57.7 | 64.2 | — | |
| Wiseman et al.year=20162019.08 | 73.4 | 77.5 | 69.8 | 66.8 | 57 | 61.5 | — | — | — | — | 62.1 | 53.9 | 57.7 | 64.2 | — | |
| Wiseman et al. (2016)2022.03 | 73.4 | 77.5 | 69.8 | 66.8 | 57 | 61.5 | — | — | — | — | 62.1 | 53.9 | 57.7 | 64.2 | — | |
| Wiseman et al.year=20152016.04 | 72.6 | 76.23 | 69.31 | 66.07 | 55.83 | 60.52 | 59.41 | 54.88 | 57.05 | 63.39 | — | — | — | — | — | |
| Clark and Manning (2015)2018.04 | 72.6 | 76.1 | 69.4 | 65.6 | 56 | 60.4 | — | — | — | — | 59.4 | 53 | 56 | 63 | — | |
| Wiseman et al. (2015)2018.04 | 72.6 | 76.2 | 69.3 | 66.2 | 55.8 | 60.5 | — | — | — | — | 59.4 | 54.9 | 57.1 | 63.4 | — | |
| Wiseman et al. (2015)2016.06 | 72.6 | 76.23 | 69.31 | 66.07 | 55.83 | 60.52 | 59.41 | 54.88 | 57.05 | — | — | — | — | 63.39 | — | |
| Wiseman et al.Version=20152017.07 | 72.6 | 76.2 | 69.3 | 66.2 | 55.8 | 60.5 | — | — | — | — | 59.4 | 54.9 | 57.1 | 63.4 | — | |
| Clark and ManningVersion=20152017.07 | 72.6 | 76.1 | 69.4 | 65.6 | 56 | 60.4 | — | — | — | — | 59.4 | 53 | 56 | 63 | — | |
| Clark and Manningyear=20152019.08 | 72.6 | 76.1 | 69.4 | 65.6 | 56 | 60.4 | — | — | — | — | 59.4 | 53 | 56 | 63 | — | |
| Wiseman et al.year=20152019.08 | 72.6 | 76.2 | 69.3 | 66.2 | 55.8 | 60.5 | — | — | — | — | 59.4 | 54.9 | 57.1 | 63.1 | — |