Joint Entity and Relation Extraction on CONLL04
93.26Entity F1multi-head EC
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| multi-head ECPre-calculated Features=false, Evaluation Protocol=relaxed2018.04 | 93.26 | 93.41 | 93.15 | 72.99 | 63.37 | 67.01 | 80.14 | — | — | |
| baseline ECExternal NLP Tools=false, Evaluation Protocol=Relaxed2018.08 | 93.26 | — | — | — | — | 67.01 | 80.14 | — | — | |
| baseline EC + ATExternal NLP Tools=false, Evaluation Protocol=Relaxed2018.08 | 93.04 | — | — | — | — | 67.99 | 80.51 | — | — | |
| Gupta et al. (2016)Pre-calculated Features=true, Evaluation Protocol=relaxed2018.04 | 92.4 | 92.5 | 92.1 | 78.5 | 63 | 69.9 | 81.15 | — | — | |
| Gupta et al. (2016)External NLP Tools=true, Evaluation Protocol=Relaxed2018.08 | 92.4 | — | — | — | — | 69.9 | 81.15 | — | — | |
| DEEPSTRUCTEvaluation protocol=w/ finetune2022.05 | 90.7 | — | — | — | — | 78.3 | — | — | — | |
| DEEPSTRUCTmode=w/ finetune2022.05 | 90.7 | — | — | — | — | 78.3 | — | — | — | |
| DEEPSTRUCTprotocol=w/ finetune2022.05 | 90.7 | — | — | — | — | 78.3 | — | — | — | |
| TANLTraining Setting=multi-task2021.01 | 90.3 | — | — | — | — | 70 | — | — | — | |
| TANL (multitask)mode=multitask2022.05 | 90.3 | — | — | — | — | 70 | — | — | — | |
| TANL2022.05 | 90.3 | — | — | — | — | 71.4 | — | — | — | |
| TANLTraining Setting=multi-dataset2021.01 | 89.8 | — | — | — | — | 72.6 | — | — | — | |
| TANLTraining Setting=standard2021.01 | 89.4 | — | — | — | — | 71.4 | — | — | — | |
| TANL2022.05 | 89.4 | — | — | — | — | 71.4 | — | — | — | |
| SpERTTraining Setting=standard2021.01 | 88.9 | — | — | — | — | 71.5 | — | — | — | |
| RSANTraining Setting=standard2021.01 | 88.9 | — | — | — | — | 71.9 | — | — | — | |
| SpERT2022.05 | 88.9 | — | — | — | — | 71.5 | — | — | — | |
| MRC4ERE2022.05 | 88.9 | — | — | — | — | 71.9 | — | — | — | |
| Zhao et al., 20202022.05 | 88.9 | — | — | — | — | 71.9 | — | — | — | |
| Gupta et al. (2016)Pre-calculated Features=false, Evaluation Protocol=relaxed2018.04 | 88.8 | 88.5 | 88.9 | 64.6 | 53.1 | 58.3 | 73.6 | — | — | |
| Gupta et al. (2016)External NLP Tools=false, Evaluation Protocol=Relaxed2018.08 | 88.8 | — | — | — | — | 58.3 | 73.6 | — | — | |
| DEEPSTRUCTEvaluation protocol=multi-task2022.05 | 88.4 | — | — | — | — | 72.8 | — | — | — | |
| DEEPSTRUCTmode=multi-task2022.05 | 88.4 | — | — | — | — | 72.8 | — | — | — | |
| DEEPSTRUCTprotocol=multi-task2022.05 | 88.4 | — | — | — | — | 72.8 | — | — | — | |
| multi-headPre-calculated Features=false, Evaluation Protocol=strict2018.04 | 83.9 | 83.75 | 84.06 | 63.75 | 60.43 | 62.04 | 72.97 | — | — | |
| Adel & Schütze (2017)Pre-calculated Features=false, Evaluation Protocol=relaxed2018.04 | 82.1 | — | — | — | — | 62.5 | 72.3 | — | — | |
| Adel and Schütze (2017)External NLP Tools=false, Evaluation Protocol=Relaxed2018.08 | 82.1 | — | — | — | — | 62.5 | 72.3 | — | — | |
| Miwa & Sasaki (2014)Pre-calculated Features=false, Evaluation Protocol=strict2018.04 | 80.7 | 81.2 | 80.2 | 76 | 50.9 | 61 | 70.85 | — | — | |
| SMADE-IEBackbone=gemini-3-flash-preview2026.06 | 62.09 | — | — | — | — | 43.69 | — | — | — | |
| CROSSAGENTIEBackbone=gemini-3-flash-preview2026.06 | 54.61 | — | — | — | — | 38.85 | — | — | — | |
| DEEPSTRUCTEvaluation protocol=zero-shot2022.05 | 48.3 | — | — | — | — | 25.8 | — | — | — | |
| DEEPSTRUCTmode=zero-shot2022.05 | 48.3 | — | — | — | — | 25.8 | — | — | — | |
| DEEPSTRUCTprotocol=zero-shot2022.05 | 48.3 | — | — | — | — | 25.8 | — | — | — | |
| GPT-3 175Bmode=zero-shot2022.05 | 34.7 | — | — | — | — | 18.1 | — | — | — | |
| GPT-3 175Bprotocol=zero-shot2022.05 | 34.7 | — | — | — | — | 18.1 | — | — | — | |
| CROSSAGENTIEBackbone=GPT-3.5-Turbo-0125, Prompting Strategy=multi-agent debate, Evaluation Protocol=zero-shot2026.06 | — | — | — | — | — | — | — | 42.5 | 29.05 | |
| SMADE-IEBackbone=GPT-3.5-Turbo-0125, Prompting Strategy=multi-agent debate, Evaluation Protocol=zero-shot2026.06 | — | — | — | — | — | — | — | 58.44 | 44.74 |