Conll
Benchmarks
Task NameDataset NameSOTA ResultTrendResults
CoNLL shared-task 2009 (test)
85.4Precision
4
CoNLL 2009 (test)
96.88Accuracy
4
CoNLL03 (train)
4.86Latency (s)
4
CoNLL 2003
96.4E2E Coverage
3
CoNLL downsampled 2003 (test)
76.96F1 (ORG)
3
CoNLL Arabic 2012 (test)
73.6MUC Precision
3
CoNLL German 2006 (test)
91.7F1 Score
3
CoNLL German 2003
86.4F1 Score
3
CoNLL Uncased (test)
87.9F1 Score
3
CoNLL 7 Small treebanks 2018 UD v2.2 (test)
87.64UPOS Accuracy
3
CoNLL 2018 Shared Task 61 Big treebanks UD v2.2 (test)
95.63UPOS
3
CoNLL 2009 (Brown)
83.5F1 Score
3
CoNLL Brown 2008
74.2F1 Score
3
CoNLL Shared Task 2017
96.41CS CAC Score
3
CoNLL 2003
94.8Recall
2
CoNLL German 2003 revised
90.3F1 Score
2
CoNLL Shared Task Big treebanks 2018 (test)
99.51Token Accuracy
2
CoNLL WSJ 2005 (test)
98.9Precision
2
CoNLL 2005 (dev)
91.92Constituents
2
CoNLL (test2)
45.1F0.5
2
CoNLL (test1)
36.1F0.5 Score
2
CoNLL 2003 (Feasibility study)
94F1 Score
1
CoNLL 2012 (test)
99.8Precision
1
CoNLL Reduced (test)
—Primary metric
0