Constituency Parsing on Penn Treebank WSJ (section 23 test)
95.8F1 ScoreBerkeley Neural
Evaluation Results
| Method | Links | |
|---|---|---|
| Berkeley NeuralTime (ms/sent.)=42 ± 3, Constraint scheduling policy=Berkeley Neural Parser2025.12 | 95.8 | |
| MetaJuLSTime (ms/sent.)=24 ± 2, Constraint scheduling policy=meta-learned MetaJuLS policy2025.12 | 95.6 | |
| A* NeuralTime (ms/sent.)=45 ± 3, Constraint scheduling policy=A* search with neural heuristics2025.12 | 95.2 | |
| Supervised RankingTime (ms/sent.)=28 ± 2, Constraint scheduling policy=supervised ranking2025.12 | 94.6 | |
| Activity-BasedTime (ms/sent.)=32 ± 2, Constraint scheduling policy=activity-based scheduling2025.12 | 94.5 | |
| VSIDS-styleTime (ms/sent.)=30 ± 2, Constraint scheduling policy=VSIDS-style conflict-driven scheduling2025.12 | 94.3 | |
| In-order parserSupervision Setting=semi-supervised, Inference Setting=reranking2017.07 | 94.2 | |
| Cost-Normalized GreedyTime (ms/sent.)=29 ± 2, Constraint scheduling policy=cost-normalized greedy2025.12 | 94.2 | |
| GPU CKYTime (ms/sent.)=38 ± 2, Constraint scheduling policy=GPU-parallelized CKY2025.12 | 94.1 | |
| Choe and CharniakSupervision Setting=semi-supervised, Inference Setting=reranking2017.07 | 93.8 | |
| Kuncoro et al.Supervision Setting=fully-supervised, Inference Setting=reranking2017.07 | 93.6 | |
| In-order parserSupervision Setting=fully-supervised, Inference Setting=reranking2017.07 | 93.6 | |
| Dyer et al.Supervision Setting=fully-supervised, Inference Setting=reranking2017.07 | 93.3 | |
| Top-down parserSupervision Setting=fully-supervised, Inference Setting=reranking2017.07 | 93.3 | |
| Bottom-up parserSupervision Setting=fully-supervised, Inference Setting=reranking2017.07 | 93.3 | |
| LSTM+A ensembleTraining Set=high-confidence corpus2014.12 | 92.8 | |
| Choe and CharniakSupervision Setting=fully-supervised, Inference Setting=reranking2017.07 | 92.6 | |
| LSTM+ATraining Set=high-confidence corpus2014.12 | 92.5 | |
| Huang & Harper (2010) ensembleTraining Set=semi-supervised2014.12 | 92.4 | |
| Shindo et al. (2012)type=Generative (G), mode=ensemble2016.02 | 92.4 | |
| RNNGtype=Generative (G), formulation=p(y | x)2016.02 | 92.4 | |
| McClosky et al. (2006)Training Set=semi-supervised2014.12 | 92.1 | |
| McClosky et al. (2006)type=Semisupervised (S)2016.02 | 92.1 | |
| Vinyals et al. (2015)type=Semisupervised (S), mode=single2016.02 | 92.1 | |
| Petrov et al. (2010) ensembleTraining Set=WSJ only2014.12 | 91.8 | |
| In-order parserSupervision Setting=fully-supervised, Inference Setting=greedy2017.07 | 91.8 | |
| Liu and ZhangSupervision Setting=fully-supervised, Inference Setting=greedy2017.07 | 91.7 | |
| HuangSupervision Setting=fully-supervised, Inference Setting=reranking2017.07 | 91.7 | |
| Charniak and JohnsonSupervision Setting=fully-supervised, Inference Setting=reranking2017.07 | 91.5 | |
| Zhu et al. (2013)Training Set=semi-supervised2014.12 | 91.3 | |
| Huang & Harper (2009)Training Set=semi-supervised2014.12 | 91.3 | |
| Zhu et al. (2013)type=Semisupervised (S)2016.02 | 91.3 | |
| Cross and HuangSupervision Setting=fully-supervised, Inference Setting=greedy2017.07 | 91.3 | |
| Bottom-up parserSupervision Setting=fully-supervised, Inference Setting=greedy2017.07 | 91.3 | |
| Dyer et al.Supervision Setting=fully-supervised, Inference Setting=greedy2017.07 | 91.2 | |
| Top-down parserSupervision Setting=fully-supervised, Inference Setting=greedy2017.07 | 91.2 | |
| Shindo et al. (2012)type=Generative (G), mode=single2016.02 | 91.1 | |
| Shindo et al.Supervision Setting=fully-supervised, Inference Setting=greedy2017.07 | 91.1 | |
| Durrett and KleinSupervision Setting=fully-supervised, Inference Setting=greedy2017.07 | 91.1 | |
| Bod (2003)type=Generative (G)2016.02 | 90.7 | |
| Vinyals et al.Supervision Setting=fully-supervised, Inference Setting=greedy2017.07 | 90.7 | |
| Watanabe and SumitaSupervision Setting=fully-supervised, Inference Setting=greedy2017.07 | 90.7 | |
| LSTM+A+D ensembleTraining Set=WSJ only2014.12 | 90.5 | |
| baseline LSTMTraining Set=BerkeleyParser corpus2014.12 | 90.5 | |
| Petrov et al. (2006)Training Set=WSJ only2014.12 | 90.4 | |
| Zhu et al. (2013)Training Set=WSJ only2014.12 | 90.4 | |
| Socher et al. (2013a)type=Discriminative (D)2016.02 | 90.4 | |
| Zhu et al. (2013)type=Discriminative (D)2016.02 | 90.4 | |
| Socher et al.Supervision Setting=fully-supervised, Inference Setting=greedy2017.07 | 90.4 | |
| Zhu et al.Supervision Setting=fully-supervised, Inference Setting=greedy2017.07 | 90.4 | |
| Petrov and Klein (2007)type=Generative (G)2016.02 | 90.1 | |
| RNNGtype=Discriminative (D), formulation=q(y | x)2016.02 | 89.8 | |
| Henderson (2004)type=Discriminative (D)2016.02 | 89.4 | |
| LSTM+A+DTraining Set=WSJ only2014.12 | 88.3 | |
| Vinyals et al. (2015)type=Discriminative (D), training=WSJ only, ensembling=none2016.02 | 88.3 |