Natural Language Understanding on GLUE SST-2, QQP, MNLI-m, MNLI-mm official (test)
94.9SST-2 AccuracyBERT_LARGE
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| BERT_LARGEnumber of parameters=335M2019.03 | 94.9 | 72.1 | 89.3 | 86.7 | 85.9 | |
| BERT_BASE2019.03 | 93.5 | 71.2 | 89.2 | 84.6 | 83.4 | |
| OpenAI GPT2019.03 | 91.3 | 70.3 | 88.5 | 82.1 | 81.4 | |
| Distilled BiLSTMSOFTnumber of parameters=0.96M2019.03 | 90.7 | 68.2 | 88.1 | 73 | 72.6 | |
| BERT ELMo baselinenumber of parameters=93.6M, source=Devlin et al., 20182019.03 | 90.4 | 64.8 | 84.7 | 76.4 | 76.1 | |
| GLUE ELMo baselinenumber of parameters=93.6M, source=Wang et al., 20182019.03 | 90.4 | 63.1 | 84.3 | 74.1 | 74.5 | |
| BiLSTM (reported by other papers)source=Aggregated2019.03 | 87.6 | — | 82.6 | 66.9 | 66.9 | |
| BiLSTM (our implementation)implementation=in-house2019.03 | 86.7 | 63.7 | 86.2 | 68.7 | 68.3 | |
| BiLSTM (reported by GLUE)source=GLUE benchmark2019.03 | 85.9 | 61.4 | 81.7 | 70.3 | 70.8 |