Reading Comprehension on DROP 1.0 (test)
92.38EMHuman Performance
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Human Performance2019.08 | 92.38 | 95.98 | |
| Giving BERT a CalculatorOperations=All, Pre-training=CoQA, Ensemble=6 models2019.08 | 78.14 | 81.78 | |
| Giving BERT a CalculatorOperations=All, Pre-training=CoQA2019.08 | 76.96 | 80.53 | |
| MTMSNModel Scale=Large2019.08 | 75.85 | 79.88 | |
| NAQANetModel=NAQANet2019.08 | 44.24 | 47.77 | |
| NAQANet2019.08 | 44.07 | 47.01 | |
| BERTModel Scale=Base2019.08 | 29.45 | 32.7 | |
| QANet+ELMo2019.08 | 27.08 | 29.67 | |
| BiDAF2019.08 | 24.75 | 27.49 | |
| Semantic Role Labeling2019.08 | 10.87 | 13.35 | |
| Heuristic Baseline2019.08 | 4.18 | 8.59 |