Question Answering on SQuAD (dev)
91F1Human
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Human2016.11 | 91 | 81.4 | — | |
| Baselinew-bits=32, e-bits=32, Size (MB)=415.4, Size-w/o-e (MB)=324.52019.09 | 88.69 | 81.54 | — | |
| Q-BERTw-bits=8, e-bits=8, Size (MB)=103.9, Size-w/o-e (MB)=81.22019.09 | 88.47 | 81.07 | — | |
| Q-BERTw-bits=4, e-bits=8, Size (MB)=63.4, Size-w/o-e (MB)=40.62019.09 | 88.36 | 80.95 | — | |
| Q-BERTw-bits=3, e-bits=8, Size (MB)=53.2, Size-w/o-e (MB)=30.52019.09 | 87.66 | 79.96 | — | |
| Q-BERTMPw-bits=2/4 MP, e-bits=8, Size (MB)=53.2, Size-w/o-e (MB)=30.52019.09 | 87.49 | 79.85 | — | |
| multi-domain debiasing frameworkSetting=Five Domains, Backbone=ELECTRA-small2026.01 | 87.4 | 79.9 | — | |
| Q-BERTMPw-bits=2/3 MP, e-bits=8, Size (MB)=48.1, Size-w/o-e (MB)=25.42019.09 | 86.95 | 79.29 | — | |
| ELECTRA-small (Baseline)Setting=Baseline, Backbone=ELECTRA-small2026.01 | 85.9 | 78.1 | — | |
| ORACLEQA Model=DCN+, Train Speedup=x3.0, Inference Speedup=x5.12018.05 | 85.1 | 76 | — | |
| ORACLEQA Model=S-Reader, Train Speedup=x6.7, Inference Speedup=x5.12018.05 | 84.3 | 74.9 | — | |
| FusionNetQA Model=DCN+2018.05 | 83.6 | 75.3 | — | |
| FULLQA Model=DCN+, Train Speedup=x1.0, Inference Speedup=x1.02018.05 | 83.1 | 74.5 | — | |
| MINIMAL (Dyn)QA Model=DCN+, Selection Method=Dynamic, Train Speedup=x3.0, Inference Speedup=x3.72018.05 | 80.6 | 72 | — | |
| FewshotQA w/ MINPROMPTNumber of training examples=128, Number of Parameters=406M2023.10 | 80.5 | — | — | |
| FULLQA Model=S-Reader, Train Speedup=x1.0, Inference Speedup=x1.02018.05 | 79.9 | 71 | — | |
| MINIMAL (Dyn)QA Model=S-Reader, Selection Method=Dynamic, Train Speedup=x6.7, Inference Speedup=x3.62018.05 | 79.8 | 70.9 | — | |
| PMRNumber of training examples=128, Number of Parameters=406M2023.10 | 79.8 | — | — | |
| Q-BERTw-bits=2, e-bits=8, Size (MB)=43.1, Size-w/o-e (MB)=20.42019.09 | 79.6 | 69.68 | — | |
| MINIMAL (Top k)QA Model=DCN+, Selection Method=Top k (k=1), Train Speedup=x3.0, Inference Speedup=x5.12018.05 | 79.2 | 70.7 | — | |
| FewshotQA w/ MINPROMPTNumber of training examples=64, Number of Parameters=406M2023.10 | 79.2 | — | — | |
| DrQAconfiguration=Document Reader Only2017.03 | 78.8 | 69.5 | — | |
| FewshotQANumber of training examples=128, Number of Parameters=406M2023.10 | 78.8 | — | — | |
| MINIMAL (Top k)QA Model=S-Reader, Selection Method=Top k (k=1), Train Speedup=x6.7, Inference Speedup=x5.12018.05 | 78.7 | 69.9 | — | |
| FastQAQA Model=DCN+2018.05 | 78.5 | 70.3 | — | |
| FewshotQA w/ MINPROMPTNumber of training examples=32, Number of Parameters=406M2023.10 | 78 | — | — | |
| FewshotQANumber of training examples=64, Number of Parameters=406M2023.10 | 77.9 | — | — | |
| BiDAF2017.03 | 77.3 | 67.7 | — | |
| DirectQw-bits=4, e-bits=8, Size (MB)=63.4, Size-w/o-e (MB)=40.62019.09 | 77.1 | 66.05 | — | |
| Multi-Perspective Matching2017.03 | 75.8 | 66.1 | — | |
| Dynamic Coattention Networks2017.03 | 75.6 | 65.4 | — | |
| RASORRepresentation=Recurrent span representations2016.11 | 74.9 | 66.4 | — | |
| FewshotQANumber of training examples=32, Number of Parameters=406M2023.10 | 73.8 | — | — | |
| FewshotQA w/ MINPROMPTNumber of training examples=16, Number of Parameters=406M2023.10 | 73.6 | — | — | |
| FG fine-grained gate + ensembleintegration=Fine-grained Gating, ensemble=true2016.11 | 73.41 | 62.38 | — | |
| SplinterNumber of training examples=128, Number of Parameters=110M2023.10 | 72.7 | — | — | |
| FewshotQANumber of training examples=16, Number of Parameters=406M2023.10 | 72.5 | — | — | |
| FG fine-grained gateintegration=Fine-grained Gating2016.11 | 71.25 | 59.95 | — | |
| DCR2016.10 | 71.2 | 62.5 | — | |
| Yu et al. (2016)2016.11 | 71.2 | 62.5 | — | |
| PMRNumber of training examples=64, Number of Parameters=406M2023.10 | 71.2 | — | — | |
| Match-LSTMVariant=Boundary2016.11 | 70.7 | 60.5 | — | |
| Splinter w/ MINPROMPTNumber of training examples=128, Number of Parameters=110M2023.10 | 70.2 | — | — | |
| Wang 20162016.10 | 70 | 59.1 | — | |
| Wang & Jiang (2016)2016.11 | 70 | 59.1 | — | |
| PMRNumber of training examples=32, Number of Parameters=406M2023.10 | 70 | — | — | |
| GA fine-grained gateintegration=Gated Attention, gating=fine-grained2016.11 | 69.83 | 58.04 | — | |
| GA word char feat concatintegration=Gated Attention, features=word + char + features (concat)2016.11 | 69.04 | 57.11 | — | |
| Splinter w/ MINPROMPTNumber of training examples=64, Number of Parameters=110M2023.10 | 68.6 | — | — | |
| GA word char concatintegration=Gated Attention, features=word + char (concat)2016.11 | 68.57 | 56.39 | — | |
| GA scalar gateintegration=Gated Attention, gating=scalar2016.11 | 68.5 | 56.2 | — | |
| Match-LSTMVariant=Sequence2016.11 | 67.7 | 54.5 | — | |
| GA wordintegration=Gated Attention, features=word2016.11 | 66.95 | 54.92 | — | |
| SplinterNumber of training examples=64, Number of Parameters=110M2023.10 | 65.2 | — | — | |
| Splinter w/ MINPROMPTNumber of training examples=32, Number of Parameters=110M2023.10 | 64.6 | — | — | |
| PMRNumber of training examples=16, Number of Parameters=406M2023.10 | 60.3 | — | — | |
| DirectQw-bits=3, e-bits=8, Size (MB)=53.2, Size-w/o-e (MB)=30.52019.09 | 59.83 | 46.77 | — | |
| SplinterNumber of training examples=32, Number of Parameters=110M2023.10 | 59.2 | — | — | |
| Splinter w/ MINPROMPTNumber of training examples=16, Number of Parameters=110M2023.10 | 58.9 | — | — | |
| SpanBERTNumber of training examples=128, Number of Parameters=110M2023.10 | 55.8 | — | — | |
| SplinterNumber of training examples=16, Number of Parameters=110M2023.10 | 54.6 | — | — | |
| Rajpurkar 20162016.10 | 51 | 39.8 | — | |
| Logistic regression baseline2016.11 | 51 | 39.8 | — | |
| SpanBERTNumber of training examples=64, Number of Parameters=110M2023.10 | 45.8 | — | — | |
| RoBERTaNumber of training examples=128, Number of Parameters=110M2023.10 | 43 | — | — | |
| KERMITtype=Insertion, zero-shot=true2019.06 | 30.3 | — | — | |
| RoBERTaNumber of training examples=64, Number of Parameters=110M2023.10 | 28.4 | — | — | |
| SpanBERTNumber of training examples=32, Number of Parameters=110M2023.10 | 25.8 | — | — | |
| BERTtype=Masking, zero-shot=true2019.06 | 18.9 | — | — | |
| SpanBERTNumber of training examples=16, Number of Parameters=110M2023.10 | 18.2 | — | — | |
| RoBERTaNumber of training examples=32, Number of Parameters=110M2023.10 | 18.2 | — | — | |
| Autoregressive (Transformer, GPT, GPT-2)type=Autoregressive, zero-shot=true2019.06 | 16.6 | — | — | |
| DirectQw-bits=2, e-bits=8, Size (MB)=43.1, Size-w/o-e (MB)=20.42019.09 | 10.32 | 4.77 | — | |
| RoBERTaNumber of training examples=16, Number of Parameters=110M2023.10 | 7.5 | — | — | |
| AutoTinyBERT-Fast-S5-432-720-6-384Speedup=10.3x2021.07 | — | — | 80 | |
| AutoTinyBERT-KD-S1Speedup=4.6x2021.07 | — | — | 87.6 | |
| AutoTinyBERT-KD-S2Speedup=9.0x2021.07 | — | — | 84.6 | |
| AutoTinyBERT-KD-S3Speedup=10.7x2021.07 | — | — | 83.3 | |
| AutoTinyBERT-KD-S4Speedup=17.0x2021.07 | — | — | 78.7 | |
| AutoTinyBERT-S5-450-636-6-384Speedup=10.8x2021.07 | — | — | 79.7 | |
| baseline (B)baseline=true2017.06 | — | 52.58 | — | |
| BERT-KD-S1Speedup=4.9x2021.07 | — | — | 86.2 | |
| BERT-KD-S2Speedup=9.8x2021.07 | — | — | 82.5 | |
| BERT-KD-S3Speedup=11.7x2021.07 | — | — | 81.6 | |
| BERT-KD-S4Speedup=17.0x2021.07 | — | — | 77.4 | |
| BERT-S4-384-1536-6-384Speedup=9.3x2021.07 | — | — | 78.5 | |
| dict, LSTM (D4)dictionary=true, aggregation=LSTM2017.06 | — | 58.78 | — | |
| dict, MP, sum (D2)dictionary=true, aggregation=mean pooling, summation=true, back-propagation=true2017.06 | — | 57.03 | — | |
| dict, MP, sum, no back-prop (D1)dictionary=true, aggregation=mean pooling, summation=true, back-propagation=false2017.06 | — | 56.27 | — | |
| dict, MP, transform and sum (D3)dictionary=true, aggregation=mean pooling, transformation=true, summation=true2017.06 | — | 58.9 | — | |
| DistilBERT-4LSpeedup=3.0x2021.07 | — | — | 81.2 | |
| DODOCompression Ratio=5x2023.10 | — | 59.1 | — | |
| DODOCompression Ratio=10x2023.10 | — | 49.8 | — | |
| FULLCompression Ratio=1x2023.10 | — | 64.5 | — | |
| GloVe (G)embeddings=GloVe2017.06 | — | 64.19 | — | |
| LMSUMMCompression Ratio=10x2023.10 | — | 30.9 | — | |
| MiniLM-4L312DSpeedup=9.8x2021.07 | — | — | 82.1 | |
| MiniLM-4L516DSpeedup=4.9x2021.07 | — | — | 85.5 | |
| MobileBERT TinySpeedup=3.6*x2021.07 | — | — | 88.6 | |
| NoDocCompression Ratio=infinity2023.10 | — | 1.4 | — |