Sequential Question Answering on SQA (test)
74.5Accuracy (All)TAPEX
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| TAPEXCategory=Pre-trained Language Models2021.07 | 74.5 | 48.4 | 76.2 | 71.9 | 76.9 | |
| MATEIntermediate pre-training (CS)=true2021.09 | 71.7 | 46.1 | — | — | — | |
| MATEIntermediate pre-training (CS)=false2021.09 | 71.6 | 46.4 | — | — | — | |
| TAPAS (Counterf. + Synthetic)Size=Large, Pre-training Data=Counterfactual + Synthetic2020.10 | 71 | 44.8 | — | — | — | |
| TAPAS (Counterfactual + Synthetic)Data=Counterfactual + Synthetic, Size=Large2020.10 | 71 | 44.8 | 80.9 | 70.6 | 64 | |
| Eisenschlos et al. (2020)Category=Pre-trained Language Models2021.07 | 71 | 44.8 | 80.9 | 70.6 | 64 | |
| TAPASIntermediate pre-training (CS)=true2021.09 | 71 | 44.8 | — | — | — | |
| TAPAS (Counterf. + Synthetic)Size=Base, Pre-training Data=Counterfactual + Synthetic2020.10 | 67.9 | 40.5 | — | — | — | |
| TAPAS (Counterfactual + Synthetic)Data=Counterfactual + Synthetic, Size=Base2020.10 | 67.9 | 40.5 | 79.3 | 67 | 61.1 | |
| TAPAS (Synthetic)Size=Base, Pre-training Data=Synthetic2020.10 | 67.4 | 39.8 | — | — | — | |
| TAPAS (Synthetic)Data=Synthetic, Size=Base2020.10 | 67.4 | 39.8 | 79.3 | 66.2 | 60.2 | |
| Herzig et al.Size=Large2020.10 | 67.2 | 40.4 | — | — | — | |
| Herzig et al. (2020)Category=Pre-trained Language Models2021.07 | 67.2 | 40.4 | 78.2 | 66 | 50.7 | |
| TAPAS2020.04 | 67.2 | 40.4 | 78.2 | 66 | 59.7 | |
| TAPASIntermediate pre-training (CS)=false2021.09 | 67.2 | 40.4 | — | — | — | |
| Yu et al. (2021b)Category=Pre-trained Language Models2021.07 | 65.4 | 38.5 | 78.4 | 65.3 | 55.1 | |
| TAPAS (Counterfactual)Size=Base, Pre-training Data=Counterfactual2020.10 | 65 | 36.5 | — | — | — | |
| TAPAS (Counterfactual)Data=Counterfactual, Size=Base2020.10 | 65 | 36.5 | 78.4 | 63.7 | 57.5 | |
| TAPAS (MASK-LM)Size=Base, Pre-training Data=MASK-LM2020.10 | 64 | 34.6 | — | — | — | |
| TAPAS (MASK-LM)Data=MASK-LM, Size=Base2020.10 | 64 | 34.6 | 79.2 | 61.2 | 55.6 | |
| BARTCategory=Pre-trained Language Models2021.07 | 58.6 | 27.8 | 65.3 | 54.1 | 57 | |
| Mueller et al.2020.10 | 55.1 | 28.1 | — | — | — | |
| Mueller et al. (2019)Category=Previous Systems2021.07 | 55.1 | 28.1 | 67.2 | 52.7 | 46.8 | |
| Müller et al. (2019)2020.04 | 55.1 | 28.1 | 67.2 | 52.7 | 46.8 | |
| Sun et al. (2019)Category=Previous Systems2021.07 | 45.6 | 13.2 | 70.3 | 42.6 | 24.8 | |
| Sun et al. (2018)2020.04 | 45.6 | 13.2 | 70.3 | 42.6 | 24.8 | |
| Iyyer et al.2020.10 | 44.7 | 12.8 | — | — | — | |
| Iyyer et al. (2017)Category=Previous Systems2021.07 | 44.7 | 12.8 | 70.4 | 41.1 | 23.6 | |
| Iyyer et al. (2017)2020.04 | 44.7 | 12.8 | 70.4 | 41.1 | 23.6 | |
| Neelakantan et al. (2017)Category=Previous Systems2021.07 | 40.2 | 11.8 | 60 | 35.9 | 25.5 | |
| Neelakantan et al. (2017)2020.04 | 40.2 | 11.8 | 60 | 35.9 | 25.5 | |
| Pasupat & Liang (2015)Category=Previous Systems2021.07 | 33.2 | 7.7 | 51.4 | 22.2 | 22.3 | |
| Pasupat and Liang (2015)2020.04 | 33.2 | 7.7 | 51.4 | 22.2 | 22.3 | |
| Liu et al. (2019)Category=Previous Systems2021.07 | — | — | 70.9 | 39.5 | — |