Question Generation on SQuAD (BLEU-4)
0.496BLEU-4M
Evaluation Results
| Method | Links | |
|---|---|---|
| Mmode=upperbound2022.09 | 0.496 | |
| bi-gram + APS + round-triptype=ensemble2022.09 | 0.401 | |
| bi-gram + round-triptype=ensemble2022.09 | 0.4 | |
| tri-gram + APS + round-triptype=ensemble2022.09 | 0.4 | |
| tri-gram + round-triptype=ensemble2022.09 | 0.398 | |
| APS + round-triptype=ensemble2022.09 | 0.397 | |
| round-tripselection_criteria=QA semantic equivalence2022.09 | 0.392 | |
| bi-gram + APStype=ensemble2022.09 | 0.384 | |
| tri-gram + APStype=ensemble2022.09 | 0.383 | |
| bi-gramselection_criteria=n-gram similarity, n=22022.09 | 0.382 | |
| tri-gramselection_criteria=n-gram similarity, n=32022.09 | 0.38 | |
| averaged prompt score (APS)selection_criteria=Prompt-based Score2022.09 | 0.38 | |
| overall prompt score (OPS)selection_criteria=Prompt-based Score2022.09 | 0.373 | |
| Mgmode=greedy2022.09 | 0.372 | |
| Msmode=sample avg2022.09 | 0.359 | |
| ERNIE-GEN Largetraining=fine-tuned, author=Xiao et al., 20212022.09 | 0.254 | |
| UniLM v2 Basetraining=fine-tuned, author=Bao et al., 20202022.09 | 0.244 | |
| UniLM Largetraining=fine-tuned, author=Bao et al., 20202022.09 | 0.228 | |
| Msmode=lowerbound2022.09 | 0.225 | |
| Zhang and Bansaltraining=fine-tuned2022.09 | 0.184 | |
| Du and Cardietraining=fine-tuned2022.09 | 0.152 |