Story Cloze Test on Story Cloze (test)
95.75AccuracyTaT
Evaluation Results
| Method | Links | |
|---|---|---|
| TaTTrain Dataset=Hellaswag2026.03 | 95.75 | |
| TaTTrain Dataset=ARC-E2026.03 | 94.98 | |
| EventBERTSize=345M, Type=Pre-trained discriminative model2022.03 | 91.33 | |
| ClarETSize=400M, Type=Pre-trained unified model2022.03 | 91.18 | |
| TaTTrain Dataset=SiQA2026.03 | 90.89 | |
| TaTTrain Dataset=CosQA2026.03 | 90.19 | |
| TaTTrain Dataset=ComQA2026.03 | 88.51 | |
| GPT-3Model Size=175B, Evaluation Protocol=Few-shot, Shots=702021.12 | 87.7 | |
| RoBERTaSize=345M, Type=Pre-trained discriminative model2022.03 | 87.1 | |
| BARTSize=400M, Type=Pre-trained unified model2022.03 | 87.01 | |
| Coherence Boosting (GPT-3 175B)alpha=-0.64, selection=validation-set optimal2021.10 | 86.85 | |
| GLaMModel Size=64B/64E, Evaluation Protocol=Few-shot, Shots=162021.12 | 86.7 | |
| GPT-3Model Size=175B, Evaluation Protocol=One-shot2021.12 | 84.7 | |
| GLaMModel Size=64B/64E, Evaluation Protocol=One-shot2021.12 | 84 | |
| GPT-3Model Size=175B, Evaluation Protocol=Zero-shot2021.12 | 83.2 | |
| Few-shot AccuracyMode=Few-shot2026.03 | 83.1 | |
| GPT-3 175Balpha=-1, protocol=Unconditional Normalization2021.10 | 82.9 | |
| TaTTrain Dataset=ARC-C2026.03 | 82.58 | |
| GLaMModel Size=64B/64E, Evaluation Protocol=Zero-shot2021.12 | 82.5 | |
| TaTTrain Dataset=OpenQA2026.03 | 81.64 | |
| GPT-3 175Balpha=0, protocol=Full-context (fmax)2021.10 | 79.16 | |
| Zero-shot AccuracyMode=Zero-shot2026.03 | 78.8 | |
| Linear ProbeTrain Dataset=Hellaswag2026.03 | 77.61 | |
| Chaturvedi et al.2018.03 | 77.6 | |
| Hidden Coherence ModelType=Task-specific model2022.03 | 77.6 | |
| Coherence Boosting (GPT-2 XL)alpha=-0.69, selection=validation-set optimal2021.10 | 76.75 | |
| val-LS-skipTraining set=validation set, Context encoding=Last Sentence (LS), Embeddings=skip-thought2018.03 | 76.5 | |
| Schwartz et al.2018.03 | 75.2 | |
| GPT-2 XL (1.6B)alpha=-1, protocol=Unconditional Normalization2021.10 | 75.09 | |
| Cai et al.2018.03 | 74.7 | |
| XLNet-Full ModelBackbone=XLNet, Fine-tuned=true2021.10 | 74.62 | |
| XLNet-ContrastiveBackbone=XLNet, Fine-tuned=true2021.10 | 72.83 | |
| val-NC-skipTraining set=validation set, Context encoding=No Context (NC), Embeddings=skip-thought2018.03 | 72.6 | |
| XLNet-PairwiseBackbone=XLNet, Fine-tuned=true2021.10 | 71.84 | |
| val-FC-skipTraining set=validation set, Context encoding=Full Context (FC), Embeddings=skip-thought2018.03 | 71.6 | |
| TaTTrain Dataset=BoolQ2026.03 | 71.41 | |
| Linear ProbeTrain Dataset=ARC-E2026.03 | 69.62 | |
| GPT-2 XL (1.6B)alpha=0, protocol=Full-context (fmax)2021.10 | 67.56 | |
| Linear ProbeTrain Dataset=ARC-C2026.03 | 66.01 | |
| GPT-2 Small (125M)alpha=-1, protocol=Unconditional Normalization2021.10 | 64.78 | |
| Coherence Boosting (GPT-2 Small)alpha=-1.02, selection=validation-set optimal2021.10 | 64.24 | |
| val-LS-GloVeTraining set=validation set, Context encoding=Last Sentence (LS), Embeddings=GloVe2018.03 | 63 | |
| trn-LS-skipTraining set=train set, Context encoding=Last Sentence (LS), Embeddings=skip-thought2018.03 | 62.7 | |
| trn-FC-skipTraining set=train set, Context encoding=Full Context (FC), Embeddings=skip-thought2018.03 | 62.6 | |
| trn-NC-skipTraining set=train set, Context encoding=No Context (NC), Embeddings=skip-thought2018.03 | 60.8 | |
| Linear ProbeTrain Dataset=ComQA2026.03 | 60.72 | |
| GPT-2 Small (125M)alpha=0, protocol=Full-context (fmax)2021.10 | 59.91 | |
| Linear ProbeTrain Dataset=BoolQ2026.03 | 59.45 | |
| Linear ProbeTrain Dataset=OpenQA2026.03 | 57.94 | |
| Linear ProbeTrain Dataset=SiQA2026.03 | 57.91 | |
| Linear ProbeTrain Dataset=CosQA2026.03 | 57.36 | |
| LCD-IBackbone=Infersent2021.10 | 52.69 | |
| LCD-GBackbone=GloVe2021.10 | 51.76 | |
| XLNet-PairwiseBackbone=XLNet, Fine-tuned=false2021.10 | 51.69 | |
| LCD-LBackbone=Language Model2021.10 | 50.09 | |
| UNCBackbone=UNC2021.10 | 49.39 |