Pronoun Resolution on XWinograd
57.7AccuracyTransformer
Evaluation Results
| Method | Links | |
|---|---|---|
| TransformerTraining=compute-matched, Training dataset=FineWeb-edu2026.05 | 57.7 | |
| DMoEModel Scale=1.7B, Training Strategy=Dynamic Mixture-of-Experts, Evaluation Protocol=Zero-shot2025.06 | 56.4 | |
| DMoEModel Scale=1.7B, Training Strategy=Dynamic Mixture-of-Experts, Evaluation Protocol=Few-shot (4-shot)2025.06 | 56.4 | |
| BLOOM + Continued Pre-trainingModel Scale=1.7B, Training Strategy=Continued Pre-training, Evaluation Protocol=Few-shot (4-shot)2025.06 | 55.8 | |
| Branch-Train-MixModel Scale=1.7B, Training Strategy=Branch-Train-Mix, Evaluation Protocol=Few-shot (4-shot)2025.06 | 55.8 | |
| Branch-Train-MixModel Scale=1.7B, Training Strategy=Branch-Train-Mix, Evaluation Protocol=Zero-shot2025.06 | 55.6 | |
| BLOOM + Continued Pre-trainingModel Scale=1.7B, Training Strategy=Continued Pre-training, Evaluation Protocol=Zero-shot2025.06 | 55.5 | |
| BLOOMModel Scale=1.7B, Training Strategy=Base, Evaluation Protocol=Few-shot (4-shot)2025.06 | 55.5 | |
| DMoEModel Scale=560M, Training Strategy=Dynamic Mixture-of-Experts, Evaluation Protocol=Zero-shot2025.06 | 55.1 | |
| DMoEModel Scale=560M, Training Strategy=Dynamic Mixture-of-Experts, Evaluation Protocol=Few-shot (4-shot)2025.06 | 55.1 | |
| BLOOMModel Scale=1.7B, Training Strategy=Base, Evaluation Protocol=Zero-shot2025.06 | 55.1 | |
| BLOOM + Continued Pre-trainingModel Scale=560M, Training Strategy=Continued Pre-training, Evaluation Protocol=Zero-shot2025.06 | 54.9 | |
| Branch-Train-MixModel Scale=560M, Training Strategy=Branch-Train-Mix, Evaluation Protocol=Zero-shot2025.06 | 54.4 | |
| Branch-Train-MixModel Scale=560M, Training Strategy=Branch-Train-Mix, Evaluation Protocol=Few-shot (4-shot)2025.06 | 54.2 | |
| BLOOM + Continued Pre-trainingModel Scale=560M, Training Strategy=Continued Pre-training, Evaluation Protocol=Few-shot (4-shot)2025.06 | 53.8 | |
| BLOOMModel Scale=560M, Training Strategy=Base, Evaluation Protocol=Zero-shot2025.06 | 53.7 | |
| SRMTraining=compute-matched, Training dataset=FineWeb-edu2026.05 | 53.41 | |
| BLOOMModel Scale=560M, Training Strategy=Base, Evaluation Protocol=Few-shot (4-shot)2025.06 | 53.3 | |
| MambaTraining=compute-matched, Training dataset=FineWeb-edu2026.05 | 51.52 |