Natural Language Inference on XNLI 1.0 (test)
89.7Accuracy (en)InfoXLM
Evaluation Results
| Method | Links | ||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| InfoXLMZero-shot=true2020.08 | 89.7 | 81.4 | 84.5 | 85.5 | 84.1 | 83.4 | 84.2 | 81.3 | 80.9 | 80.4 | 80.8 | 78.9 | 80.9 | 77.9 | 74.8 | 73.7 | |
| INFOXLMEvaluation Protocol=Cross-lingual Transfer, Model Scale=LARGE2020.12 | 89.7 | 81.4 | 84.5 | 85.5 | 84.1 | 83.4 | 84.2 | 81.3 | 80.9 | 80.4 | 80.8 | 78.9 | 80.9 | 77.9 | 74.8 | 73.7 | |
| R4FModel=XLM-R Large, Zero-shot=true2020.08 | 89.6 | 81.4 | 84.7 | 85.2 | 84.2 | 83.6 | 84.6 | 82.5 | 80.3 | 80.5 | 80.9 | 79.2 | 80.6 | 78.2 | 72.7 | 73.9 | |
| ERNIE-MEvaluation Protocol=Translate-Train-All, Model Scale=LARGE2020.12 | 89.5 | 84.2 | 86.5 | 86.9 | 86.1 | 86 | 86.8 | 84.1 | 83.8 | 84.1 | 84.5 | 82.1 | 83.5 | 81.1 | 79.4 | 77.9 | |
| R3FModel=XLM-R Large, Zero-shot=true2020.08 | 89.4 | 81.2 | 84.2 | 85.1 | 83.7 | 83.6 | 84.6 | 82.3 | 80.7 | 80.6 | 81.1 | 79.4 | 80.1 | 77.3 | 72.6 | 74.2 | |
| ERNIE-MEvaluation Protocol=Cross-lingual Transfer, Model Scale=LARGE2020.12 | 89.3 | 82 | 85.1 | 85.7 | 84.4 | 83.7 | 84.5 | 82 | 81.2 | 81.2 | 81.9 | 79.2 | 81 | 78.6 | 76.2 | 75.4 | |
| XLM-R LargeModel=XLM-R Large, Zero-shot=true2020.08 | 89.1 | 80.9 | 84.1 | 85.1 | 83.9 | 82.9 | 84 | 81.2 | 79.6 | 79.8 | 80.8 | 78.1 | 80.2 | 76.9 | 73.9 | 73.8 | |
| XLM-REvaluation Protocol=Cross-lingual Transfer, Model Scale=LARGE2020.12 | 89.1 | 80.9 | 84.1 | 85.1 | 83.9 | 82.9 | 84 | 81.2 | 79.6 | 79.8 | 80.8 | 78.1 | 80.2 | 76.9 | 73.9 | 73.8 | |
| XLM-REvaluation Protocol=Translate-Train-All, Model Scale=LARGE2020.12 | 89.1 | 83.6 | 85.1 | 86.6 | 85.7 | 85.3 | 85.9 | 83.5 | 83.2 | 83.1 | 83.7 | 81.5 | 83.7 | 81.6 | 78 | 78.1 | |
| VECOEvaluation Protocol=Translate-Train-All, Model Scale=LARGE2020.12 | 88.9 | 83 | 82.4 | 86 | 84.7 | 85.3 | 86.2 | 85.8 | 80.1 | 83 | 77.2 | 80.9 | 82.8 | 75.3 | 83.1 | 83 | |
| VECOEvaluation Protocol=Cross-lingual Transfer, Model Scale=LARGE2020.12 | 88.2 | 79.9 | 79.2 | 83.1 | 82.9 | 81.2 | 84.2 | 82.8 | 76.2 | 80.3 | 74.3 | 77 | 78.4 | 71.3 | 80.4 | 79.1 | |
| INFOXLMEvaluation Protocol=Cross-lingual Transfer, Model Scale=Base2020.12 | 86.4 | 76.2 | 80.6 | 80.8 | 78.9 | 77.8 | 78.9 | 77.6 | 75.6 | 74 | 77 | 73.7 | 76.7 | 72 | 66.4 | 67.1 | |
| ERNIE-MEvaluation Protocol=Translate-Train-All, Model Scale=Base2020.12 | 86.2 | 80.6 | 82.5 | 83.8 | 82.6 | 82.4 | 83.4 | 80.2 | 80.6 | 80.5 | 81.1 | 79.2 | 80.5 | 77.7 | 75 | 73.3 | |
| INFOXLMEvaluation Protocol=Translate-Train-All, Model Scale=Base2020.12 | 86.1 | 79.7 | 82 | 82.8 | 81.8 | 80.9 | 82 | 80.2 | 79 | 78.8 | 80.5 | 78.3 | 80.5 | 77.4 | 73 | 71.6 | |
| XLM-R BaseModel=XLM-R Base, Zero-shot=true2020.08 | 85.8 | 76.2 | 79.7 | 80.7 | 78.7 | 77.5 | 79.6 | 78.1 | 74.2 | 73.8 | 76.5 | 74.6 | 76.7 | 72.4 | 66.5 | 68.3 | |
| XLM-REvaluation Protocol=Cross-lingual Transfer, Model Scale=Base2020.12 | 85.8 | 76.2 | 79.7 | 80.7 | 78.7 | 77.5 | 79.6 | 78.1 | 74.2 | 73.8 | 76.5 | 74.6 | 76.7 | 72.4 | 66.5 | 68.3 | |
| UnicoderEvaluation Protocol=Translate-Train-All, Model Scale=Base2020.12 | 85.6 | 78.5 | 81.1 | 82.3 | 80.9 | 79.5 | 81.4 | 79.7 | 76.8 | 78.2 | 77.9 | 77.1 | 80.5 | 73.4 | 73.8 | 69.6 | |
| ERNIE-MEvaluation Protocol=Cross-lingual Transfer, Model Scale=Base2020.12 | 85.5 | 77.3 | 80.1 | 81.2 | 79.2 | 79.1 | 80.4 | 78.1 | 76.8 | 76.3 | 78.3 | 75.8 | 77.4 | 72.9 | 69.5 | 68.8 | |
| XLM-REvaluation Protocol=Translate-Train-All, Model Scale=Base2020.12 | 85.4 | 79.1 | 81.4 | 82.2 | 80.3 | 80.4 | 81.3 | 79.7 | 78.6 | 77.3 | 79.7 | 77.9 | 80.2 | 76.1 | 73.1 | 73 | |
| UnicoderEvaluation Protocol=Cross-lingual Transfer, Model Scale=Base2020.12 | 85.1 | 75.4 | 79 | 79.4 | 77.8 | 77.2 | 77.2 | 76.3 | 72.8 | 73.5 | 76.4 | 73.6 | 76.2 | 69.4 | 69.7 | 66.7 | |
| XLMEvaluation Protocol=Cross-lingual Transfer, Model Scale=Base2020.12 | 85 | 75.1 | 78.7 | 78.9 | 77.8 | 76.6 | 77.4 | 75.3 | 72.5 | 73.1 | 76.1 | 73.2 | 76.5 | 69.6 | 68.4 | 67.3 | |
| XLMEvaluation Protocol=Translate-Train-All, Model Scale=Base2020.12 | 85 | 77.8 | 80.8 | 81.3 | 80.3 | 79.1 | 80.9 | 78.3 | 75.6 | 77.6 | 78.5 | 76 | 79.5 | 72.9 | 72.8 | 68.5 | |
| TokAlign++Backbone=Pythia 6.9B, Tuning Stage=1k steps tuning on multilingual corpus2026.05 | 54.6 | — | — | — | 36.7 | — | — | — | — | 40.4 | 42.9 | 40.1 | 42.7 | — | — | 36.3 | |
| PythiaBackbone=Pythia 6.9B, Tuning Stage=Vanilla (Zero-shot)2026.05 | 54.4 | — | — | — | 39 | — | — | — | — | 39.3 | 39.3 | 39.8 | 46.2 | — | — | 36.4 | |
| TokAlignBackbone=Pythia 6.9B, Tuning Stage=1k steps tuning on multilingual corpus2026.05 | 53.7 | — | — | — | 34.9 | — | — | — | — | 39.1 | 40.1 | 39.2 | 41.7 | — | — | 36.1 | |
| ZeTTBackbone=Pythia 6.9B, Tuning Stage=Initialization without any tuning2026.05 | 53.2 | — | — | — | 35.8 | — | — | — | — | 34.6 | 33.9 | 33.4 | 34.1 | — | — | 32.3 | |
| TokAlign++Backbone=Pythia 6.9B, Tuning Stage=Initialization without any tuning2026.05 | 53.2 | — | — | — | 36.7 | — | — | — | — | 34.7 | 35.5 | 34.6 | 35 | — | — | 34.4 | |
| TokAlignBackbone=Pythia 6.9B, Tuning Stage=Initialization without any tuning2026.05 | 52.6 | — | — | — | 35.1 | — | — | — | — | 34.3 | 33.8 | 34.5 | 34.4 | — | — | 33.2 | |
| PythiaBackbone=Pythia 1B, Tuning Stage=Vanilla (Zero-shot)2026.05 | 51 | — | — | — | 37.8 | — | — | — | — | 35.9 | 37 | 34.8 | 42.6 | — | — | 34.7 | |
| TokAlign++Backbone=Pythia 1B, Tuning Stage=1k steps tuning on multilingual corpus2026.05 | 50.8 | — | — | — | 40.6 | — | — | — | — | 38.4 | 39.4 | 37.6 | 43.2 | — | — | 35.6 | |
| ZeTTBackbone=Pythia 6.9B, Tuning Stage=1k steps tuning on multilingual corpus2026.05 | 50.6 | — | — | — | 36.4 | — | — | — | — | 37.4 | 38.5 | 35.3 | 38.7 | — | — | 34.3 | |
| TokAlign++Backbone=Pythia 1B, Tuning Stage=Initialization without any tuning2026.05 | 49.9 | — | — | — | 37.4 | — | — | — | — | 34 | 35.1 | 34.2 | 33.9 | — | — | 34.7 | |
| TokAlignBackbone=Pythia 1B, Tuning Stage=1k steps tuning on multilingual corpus2026.05 | 49.1 | — | — | — | 38.6 | — | — | — | — | 36.1 | 38.5 | 34.9 | 39.4 | — | — | 34.3 | |
| TokAlignBackbone=Pythia 1B, Tuning Stage=Initialization without any tuning2026.05 | 48.1 | — | — | — | 35.7 | — | — | — | — | 32.8 | 33.4 | 32.9 | 34.7 | — | — | 32.8 | |
| ZeTTBackbone=Pythia 1B, Tuning Stage=1k steps tuning on multilingual corpus2026.05 | 48 | — | — | — | 39.8 | — | — | — | — | 35.6 | 37.2 | 34.8 | 38.7 | — | — | 34.7 | |
| FocusBackbone=Pythia 6.9B, Tuning Stage=1k steps tuning on multilingual corpus2026.05 | 47 | — | — | — | 35.1 | — | — | — | — | 35.1 | 35.4 | 34.6 | 36.6 | — | — | 33.7 | |
| ZeTTBackbone=Pythia 1B, Tuning Stage=Initialization without any tuning2026.05 | 46.2 | — | — | — | 34.9 | — | — | — | — | 33 | 34.2 | 33.4 | 33.9 | — | — | 33.7 | |
| FocusBackbone=Pythia 1B, Tuning Stage=1k steps tuning on multilingual corpus2026.05 | 44.4 | — | — | — | 36.7 | — | — | — | — | 34.5 | 34.6 | 33.9 | 35.3 | — | — | 34.5 | |
| FocusBackbone=Pythia 1B, Tuning Stage=Initialization without any tuning2026.05 | 33.8 | — | — | — | 34.6 | — | — | — | — | 35.4 | 34.4 | 33.8 | 34.7 | — | — | 33.9 | |
| FocusBackbone=Pythia 6.9B, Tuning Stage=Initialization without any tuning2026.05 | 33.1 | — | — | — | 32.7 | — | — | — | — | 33.4 | 33 | 33.3 | 33.7 | — | — | 32.1 | |
| BLOOMModel Scale=560M, Training Strategy=Base, Evaluation Protocol=Zero-shot2025.06 | — | 36.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| BLOOMModel Scale=560M, Training Strategy=Base, Evaluation Protocol=Few-shot (4-shot)2025.06 | — | 34.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| BLOOMModel Scale=1.7B, Training Strategy=Base, Evaluation Protocol=Zero-shot2025.06 | — | 39.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| BLOOMModel Scale=1.7B, Training Strategy=Base, Evaluation Protocol=Few-shot (4-shot)2025.06 | — | 37.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| BLOOM + Continued Pre-trainingModel Scale=560M, Training Strategy=Continued Pre-training, Evaluation Protocol=Zero-shot2025.06 | — | 37.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| BLOOM + Continued Pre-trainingModel Scale=560M, Training Strategy=Continued Pre-training, Evaluation Protocol=Few-shot (4-shot)2025.06 | — | 34.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| BLOOM + Continued Pre-trainingModel Scale=1.7B, Training Strategy=Continued Pre-training, Evaluation Protocol=Zero-shot2025.06 | — | 39.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| BLOOM + Continued Pre-trainingModel Scale=1.7B, Training Strategy=Continued Pre-training, Evaluation Protocol=Few-shot (4-shot)2025.06 | — | 37.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Branch-Train-MixModel Scale=560M, Training Strategy=Branch-Train-Mix, Evaluation Protocol=Zero-shot2025.06 | — | 37.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Branch-Train-MixModel Scale=560M, Training Strategy=Branch-Train-Mix, Evaluation Protocol=Few-shot (4-shot)2025.06 | — | 35.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Branch-Train-MixModel Scale=1.7B, Training Strategy=Branch-Train-Mix, Evaluation Protocol=Zero-shot2025.06 | — | 39.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Branch-Train-MixModel Scale=1.7B, Training Strategy=Branch-Train-Mix, Evaluation Protocol=Few-shot (4-shot)2025.06 | — | 36.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DMoEModel Scale=560M, Training Strategy=Dynamic Mixture-of-Experts, Evaluation Protocol=Zero-shot2025.06 | — | 37.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DMoEModel Scale=560M, Training Strategy=Dynamic Mixture-of-Experts, Evaluation Protocol=Few-shot (4-shot)2025.06 | — | 35.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DMoEModel Scale=1.7B, Training Strategy=Dynamic Mixture-of-Experts, Evaluation Protocol=Zero-shot2025.06 | — | 39.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DMoEModel Scale=1.7B, Training Strategy=Dynamic Mixture-of-Experts, Evaluation Protocol=Few-shot (4-shot)2025.06 | — | 37.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | — |