Paraphrase Identification on PAWS-X
91.7AccuracyByT5
Evaluation Results
| Method | Links | |
|---|---|---|
| ByT5Scale=XXL, Evaluation Protocol=Translate-train2021.05 | 91.7 | |
| mT5Scale=XXL, Evaluation Protocol=Translate-train2021.05 | 91.5 | |
| mT5Scale=Large, Evaluation Protocol=Translate-train2021.05 | 91.3 | |
| mT5Scale=XL, Evaluation Protocol=Translate-train2021.05 | 91 | |
| ByT5Scale=Large, Evaluation Protocol=Translate-train2021.05 | 90.6 | |
| mT5Scale=Base, Evaluation Protocol=Translate-train2021.05 | 90.5 | |
| ByT5Scale=XL, Evaluation Protocol=Translate-train2021.05 | 90.5 | |
| ByT5Scale=XXL, Evaluation Protocol=Cross-lingual zero-shot transfer2021.05 | 90.1 | |
| mT5Scale=XXL, Evaluation Protocol=Cross-lingual zero-shot transfer2021.05 | 90 | |
| ByT5Scale=Base, Evaluation Protocol=Translate-train2021.05 | 89.8 | |
| mT5Scale=XL, Evaluation Protocol=Cross-lingual zero-shot transfer2021.05 | 89.6 | |
| mT5Scale=Large, Evaluation Protocol=Cross-lingual zero-shot transfer2021.05 | 88.9 | |
| ByT5Scale=Small, Evaluation Protocol=Translate-train2021.05 | 88.6 | |
| ByT5Scale=XL, Evaluation Protocol=Cross-lingual zero-shot transfer2021.05 | 88.6 | |
| mT5Scale=Small, Evaluation Protocol=Translate-train2021.05 | 87.7 | |
| ByT5Scale=Large, Evaluation Protocol=Cross-lingual zero-shot transfer2021.05 | 87.4 | |
| mT5Scale=Base, Evaluation Protocol=Cross-lingual zero-shot transfer2021.05 | 86.4 | |
| ByT5Scale=Base, Evaluation Protocol=Cross-lingual zero-shot transfer2021.05 | 86.3 | |
| EXAONE-2.4B + SCRIPTBackbone=EXAONE-2.4B, Enhancement=SCRIPT2026.04 | 85.9 | |
| EXAONE-2.4BBackbone=EXAONE-2.4B2026.04 | 85.24 | |
| ByT5Scale=Small, Evaluation Protocol=Cross-lingual zero-shot transfer2021.05 | 84 | |
| mT5Scale=Small, Evaluation Protocol=Cross-lingual zero-shot transfer2021.05 | 82.4 | |
| KoGPT3-1.2B + SCRIPTBackbone=KoGPT3-1.2B, Enhancement=SCRIPT2026.04 | 79.95 | |
| KoGPT3-1.2BBackbone=KoGPT3-1.2B2026.04 | 77.4 | |
| KoGPT2base + SCRIPTBackbone=KoGPT2base, Enhancement=SCRIPT2026.04 | 76.61 | |
| KoGPT2baseBackbone=KoGPT2base2026.04 | 76.33 | |
| BERTbase + SCRIPTBackbone=BERTbase, Enhancement=SCRIPT2026.04 | 73.68 | |
| KOMBObaseBackbone=KOMBObase2026.04 | 73.4 | |
| BERTbaseBackbone=BERTbase2026.04 | 72.38 | |
| Parallel Only (OPENSEAL)Param=7B2026.02 | 64.7 | |
| Sailor2Param=8B2026.02 | 64.53 | |
| SEA-LION v3.5Param=8B2026.02 | 63.58 | |
| SeaLLM v3Param=7B2026.02 | 60.18 | |
| MultilingualParam=7B2026.02 | 60.08 | |
| mGPT13Bk-shot=162022.04 | 55.1 | |
| DMoEModel Scale=1.7B, Training Strategy=Dynamic Mixture-of-Experts, Evaluation Protocol=Zero-shot2025.06 | 54.1 | |
| BLOOMModel Scale=1.7B, Training Strategy=Base, Evaluation Protocol=Zero-shot2025.06 | 53.9 | |
| BLOOM + Continued Pre-trainingModel Scale=1.7B, Training Strategy=Continued Pre-training, Evaluation Protocol=Zero-shot2025.06 | 53.7 | |
| Branch-Train-MixModel Scale=1.7B, Training Strategy=Branch-Train-Mix, Evaluation Protocol=Zero-shot2025.06 | 53.5 | |
| DMoEModel Scale=560M, Training Strategy=Dynamic Mixture-of-Experts, Evaluation Protocol=Zero-shot2025.06 | 53.2 | |
| Branch-Train-MixModel Scale=560M, Training Strategy=Branch-Train-Mix, Evaluation Protocol=Zero-shot2025.06 | 53.1 | |
| mGPT1.3Bk-shot=02022.04 | 53.1 | |
| BLOOM + Continued Pre-trainingModel Scale=560M, Training Strategy=Continued Pre-training, Evaluation Protocol=Zero-shot2025.06 | 52.9 | |
| DMoEModel Scale=560M, Training Strategy=Dynamic Mixture-of-Experts, Evaluation Protocol=Few-shot (4-shot)2025.06 | 52.2 | |
| DMoEModel Scale=1.7B, Training Strategy=Dynamic Mixture-of-Experts, Evaluation Protocol=Few-shot (4-shot)2025.06 | 52.2 | |
| mGPT1.3Bk-shot=42022.04 | 52.2 | |
| mGPT1.3Bk-shot=162022.04 | 52.2 | |
| BLOOM + Continued Pre-trainingModel Scale=1.7B, Training Strategy=Continued Pre-training, Evaluation Protocol=Few-shot (4-shot)2025.06 | 51.8 | |
| BLOOM + Continued Pre-trainingModel Scale=560M, Training Strategy=Continued Pre-training, Evaluation Protocol=Few-shot (4-shot)2025.06 | 51.6 | |
| mGPT13Bk-shot=42022.04 | 51.6 | |
| BLOOMModel Scale=560M, Training Strategy=Base, Evaluation Protocol=Zero-shot2025.06 | 51.5 | |
| mGPT13Bk-shot=02022.04 | 51.5 | |
| Branch-Train-MixModel Scale=560M, Training Strategy=Branch-Train-Mix, Evaluation Protocol=Few-shot (4-shot)2025.06 | 51.4 | |
| Branch-Train-MixModel Scale=1.7B, Training Strategy=Branch-Train-Mix, Evaluation Protocol=Few-shot (4-shot)2025.06 | 51.3 | |
| mGPT1.3Bk-shot=12022.04 | 51.3 | |
| BLOOMModel Scale=560M, Training Strategy=Base, Evaluation Protocol=Few-shot (4-shot)2025.06 | 51.1 | |
| mGPT13Bk-shot=12022.04 | 50.6 | |
| BLOOMModel Scale=1.7B, Training Strategy=Base, Evaluation Protocol=Few-shot (4-shot)2025.06 | 50.5 | |
| XGLM1.7Bk-shot=02022.04 | 50.3 | |
| XGLM7.5Bk-shot=02022.04 | 50.1 | |
| XGLM7.5Bk-shot=12022.04 | 46.4 | |
| XGLM1.7Bk-shot=12022.04 | 45.9 | |
| XGLM1.7Bk-shot=42022.04 | 45.9 | |
| XGLM7.5Bk-shot=42022.04 | 45.3 | |
| XGLM7.5Bk-shot=162022.04 | 44.9 | |
| XGLM1.7Bk-shot=162022.04 | 44.2 |