Natural Language Inference on GLUE (MNLI, QNLI, RTE), SciTail, and SNLI (test)
89.88MNLI AccθA fine-tune
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| θA fine-tuneEncoder=Source (t5-3b), B (Number of alignment batches)=All, Evaluation Protocol=Fine-tuning2026.02 | 89.88 | 95.78 | 87.73 | 92.45 | 90.45 | 91.26 | — | |
| θB fine-tuneEncoder=Target (t5-lg), B (Number of alignment batches)=All, Evaluation Protocol=Fine-tuning2026.02 | 86.34 | 92.14 | 81.23 | 91.12 | 88.75 | 87.92 | — | |
| THESEUSEncoder=Transported (θB + τA), B (Number of alignment batches)=100, Evaluation Protocol=Linear Probing2026.02 | 83.6 | 90.7 | 83.8 | 89 | 82.1 | 85.84 | 3.93 | |
| THESEUSEncoder=Transported (θB + τA), B (Number of alignment batches)=50, Evaluation Protocol=Linear Probing2026.02 | 81.81 | 89.1 | 80.9 | 88.3 | 80.3 | 84.08 | 7.3 | |
| θBEncoder=Target (t5-lg), B (Number of alignment batches)=100, Evaluation Protocol=Linear Probing2026.02 | 81.04 | 84.55 | 81.59 | 82.36 | 80.02 | 81.91 | 0 | |
| θAEncoder=Source (t5-3b), B (Number of alignment batches)=100, Evaluation Protocol=Linear Probing2026.02 | 80.14 | 90.59 | 79.42 | 82.06 | 80.78 | 82.6 | 0.69 | |
| THESEUSEncoder=Transported (θB + τA), B (Number of alignment batches)=20, Evaluation Protocol=Linear Probing2026.02 | 78.67 | 86.09 | 74 | 84.74 | 76.04 | 79.91 | 22.04 | |
| θBEncoder=Target (t5-lg), B (Number of alignment batches)=50, Evaluation Protocol=Linear Probing2026.02 | 78.58 | 81.79 | 76.17 | 75.15 | 72.2 | 76.78 | 0 | |
| θAEncoder=Source (t5-3b), B (Number of alignment batches)=50, Evaluation Protocol=Linear Probing2026.02 | 74.91 | 82.43 | 74.01 | 68.67 | 72.61 | 74.53 | -2.25 | |
| θAEncoder=Source (t5-3b), B (Number of alignment batches)=20, Evaluation Protocol=Linear Probing2026.02 | 66.84 | 80.85 | 71.84 | 69.4 | 45.06 | 66.8 | 8.39 | |
| θBEncoder=Target (t5-lg), B (Number of alignment batches)=20, Evaluation Protocol=Linear Probing2026.02 | 64.87 | 66.61 | 64.62 | 49.85 | 43.42 | 57.87 | 0 |