Natural Language Inference on MNLI
90.8Accuracy (matched)Vanilla
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| VanillaTraining=Standard Training, Model=RoBERTa2020.10 | 90.8 | 90.6 | — | |
| InfoBERTTraining=Adversarial Training, Model=RoBERTa2020.10 | 90.7 | 90.4 | — | |
| InfoBERTTraining=Standard Training, Model=RoBERTa2020.10 | 90.5 | 90.4 | — | |
| FreeLBTraining=Adversarial Training, Model=RoBERTa2020.10 | 90.1 | 90.3 | — | |
| InfoBERTTraining=Adversarial Training, Model=BERT2020.10 | 87.2 | 87.2 | — | |
| FreeLBTraining=Adversarial Training, Model=BERT2020.10 | 86.9 | 86.5 | — | |
| RLLR_MIXEDBackbone=Mistral 7B2024.05 | 86.8 | 88.8 | — | |
| VanillaTraining=Standard Training, Model=BERT2020.10 | 86.7 | 86.4 | — | |
| RLLRBackbone=Mistral 7B2024.05 | 86.6 | 88.9 | — | |
| InfoBERTTraining=Standard Training, Model=BERT2020.10 | 86.2 | 86 | — | |
| RLLRModel=LLaMA2, Model Size=13B2024.05 | 86.1 | 87.8 | — | |
| RLLRMIXEDModel=LLaMA2, Model Size=13B2024.05 | 86.1 | 87.5 | — | |
| SMP-SRemaining Weights=50%, New Params Per Task=θm, Trainable Params=85M, Knowledge Distillation=true2022.10 | 85.7 | 85.5 | — | |
| RLLRBackbone=Baichuan2 7B2024.05 | 85.7 | 85.8 | — | |
| RLHFModel=LLaMA2, Model Size=13B2024.05 | 85.6 | 87.2 | — | |
| RLLR_MIXEDBackbone=Baichuan2 7B2024.05 | 85.5 | 85.7 | — | |
| SFT w. rat.Backbone=Mistral 7B2024.05 | 85.4 | 87.6 | — | |
| RLHFBackbone=Mistral 7B2024.05 | 85.4 | 87.8 | — | |
| SFT w. rat.Model=LLaMA2, Model Size=13B2024.05 | 85.4 | 87.1 | — | |
| SMP-LRemaining Weights=50%, New Params Per Task=θm, Trainable Params=85M, Knowledge Distillation=true2022.10 | 85.3 | 85.6 | — | |
| RLLRBackbone=LLaMA2 7B2024.05 | 85.1 | 85.9 | — | |
| RLLR_MIXEDBackbone=LLaMA2 7B2024.05 | 85.1 | 85.9 | — | |
| RLLRModel=LLaMA2, Model Size=7B2024.05 | 85.1 | 85.9 | — | |
| RLLRMIXEDModel=LLaMA2, Model Size=7B2024.05 | 85.1 | 85.9 | — | |
| BERTBackbone=BERT-base, Evaluation Protocol=Fine-tuned2021.06 | 84.8 | 83.1 | — | |
| BERT + AlignerBackbone=BERT-base, Evaluation Protocol=Fine-tuned, Alignment Feature=Conditional alignment probability2021.06 | 84.8 | 83.5 | — | |
| SFT w. rat.Backbone=Baichuan2 7B2024.05 | 84.8 | 85 | — | |
| SFTBackbone=Mistral 7B2024.05 | 84.7 | 87.5 | — | |
| BERT baseRemaining Weights=100%, New Params Per Task=110M, Trainable Params=110M, Knowledge Distillation=false2022.10 | 84.5 | 84.9 | — | |
| RLHFBackbone=Baichuan2 7B2024.05 | 84.5 | 85.3 | — | |
| CAPRemaining Weights=50%, New Params Per Task=42.5M + θm, Trainable Params=170M, Knowledge Distillation=true2022.10 | 83.8 | 84.2 | — | |
| SMP-SRemaining Weights=10%, New Params Per Task=θm, Trainable Params=85M, Knowledge Distillation=true2022.10 | 83.7 | 83.6 | — | |
| RLHFBackbone=LLaMA2 7B2024.05 | 83.6 | 85 | — | |
| RLLRBackbone=ChatGLM3 6B2024.05 | 83.6 | 84.6 | — | |
| RLHFModel=LLaMA2, Model Size=7B2024.05 | 83.6 | 85 | — | |
| SFTBackbone=LLaMA2 7B2024.05 | 83.5 | 85.1 | — | |
| SFT w. rat.Backbone=LLaMA2 7B2024.05 | 83.5 | 85 | — | |
| RLLR_MIXEDBackbone=ChatGLM3 6B2024.05 | 83.5 | 84.6 | — | |
| SFTModel=LLaMA2, Model Size=7B2024.05 | 83.5 | 85.1 | — | |
| SFT w. rat.Model=LLaMA2, Model Size=7B2024.05 | 83.5 | 85 | — | |
| SMP-LRemaining Weights=10%, New Params Per Task=θm, Trainable Params=85M, Knowledge Distillation=true2022.10 | 83.1 | 83.1 | — | |
| SFTModel=LLaMA2, Model Size=13B2024.05 | 83.1 | 85.2 | — | |
| SFTBackbone=Baichuan2 7B2024.05 | 82.9 | 84.1 | — | |
| SFT w. rat.Backbone=ChatGLM3 6B2024.05 | 82.8 | 84.2 | — | |
| RLHFBackbone=ChatGLM3 6B2024.05 | 82.8 | 84.3 | — | |
| SMP-SRemaining Weights=10%, New Params Per Task=θm, Trainable Params=85M, Knowledge Distillation=false2022.10 | 82.5 | 82.3 | — | |
| MovementRemaining Weights=50%, New Params Per Task=42.5M + θm, Trainable Params=170M, Knowledge Distillation=true2022.10 | 82.5 | 82.9 | — | |
| SMP-LRemaining Weights=10%, New Params Per Task=θm, Trainable Params=85M, Knowledge Distillation=false2022.10 | 82 | 82.3 | — | |
| CAPRemaining Weights=10%, New Params Per Task=8.5M + θm, Trainable Params=170M, Knowledge Distillation=true2022.10 | 82 | 82.9 | — | |
| SMP-SRemaining Weights=3%, New Params Per Task=θm, Trainable Params=85M, Knowledge Distillation=true2022.10 | 81.8 | 82 | — | |
| SFTBackbone=ChatGLM3 6B2024.05 | 81.8 | 83.9 | — | |
| Soft-MovementRemaining Weights=10%, New Params Per Task=8.5M + θm, Trainable Params=170M, Knowledge Distillation=true2022.10 | 81.2 | 81.8 | — | |
| SMP-SRemaining Weights=3%, New Params Per Task=θm, Trainable Params=85M, Knowledge Distillation=false2022.10 | 80.9 | 81.1 | — | |
| SMP-LRemaining Weights=3%, New Params Per Task=θm, Trainable Params=85M, Knowledge Distillation=true2022.10 | 80.8 | 81.2 | — | |
| Soft-MovementRemaining Weights=10%, New Params Per Task=8.5M + θm, Trainable Params=170M, Knowledge Distillation=false2022.10 | 80.7 | 81.1 | — | |
| SMP-LRemaining Weights=3%, New Params Per Task=θm, Trainable Params=85M, Knowledge Distillation=false2022.10 | 80.6 | 81 | — | |
| MovementRemaining Weights=10%, New Params Per Task=8.5M + θm, Trainable Params=170M, Knowledge Distillation=true2022.10 | 80.1 | 80.4 | — | |
| CAPRemaining Weights=3%, New Params Per Task=2.6M + θm, Trainable Params=170M, Knowledge Distillation=true2022.10 | 80.1 | 81.3 | — | |
| Soft-MovementRemaining Weights=3%, New Params Per Task=2.6M + θm, Trainable Params=170M, Knowledge Distillation=true2022.10 | 79.5 | 80.1 | — | |
| MovementRemaining Weights=10%, New Params Per Task=8.5M + θm, Trainable Params=170M, Knowledge Distillation=false2022.10 | 79.3 | 79.5 | — | |
| Soft-MovementRemaining Weights=3%, New Params Per Task=2.6M + θm, Trainable Params=170M, Knowledge Distillation=false2022.10 | 79 | 79.6 | — | |
| L0-regularizationRemaining Weights=10%, New Params Per Task=8.5M + θm, Trainable Params=170M, Knowledge Distillation=true2022.10 | 78.7 | 79.7 | — | |
| MagnitudeRemaining Weights=10%, New Params Per Task=8.5M + θm, Trainable Params=85M, Knowledge Distillation=true2022.10 | 78.3 | 79.3 | — | |
| RLLRBackbone=Bloom 7B2024.05 | 77.9 | 81.3 | — | |
| RLLRModel=Bloom, Model Size=7B2024.05 | 77.9 | 81.3 | — | |
| RLLR_MIXEDBackbone=Bloom 7B2024.05 | 77.8 | 80.7 | — | |
| RLLRMIXEDModel=Bloom, Model Size=7B2024.05 | 77.8 | 80.7 | — | |
| RLHFBackbone=Bloom 7B2024.05 | 77 | 80 | — | |
| RLHFModel=Bloom, Model Size=7B2024.05 | 77 | 80 | — | |
| MovementRemaining Weights=3%, New Params Per Task=2.6M + θm, Trainable Params=170M, Knowledge Distillation=true2022.10 | 76.5 | 77.4 | — | |
| SFT w. rat.Backbone=Bloom 7B2024.05 | 76.5 | 80.7 | — | |
| SFT w. rat.Model=Bloom, Model Size=7B2024.05 | 76.5 | 80.7 | — | |
| MovementRemaining Weights=3%, New Params Per Task=2.6M + θm, Trainable Params=170M, Knowledge Distillation=false2022.10 | 76.1 | 76.7 | — | |
| SFTBackbone=Bloom 7B2024.05 | 75.8 | 78.5 | — | |
| SFTModel=Bloom, Model Size=7B2024.05 | 75.8 | 78.5 | — | |
| RLLRMIXEDModel=Bloom, Model Size=3B2024.05 | 74.7 | 76.4 | — | |
| RLLRModel=Bloom, Model Size=3B2024.05 | 74.6 | 76.7 | — | |
| RLHFModel=Bloom, Model Size=3B2024.05 | 74.2 | 75.6 | — | |
| SFT w. rat.Model=Bloom, Model Size=3B2024.05 | 73.4 | 75.1 | — | |
| SFTModel=Bloom, Model Size=3B2024.05 | 73.3 | 74.7 | — | |
| 32-shot tuningTraining=32-shot (FS)2023.05 | — | — | 41.7 | |
| AdamWRuntime (min)=8.12026.04 | — | — | 56.2 | |
| Base modelRuntime (min)=N/A2026.04 | — | — | 56.2 | |
| BLURRuntime (min)=8.62026.04 | — | — | 56.2 | |
| CENTRALISEDBackbone=RoBERTa-base, Setting=Balanced (i.i.d.)2025.01 | — | — | 86.0316 | |
| CHILD-TUNINGDTraining Dataset=MNLI2022.11 | — | — | 78.01 | |
| CHILD-TUNINGDTraining Dataset=SNLI2022.11 | — | — | 67.26 | |
| ColD-Fusion2022.12 | — | — | 87.14 | |
| Coordinate LowRank-LRBackbone=RoBERTa-large2026.03 | — | — | 57.2 | |
| Demonstration learningMode=Zero-shot with context2023.05 | — | — | 46.1 | |
| DPS DenseTraining Dataset=MNLI2022.11 | — | — | 77.59 | |
| DPS DenseTraining Dataset=SNLI2022.11 | — | — | 67.32 | |
| DPS MixTraining Dataset=MNLI2022.11 | — | — | 77.75 | |
| DPS MixTraining Dataset=SNLI2022.11 | — | — | 66.81 | |
| Evolved LLM Θ′Update Source=OpenOrca, Evaluation Protocol=Zero-shot2026.05 | — | — | 32.3 | |
| Evolved LLM Θ′Update Source=AlpacaGPT4, Evaluation Protocol=Zero-shot2026.05 | — | — | 31.6 | |
| Evolved LLM Θ′Update Source=OpenPlatypus, Evaluation Protocol=Zero-shot2026.05 | — | — | 31.9 | |
| FastGASAnnotation budget (|L|)=182024.06 | — | — | 44.53 | |
| FEDAVGBackbone=RoBERTa-base, Adapter=LoRA, Setting=Balanced (i.i.d.)2025.01 | — | — | 85.2471 | |
| FFABackbone=RoBERTa-base, Adapter=LoRA, Setting=Balanced (i.i.d.)2025.01 | — | — | 84.1773 |