Paraphrase Identification on QQP
91.7AccuracySMP-S
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SMP-SRemaining Weights=50%, Knowledge Distillation=true2022.10 | 91.7 | 88.8 | |
| CAPRemaining Weights=50%, Knowledge Distillation=true2022.10 | 91.6 | 88.6 | |
| SMP-LRemaining Weights=50%, Knowledge Distillation=true2022.10 | 91.6 | 88.7 | |
| BERT baseRemaining Weights=100%, New Params Per Task=110M, Trainable Params=110M, Knowledge Distillation=false2022.10 | 91.4 | 88.4 | |
| ALIGNModel Variant=large, Model Parameters=355M2023.07 | 91.3 | — | |
| MUPPET2022.12 | 91.25 | — | |
| ColD-Fusion2022.12 | 91.22 | — | |
| DenseSparsity (%)=0, Backbone=BERT-base2023.05 | 91.2 | 88.1 | |
| MovementRemaining Weights=50%, Knowledge Distillation=true2022.10 | 91 | 87.8 | |
| SMP-LRemaining Weights=10%, Knowledge Distillation=true2022.10 | 91 | 87.9 | |
| SMP-SRemaining Weights=10%, Knowledge Distillation=true2022.10 | 91 | 87.9 | |
| PDPSparsity (%)=90, Backbone=BERT-base2023.05 | 91 | 88 | |
| BERT + AlignerBackbone=BERT-base, Evaluation Protocol=Fine-tuned, Alignment Feature=Conditional alignment probability2021.06 | 90.9 | — | |
| POFASparsity (%)=90, Backbone=BERT-base2023.05 | 90.9 | 87.7 | |
| Multitask2022.12 | 90.89 | — | |
| BERTBackbone=BERT-base, Evaluation Protocol=Fine-tuned2021.06 | 90.8 | — | |
| SMP-LRemaining Weights=10%, Knowledge Distillation=false2022.10 | 90.8 | 87.7 | |
| SMP-SRemaining Weights=10%, Knowledge Distillation=false2022.10 | 90.8 | 87.6 | |
| Finetune2022.12 | 90.72 | — | |
| CAPRemaining Weights=10%, Knowledge Distillation=true2022.10 | 90.7 | 87.4 | |
| Tay et al. (2022)Mode=Fully-supervised, Model Parameters=4x of UD (~44B)2022.11 | 90.6 | — | |
| Soft-MovementRemaining Weights=10%, Knowledge Distillation=false2022.10 | 90.5 | 87.1 | |
| SMP-SRemaining Weights=3%, Knowledge Distillation=true2022.10 | 90.5 | 87.4 | |
| UD+-XXLMode=Fully-supervised, Backbone=T5-XXL encoder, Model Parameters=11B2022.11 | 90.44 | — | |
| SMP-SRemaining Weights=3%, Knowledge Distillation=false2022.10 | 90.3 | 87.1 | |
| SMP-LRemaining Weights=3%, Knowledge Distillation=false2022.10 | 90.2 | 87 | |
| Soft-MovementRemaining Weights=10%, Knowledge Distillation=true2022.10 | 90.2 | 86.8 | |
| CAPRemaining Weights=3%, Knowledge Distillation=true2022.10 | 90.2 | 86.7 | |
| MVPSparsity (%)=90, Backbone=BERT-base2023.05 | 90.2 | 86.8 | |
| SMP-LRemaining Weights=3%, Knowledge Distillation=true2022.10 | 90.1 | 87 | |
| ALIGNModel Variant=base, Model Parameters=125M2023.07 | 90.1 | — | |
| OptGSparsity (%)=90, Backbone=BERT-base2023.05 | 89.8 | 86.2 | |
| MovementRemaining Weights=10%, Knowledge Distillation=true2022.10 | 89.7 | 86.2 | |
| Soft-MovementRemaining Weights=3%, Knowledge Distillation=false2022.10 | 89.3 | 85.6 | |
| MovementRemaining Weights=10%, Knowledge Distillation=false2022.10 | 89.1 | 85.5 | |
| Soft-MovementRemaining Weights=3%, Knowledge Distillation=true2022.10 | 89.1 | 85.5 | |
| RLLR_MIXEDBackbone=Mistral 7B2024.05 | 88.7 | — | |
| RLHFBackbone=Mistral 7B2024.05 | 88.3 | — | |
| RLLRBackbone=Mistral 7B2024.05 | 88.3 | — | |
| RLLRBackbone=LLaMA2 7B2024.05 | 88.2 | — | |
| L0-regularizationRemaining Weights=10%, Knowledge Distillation=true2022.10 | 88.1 | 82.8 | |
| RLLR_MIXEDBackbone=LLaMA2 7B2024.05 | 88 | — | |
| SFT w. rat.Backbone=LLaMA2 7B2024.05 | 87.9 | — | |
| RLHFBackbone=LLaMA2 7B2024.05 | 87.6 | — | |
| RLLR_MIXEDBackbone=Baichuan2 7B2024.05 | 87.5 | — | |
| FLAN-T5Model Variant=xlarge, Model Parameters=3B2023.07 | 87.4 | — | |
| RLLRBackbone=Baichuan2 7B2024.05 | 87.4 | — | |
| SFT w. rat.Backbone=Baichuan2 7B2024.05 | 87 | — | |
| FLAN-T5Model Variant=large, Model Parameters=780M2023.07 | 86.8 | — | |
| RLHFBackbone=Baichuan2 7B2024.05 | 86.5 | — | |
| MovementRemaining Weights=3%, Knowledge Distillation=true2022.10 | 86.1 | 81.5 | |
| SFT w. rat.Backbone=Mistral 7B2024.05 | 86.1 | — | |
| SFTBackbone=Mistral 7B2024.05 | 85.9 | — | |
| RLLR_MIXEDBackbone=ChatGLM3 6B2024.05 | 85.7 | — | |
| MovementRemaining Weights=3%, Knowledge Distillation=false2022.10 | 85.6 | 81 | |
| SFTBackbone=LLaMA2 7B2024.05 | 85.5 | — | |
| RLLRBackbone=ChatGLM3 6B2024.05 | 85.5 | — | |
| SFTBackbone=ChatGLM3 6B2024.05 | 85 | — | |
| RLHFBackbone=ChatGLM3 6B2024.05 | 85 | — | |
| SFT w. rat.Backbone=ChatGLM3 6B2024.05 | 84.9 | — | |
| SFTBackbone=Baichuan2 7B2024.05 | 84.9 | — | |
| RLLRBackbone=Bloom 7B2024.05 | 84.3 | — | |
| RLLR_MIXEDBackbone=Bloom 7B2024.05 | 84.3 | — | |
| RLHFBackbone=Bloom 7B2024.05 | 84 | — | |
| SFT w. rat.Backbone=Bloom 7B2024.05 | 83.6 | — | |
| SFTBackbone=Bloom 7B2024.05 | 83.4 | — | |
| GENICLBackbone=LLaMA-7B, Evaluation Setting=Few-shot (ICL)2025.05 | 82 | — | |
| SBERTBackbone=LLaMA-7B, Evaluation Setting=Few-shot (ICL)2025.05 | 81.7 | — | |
| EPRBackbone=LLaMA-7B, Evaluation Setting=Few-shot (ICL)2025.05 | 81.7 | — | |
| LLM-RBackbone=LLaMA-7B, Evaluation Setting=Few-shot (ICL)2025.05 | 80.9 | — | |
| BM25Backbone=LLaMA-7B, Evaluation Setting=Few-shot (ICL)2025.05 | 80.3 | — | |
| MagnitudeRemaining Weights=10%, Knowledge Distillation=true2022.10 | 79.8 | 75.9 | |
| fs-X-ICL (ChatGPT)Model=Llama2 70B2023.11 | 77.6 | — | |
| fs-X-ICL (ChatGPT)Model=Zephyr 7B2023.11 | 77.3 | — | |
| E5baseBackbone=LLaMA-7B, Evaluation Setting=Few-shot (ICL)2025.05 | 77.3 | — | |
| CBDSBackbone=LLaMA-7B, Evaluation Setting=Few-shot (ICL)2025.05 | 64.2 | — | |
| Zero-shotBackbone=LLaMA-7B, Evaluation Setting=Zero-shot2025.05 | 57.9 | — | |
| RandomBackbone=LLaMA-7B, Evaluation Setting=Few-shot (ICL)2025.05 | 54 | — | |
| BERTbackbone=BERT2019.11 | — | 91.3 | |
| BERT+DLbackbone=BERT, loss=Dice Loss2019.11 | — | 91.92 | |
| BERT+DSCbackbone=BERT, loss=DSC2019.11 | — | 92.11 | |
| BERT+FLbackbone=BERT, loss=Focal Loss2019.11 | — | 91.86 | |
| XLNetbackbone=XLNet2019.11 | — | 91.8 | |
| XLNet+DLbackbone=XLNet, loss=Dice Loss2019.11 | — | 92.39 | |
| XLNet+DSCbackbone=XLNet, loss=DSC2019.11 | — | 92.6 | |
| XLNet+FLbackbone=XLNet, loss=Focal Loss2019.11 | — | 92.31 |