Sentiment Analysis on SST-2 (test)
97.1AccuracyPrevious SOTA
Evaluation Results
| Method | Links | |
|---|---|---|
| Previous SOTA2020.05 | 97.1 | |
| RoBERTa + SKEPModel Size=large2020.05 | 97 | |
| RoBERTa + SKEPModel Size=base2020.05 | 96.7 | |
| LM-Cocktail2Fine-tune on=SST22023.11 | 96.56 | |
| LM-Cocktail10Fine-tune on=SST22023.11 | 96.56 | |
| RoBERTaModel Size=large2020.05 | 96.5 | |
| CLS-based FT + TAPTBackbone=RoBERTa-LARGE, Supervision Setting=Full, Evaluation Protocol=CLS-based Fine-tuning, Continued Pre-training=TAPT2023.05 | 96 | |
| PGINum. of Samples=1821, Cost (USD)=7.33, Time (Mins)=122022.12 | 95.77 | |
| Non-AlignedHarmful ratio (p)=clean, Sample number (n)=10002024.02 | 95.6 | |
| Fine-tunedFine-tune on=SST22023.11 | 95.53 | |
| Prompt-based FT (hard) + PCPBackbone=RoBERTa-LARGE, Supervision Setting=Full, Evaluation Protocol=Prompt-based Fine-tuning (hard), Continued Pre-training=PCP2023.05 | 95.5 | |
| Prompt-based FT (hard)Backbone=RoBERTa-LARGE, Supervision Setting=Full, Evaluation Protocol=Prompt-based Fine-tuning (hard), Continued Pre-training=None2023.05 | 95.2 | |
| CLS-based FTBackbone=RoBERTa-LARGE, Supervision Setting=Full, Evaluation Protocol=CLS-based Fine-tuning, Continued Pre-training=None2023.05 | 95.1 | |
| Fine-tuning (full)Backbone=RoBERTa-large, Training Samples (K)=full, Evaluation Protocol=Full fine-tuning2021.08 | 95 | |
| VaccineHarmful ratio (p)=0.2, Sample number (n)=10002024.02 | 95 | |
| UnpatchedModel=Llama2026.06 | 95 | |
| RoBERTaModel Size=base2020.05 | 94.9 | |
| SFTHarmful ratio (p)=0.05, Sample number (n)=10002024.02 | 94.8 | |
| VlguardHarmful ratio (p)=clean, Sample number (n)=10002024.02 | 94.8 | |
| VlguardHarmful ratio (p)=0.01, Sample number (n)=10002024.02 | 94.8 | |
| Extra AnnotationEvaluation Protocol=ICL, Number of shots (K)=Increased2023.07 | 94.7 | |
| Non-AlignedHarmful ratio (p)=0.01, Sample number (n)=10002024.02 | 94.6 | |
| Non-AlignedHarmful ratio (p)=0.1, Sample number (n)=10002024.02 | 94.6 | |
| VlguardHarmful ratio (p)=0.05, Sample number (n)=10002024.02 | 94.6 | |
| VlguardHarmful ratio (p)=0.1, Sample number (n)=10002024.02 | 94.6 | |
| VlguardHarmful ratio (p)=0.2, Sample number (n)=10002024.02 | 94.6 | |
| GENICLBackbone=Qwen2.5-3B2025.05 | 94.6 | |
| CoTAMEvaluation Protocol=ICL, Number of shots (K)=32023.07 | 94.5 | |
| GENICLBackbone=LLaMA-3.2-3B2025.05 | 94.5 | |
| Non-AlignedHarmful ratio (p)=0.2, Sample number (n)=10002024.02 | 94.4 | |
| SFTHarmful ratio (p)=0.01, Sample number (n)=10002024.02 | 94.4 | |
| SFTHarmful ratio (p)=0.1, Sample number (n)=10002024.02 | 94.4 | |
| Prompt-based FT (soft) + PCPBackbone=RoBERTa-LARGE, Supervision Setting=Full, Evaluation Protocol=Prompt-based Fine-tuning (soft), Continued Pre-training=PCP2023.05 | 94.3 | |
| FlipDA++Evaluation Protocol=ICL, Number of shots (K)=32023.07 | 94.3 | |
| LLM-RBackbone=Qwen2.5-3B2025.05 | 94.3 | |
| LLM Pseudo LabelEvaluation Protocol=ICL, Number of shots (K)=32023.07 | 94.2 | |
| SFTHarmful ratio (p)=clean, Sample number (n)=10002024.02 | 94.2 | |
| SFTHarmful ratio (p)=0.2, Sample number (n)=10002024.02 | 94.2 | |
| BaseEvaluation Protocol=ICL, Number of shots (K)=32023.07 | 94 | |
| COTDAEvaluation Protocol=ICL, Number of shots (K)=32023.07 | 94 | |
| Non-AlignedHarmful ratio (p)=0.05, Sample number (n)=10002024.02 | 94 | |
| GENICLBackbone=Vicuna-13B2025.05 | 94 | |
| FTModel=Llama2026.06 | 94 | |
| FT_l2Model=Llama2026.06 | 94 | |
| FT_linfModel=Llama2026.06 | 94 | |
| PatcherModel=Llama2026.06 | 94 | |
| Prompt-based FT (soft) + PCPBackbone=RoBERTa-LARGE, Supervision Setting=Semi-supervised (16-shot), Evaluation Protocol=Prompt-based Fine-tuning (soft), Continued Pre-training=PCP2023.05 | 93.9 | |
| KENTrainable params=80M2024.02 | 93.8 | |
| VaccineHarmful ratio (p)=0.1, Sample number (n)=10002024.02 | 93.8 | |
| LLM-RBackbone=Vicuna-13B2025.05 | 93.8 | |
| Prompt-based FT (hard) + PCPBackbone=RoBERTa-LARGE, Supervision Setting=Semi-supervised (16-shot), Evaluation Protocol=Prompt-based Fine-tuning (hard), Continued Pre-training=PCP2023.05 | 93.6 | |
| Human LabeledNum. of Samples=67349, Cost (USD)=4800 - 6700, Time (Mins)=227402022.12 | 93.52 | |
| DARTBackbone=RoBERTa-large, Training Samples (K)=16, Evaluation Protocol=Few-shot fine-tuning2021.08 | 93.5 | |
| Prompt-based FT (hard) + TAPTBackbone=RoBERTa-LARGE, Supervision Setting=Full, Evaluation Protocol=Prompt-based Fine-tuning (hard), Continued Pre-training=TAPT2023.05 | 93.5 | |
| Bert-baseTrainable params=109M2024.02 | 93.37 | |
| HybridTrainable params=94M2024.02 | 93.23 | |
| VaccineHarmful ratio (p)=0.05, Sample number (n)=10002024.02 | 93 | |
| LLM-RBackbone=LLaMA-3.2-3B2025.05 | 93 | |
| FinePruningModel=Llama2026.06 | 93 | |
| BAERASERModel=Llama2026.06 | 93 | |
| MudjackingModel=Llama2026.06 | 93 | |
| SPPModel=Llama2026.06 | 93 | |
| MENDModel=Llama2026.06 | 93 | |
| ROMEModel=Llama2026.06 | 93 | |
| KENTrainable params=63M2024.02 | 92.9 | |
| GENICLBackbone=GPT-Neo 2.7B2025.05 | 92.7 | |
| VaccineHarmful ratio (p)=clean, Sample number (n)=10002024.02 | 92.6 | |
| VaccineHarmful ratio (p)=0.01, Sample number (n)=10002024.02 | 92.6 | |
| LLM-RBackbone=GPT-Neo 2.7B2025.05 | 92.6 | |
| LKD+Init.Evaluation Protocol=Adapted Prompting (AP)2023.05 | 92.5 | |
| LKD+Init.Evaluation Protocol=Fine-tuning (FT)2023.05 | 92.5 | |
| LM-BFFBackbone=RoBERTa-large, Training Samples (K)=16, Evaluation Protocol=Few-shot fine-tuning2021.08 | 92.3 | |
| P-TuningBackbone=RoBERTa-large, Training Samples (K)=16, Evaluation Protocol=Few-shot fine-tuning2021.08 | 92.2 | |
| HybridNTTrainable params=94M2024.02 | 92.2 | |
| LKDEvaluation Protocol=Fine-tuning (FT)2023.05 | 92 | |
| OneShotModel=Llama2026.06 | 92 | |
| HybridTrainable params=66M2024.02 | 91.97 | |
| ERMTraining Source=SST22023.05 | 91.85 | |
| ZoVHModel=OPT-13B, N=1, Forward pass budget=10,0002026.05 | 91.55 | |
| RETROPROMPTShot=16-shot, Source Domain=MR2022.05 | 91.4 | |
| RETROPROMPTshot=16-shot, source_dataset=MR, domain_role=Target Domain2025.12 | 91.4 | |
| MeZOModel=OPT-13B, Forward pass budget=10,0002026.05 | 91.13 | |
| ZoVHModel=OPT-13B, N=2, Forward pass budget=10,0002026.05 | 90.9 | |
| Gordon et al. (2020)Trainable params=66M2024.02 | 90.8 | |
| HybridNTTrainable params=66M2024.02 | 90.71 | |
| ZoVHModel=OPT-1.3B, N=2, Forward pass budget=10,0002026.05 | 90.63 | |
| Q-DiversityTraining Source=SST22023.05 | 90.62 | |
| LKDEvaluation Protocol=Adapted Prompting (AP)2023.05 | 90.6 | |
| No ExampleEvaluation Protocol=ICL, Number of shots (K)=02023.07 | 90.5 | |
| MeZOModel=OPT-1.3B, Forward pass budget=10,0002026.05 | 90.41 | |
| Sajjad et al. (2020)Trainable params=66M2024.02 | 90.3 | |
| ZoVHModel=OPT-1.3B, N=1, Forward pass budget=10,0002026.05 | 90.29 | |
| Coherence Boosting (GPT-3 175B)alpha=-0.542021.10 | 89.84 | |
| LM-BFF (D-demo)Shot=16-shot, Source Domain=MR2022.05 | 89.3 | |
| LM-BFF (D-demo)shot=16-shot, source_dataset=MR, domain_role=Target Domain2025.12 | 89.3 | |
| PGDANum. of Samples=6000, Cost (USD)=22.63, Time (Mins)=27†2022.12 | 89.29 | |
| LM-BFF (man)Shot=16-shot, Source Domain=MR2022.05 | 88.9 | |
| LM-BFF (man)shot=16-shot, source_dataset=MR, domain_role=Target Domain2025.12 | 88.9 | |
| TANDEMModel Size=500M, Base Model=Qwen-2, Training Steps=2000, Batch Size=32, Context Length=512, K=20, E=102026.06 | 88.7 | |
| EWCHarmful ratio (p)=clean, Sample number (n)=10002024.02 | 88.6 |