Sentiment Analysis on IMDB (test)
97.42AccuracyDV-ngrams-cosine + NB-weighted BON
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| DV-ngrams-cosine + NB-weighted BONTraining split=Original 25K, Status=Incorrect previously reported results2022.05 | 97.42 | — | — | — | — | — | |
| LS-largeFine-tuning=true, Model Size=large, Pre-training=RoBERTa-large2021.07 | 96.8 | — | — | — | — | — | |
| SOTA2024.06 | 96.68 | — | — | — | — | — | |
| roberta-largeTraining Strategy=Combined extra+train2024.03 | 96.68 | — | — | — | — | — | |
| roberta-largeTraining Strategy=Baseline train2024.03 | 96.54 | — | — | — | — | — | |
| RoBERTa-largeFine-tuning=true, Model Size=large2021.07 | 96.5 | — | — | — | — | — | |
| Fine-tuning*Backbone=RoBERTa-large, Setting=Supervised2023.05 | 96.4 | — | — | — | — | — | |
| LS-baseFine-tuning=true, Model Size=base, Pre-training=RoBERTa-base2021.07 | 96 | — | — | — | — | — | |
| DV-ngrams-cosine with NB sub-sampling + RoBERTaTraining split=Suchin et al. 2020 (20K/5K), NB sub-sampling=true2022.05 | 95.94 | — | — | — | — | — | |
| DV-ngrams-cosine + RoBERTaTraining split=Suchin et al. 2020 (20K/5K)2022.05 | 95.92 | — | — | — | — | — | |
| RoBERTaTraining split=Suchin et al. 2020 (20K/5K)2022.05 | 95.79 | — | — | — | — | — | |
| Longformer-baseFine-tuning=true, Model Size=base2021.07 | 95.7 | — | — | — | — | — | |
| RoBERTa-baseFine-tuning=true, Model Size=base2021.07 | 95.3 | — | — | — | — | — | |
| XLNet-LargeParameter=340M, Training Set=Original (O)2021.06 | 95.3 | — | — | — | — | — | |
| roberta-baseTraining Strategy=Combined extra+train2024.03 | 95.23 | — | — | — | — | — | |
| SENTECONInterpretable?=Yes, M_theta=Fine-tuned, Base Lexicon=LIWC (L)2023.05 | 95.1 | — | — | — | — | — | |
| MPNetInterpretable?=No, M_theta=Fine-tuned2023.05 | 95.1 | — | — | — | — | — | |
| bert-largeTraining Strategy=Combined extra+train2024.03 | 95.03 | — | — | — | — | — | |
| SENTECON+Interpretable?=Yes, M_theta=Fine-tuned, Base Lexicon=LIWC (L)2023.05 | 95 | — | — | — | — | — | |
| SENTECON+Interpretable?=Yes, M_theta=Fine-tuned, Base Lexicon=Empath (E)2023.05 | 95 | — | — | — | — | — | |
| roberta-largeTraining Strategy=LlamBERT train&extra2024.03 | 94.98 | — | — | — | — | — | |
| XLNet-LargeParameter=340M, Training Set=Augmented (AC)2021.06 | 94.9 | — | — | — | — | — | |
| SENTECONInterpretable?=Yes, M_theta=Fine-tuned, Base Lexicon=Empath (E)2023.05 | 94.9 | — | — | — | — | — | |
| roberta-largeTraining Strategy=LlamBERT train2024.03 | 94.83 | — | — | — | — | — | |
| UniMC*Backbone=ALBERT-xxlarge, Labeled=true, Setting=Zero-shot2023.05 | 94.8 | — | — | — | — | — | |
| roberta-baseTraining Strategy=Baseline train2024.03 | 94.74 | — | — | — | — | — | |
| CLINEmode=fine-tuned2021.07 | 94.5 | — | — | — | — | — | |
| bert-largeTraining Strategy=Baseline train2024.03 | 94.29 | — | — | — | — | — | |
| roberta-baseTraining Strategy=LlamBERT train&extra2024.03 | 94.28 | — | — | — | — | — | |
| Virtual2017.08 | 94.1 | — | — | — | — | — | |
| oh-LSTM2017.08 | 94.1 | — | — | — | — | — | |
| ROBERTa-LargeParameter=355M, Training Set=Augmented (AC)2021.06 | 94.1 | — | — | — | — | — | |
| bert-largeTraining Strategy=LlamBERT train&extra2024.03 | 94.07 | — | — | — | — | — | |
| XLNet-LargeParameter=340M, Training Set=Combined (C)2021.06 | 93.9 | — | — | — | — | — | |
| TRNN2017.08 | 93.8 | — | — | — | — | — | |
| BERTBackbone=BERT-base-uncased2022.03 | 93.8 | — | — | — | 1 | — | |
| DV-ngrams-cosine + NB-weighted BONTraining split=Original 25K, Status=Re-evaluated in this paper2022.05 | 93.68 | — | — | — | — | — | |
| RoBERTamode=fine-tuned2021.07 | 93.6 | — | — | — | — | — | |
| ROBERTa-LargeParameter=355M, Training Set=Combined (C)2021.06 | 93.6 | — | — | — | — | — | |
| UniMC (Rerun)Backbone=ALBERT-xxlarge, Labeled=true, Setting=Zero-shot2023.05 | 93.6 | — | — | — | — | — | |
| roberta-baseTraining Strategy=LlamBERT train2024.03 | 93.53 | — | — | — | — | — | |
| bert-baseTraining Strategy=Combined extra+train2024.03 | 93.47 | — | — | — | — | — | |
| BERT-basemode=fine-tuning2019.10 | 93.46 | — | — | — | — | — | |
| ROBERTa-LargeParameter=355M, Training Set=Original (O)2021.06 | 93.4 | — | — | — | — | — | |
| SSTuning-ALBERTBackbone=ALBERT-xxlarge, Labeled=false, Setting=Zero-shot2023.05 | 93.4 | — | — | — | — | — | |
| DV-ngrams-cosine with NB sub-samplingTraining split=Suchin et al. 2020 (20K/5K), NB sub-sampling=true2022.05 | 93.36 | — | — | — | — | — | |
| bert-largeTraining Strategy=LlamBERT train2024.03 | 93.31 | — | — | — | — | — | |
| Parallelized LMUNumber of parameters (Millions)=342021.02 | 93.2 | — | — | — | — | — | |
| TR-BERTBackbone=BERT-base-uncased2022.03 | 93.2 | — | — | — | 2.9 | — | |
| DV-ngrams-cosineTraining split=Original 25K2022.05 | 93.13 | — | — | — | — | — | |
| SSTuning-largeBackbone=RoBERTa-large, Labeled=false, Setting=Zero-shot2023.05 | 93 | — | — | — | — | — | |
| Baseline BERTLayer Configuration={12, sf}2023.01 | 93 | — | — | — | — | — | |
| ContextFirstLayer Configuration=sfsf{10,s}:{10,f}2023.01 | 93 | — | — | — | — | — | |
| SparseQueriesLayer Configuration={4,sf}:{8,sf}2023.01 | 93 | — | — | — | — | — | |
| bmLSTM2017.08 | 92.9 | — | — | — | — | — | |
| DistilBERTBackbone=BERT-base-uncased2022.03 | 92.9 | — | — | — | 2 | — | |
| LSTMNumber of parameters (Millions)=752021.02 | 92.88 | — | — | — | — | — | |
| DistilBERTmode=fine-tuning2019.10 | 92.82 | — | — | — | — | — | |
| DistilBERTNumber of parameters (Millions)=662021.02 | 92.82 | — | — | — | — | — | |
| SA-LSTM2017.08 | 92.8 | — | — | — | — | — | |
| bert-baseTraining Strategy=LlamBERT train&extra2024.03 | 92.76 | — | — | — | — | — | |
| Ensemble of RNNs and NB-SVM2016.11 | 92.6 | — | — | — | — | — | |
| distilbert-baseTraining Strategy=Combined extra+train2024.03 | 92.53 | — | — | — | — | — | |
| bert-baseTraining Strategy=Baseline train2024.03 | 92.35 | — | — | — | — | — | |
| 2 layer sequential BoW CNNnumber of layers=22016.11 | 92.3 | — | — | — | — | — | |
| BERTmode=fine-tuned2021.07 | 92.2 | — | — | — | — | — | |
| PoWER-BERTBackbone=BERT-base-uncased2022.03 | 92.2 | — | — | — | 1.7 | — | |
| distilbert-baseTraining Strategy=LlamBERT train&extra2024.03 | 92.12 | — | — | — | — | — | |
| Funnel Transformer2023.01 | 92 | — | — | — | — | — | |
| SparseQueriesLayer Configuration={3,sf}:{9,sf}2023.01 | 92 | — | — | — | — | — | |
| BCN+Char+CoVeembeddings=GloVe, character n-gram, CoVe2017.08 | 91.8 | — | — | — | — | — | |
| WWM-BERT-LargeParameter=335M, Training Set=Augmented (AC)2021.06 | 91.8 | — | — | — | — | — | |
| AdapLeRBackbone=BERT-base-uncased2022.03 | 91.7 | — | — | — | 3.21 | — | |
| BSRBackbone=BERT-base, Memory regime=high memory, p=64, epsilon=82026.05 | 91.68 | — | — | — | — | — | |
| BandMFBackbone=BERT-base, Memory regime=high memory, p=64, epsilon=82026.05 | 91.65 | — | — | — | — | — | |
| ROBERTa-LargeParameter=355M, Training Set=Counterfactual (CF)2021.06 | 91.6 | — | — | — | — | — | |
| bert-baseTraining Strategy=LlamBERT train2024.03 | 91.58 | — | — | — | — | — | |
| BLTBackbone=BERT-base, Memory regime=high memory, epsilon=82026.05 | 91.48 | — | — | — | — | — | |
| BandInvMFBackbone=BERT-base, Memory regime=high memory, p=64, epsilon=82026.05 | 91.41 | — | — | — | — | — | |
| Densely-connected 4-layer QRNNnumber of layers=4, units per layer=256, filter width (k)=22016.11 | 91.4 | — | — | 150 | — | — | |
| NB-weighted BONTraining split=Original 25K2022.05 | 91.29 | — | — | — | — | — | |
| γ-BIFRBackbone=BERT-base, Memory regime=high memory, p=64, epsilon=82026.05 | 91.28 | — | — | — | — | — | |
| distilbert-baseTraining Strategy=Baseline train2024.03 | 91.23 | — | — | — | — | — | |
| NBSVM-bi2016.11 | 91.2 | — | — | — | — | — | |
| WWM-BERT-LargeParameter=335M, Training Set=Original (O)2021.06 | 91.2 | — | — | — | — | — | |
| RAdamGeneration=Gen 52026.04 | 91.2 | — | — | — | — | — | |
| BISRBackbone=BERT-base, Memory regime=high memory, p=64, epsilon=82026.05 | 91.17 | — | — | — | — | — | |
| Densely-connected 4-layer QRNNnumber of layers=4, units per layer=256, filter width (k)=42016.11 | 91.1 | — | — | 160 | — | — | |
| TE-MNLIBackbone=BART-large, Labeled=true, Setting=Zero-shot2023.05 | 91.1 | — | — | — | — | — | |
| Lookahead+AdamWGeneration=Gen 52026.04 | 91.1 | — | — | — | — | — | |
| DP-λCGDBackbone=BERT-base, Memory regime=high memory, p=2, epsilon=82026.05 | 91.05 | — | — | — | — | — | |
| WWM-BERT-LargeParameter=335M, Training Set=Combined (C)2021.06 | 91 | — | — | — | — | — | |
| SparseQueriesLayer Configuration={1,sf}:{11,sf}2023.01 | 91 | — | — | — | — | — | |
| SparseQueriesLayer Configuration={2,sf}:{10,sf}2023.01 | 91 | — | — | — | — | — | |
| BLTBackbone=BERT-base, Memory regime=high memory, epsilon=42026.05 | 91 | — | — | — | — | — | |
| BandMFBackbone=BERT-base, Memory regime=high memory, p=64, epsilon=42026.05 | 90.96 | — | — | — | — | — | |
| Densely-connected 4-layer LSTMnumber of layers=4, optimization=cuDNN optimized, units per layer=2562016.11 | 90.9 | — | — | 480 | — | — | |
| XLNet-LargeParameter=340M, Training Set=Counterfactual (CF)2021.06 | 90.8 | — | — | — | — | — | |
| AdamGeneration=Gen 32026.04 | 90.8 | — | — | — | — | — | |
| BandInvMFBackbone=BERT-base, Memory regime=high memory, p=64, epsilon=42026.05 | 90.8 | — | — | — | — | — |