Sentiment Classification on IMDB
95.79AccuracySentriLlama 3.2 (3B) Instruct
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| SentriLlama 3.2 (3B) InstructAdded Params (M)=0.003–1.2, Extra LM call=true2026.01 | 95.79 | — | — | |
| DeBERTa V3 LargeAdded Params (M)=418, Extra LM call=true2026.01 | 95.34 | — | — | |
| Multi-head self-attnAdded Params (M)=35–35.5, Extra LM call=false2026.01 | 95.15 | — | — | |
| Scoring attentionAdded Params (M)=0.10–0.11, Extra LM call=false2026.01 | 95.05 | — | — | |
| StandardModel=RoBERTa-base2021.12 | 94.9 | — | — | |
| RoBERTa LargeAdded Params (M)=355, Extra LM call=true2026.01 | 94.3 | — | — | |
| ROBERTa-CLTraining supervision type=clean labels, Method categorization=Fully-supervised2020.10 | 94.26 | — | — | |
| Direct poolingAdded Params (M)=0.003–0.018, Extra LM call=false2026.01 | 94.07 | — | — | |
| BERT#Params (M)=1102026.06 | 93.8 | — | 93.8 | |
| BERT-Base2025.11 | 93.2 | — | — | |
| StandardModel=BERT-base-uncased2021.12 | 93.1 | — | — | |
| RoBERTa-BiLSTM#Params (M)=>1252026.06 | 92.36 | — | 92.35 | |
| Llama 3.2 (3B) Chain-of-ThoughtAdded Params (M)=0, Extra LM call=true2026.01 | 91.54 | — | — | |
| SemImagebackbone=ResNet-182025.11 | 91.5 | — | — | |
| GPT-3#Params (M)=1750002026.06 | 90.76 | — | 91.19 | |
| COSINEMethod categorization=Framework2020.10 | 90.54 | — | — | |
| ANYSIMLITE#Params (M)=1.12026.06 | 88.68 | — | 88.68 | |
| HAN2025.11 | 88.1 | — | — | |
| FreeLBMethod categorization=Baseline2020.10 | 88.04 | — | — | |
| Non-DPModel=LogReg2026.05 | 88 | — | — | |
| TextCNN2025.11 | 87 | — | — | |
| SMARTMethod categorization=Baseline2020.10 | 86.98 | — | — | |
| MixupMethod categorization=Baseline2020.10 | 86.92 | — | — | |
| Self-ensembleMethod categorization=Baseline2020.10 | 86.72 | — | — | |
| MULI (logits, Llama-3.2-3B)Added Params (M)=0.13–0.77, Extra LM call=false2026.01 | 86.5 | — | — | |
| Syn. KFACModel=LogReg, epsilon=8, Mode=Synthetic noise2026.05 | 86 | — | — | |
| Pub. KFACModel=LogReg, epsilon=8, Mode=Public data preconditioner2026.05 | 86 | — | — | |
| Syn. KFACModel=LogReg, epsilon=2.8, Mode=Synthetic noise2026.05 | 85.9 | — | — | |
| Pub. KFACModel=LogReg, epsilon=2.8, Mode=Public data preconditioner2026.05 | 85.8 | — | — | |
| Syn. KFACModel=LogReg, epsilon=1.5, Mode=Synthetic noise2026.05 | 85.5 | — | — | |
| Pub. KFACModel=LogReg, epsilon=1.5, Mode=Public data preconditioner2026.05 | 85.5 | — | — | |
| Syn. KFACModel=LogReg, epsilon=1, Mode=Synthetic noise2026.05 | 85.1 | — | — | |
| Pub. KFACModel=LogReg, epsilon=1, Mode=Public data preconditioner2026.05 | 85.1 | — | — | |
| USTMethod categorization=Baseline2020.10 | 84.56 | — | — | |
| RIFTModel=RoBERTa-base2021.12 | 84.2 | — | — | |
| InitMethod categorization=Framework2020.10 | 83.58 | — | — | |
| Syn. KFACModel=LogReg, epsilon=0.5, Mode=Synthetic noise2026.05 | 83.5 | — | — | |
| Pub. KFACModel=LogReg, epsilon=0.5, Mode=Public data preconditioner2026.05 | 83.5 | — | — | |
| DP-SGDModel=LogReg, epsilon=82026.05 | 83.2 | — | — | |
| DP-SGDModel=LogReg, epsilon=1.52026.05 | 83.1 | — | — | |
| DP-SGDModel=LogReg, epsilon=2.82026.05 | 83.1 | — | — | |
| DenoiseMethod categorization=Baseline2020.10 | 82.9 | — | — | |
| DP-SGDModel=LogReg, epsilon=12026.05 | 82.9 | — | — | |
| DP-SGDModel=LogReg, epsilon=0.52026.05 | 82.1 | — | — | |
| Adv-PTWDModel=RoBERTa-base2021.12 | 80.7 | — | — | |
| Adv-BaseModel=RoBERTa-base2021.12 | 80.1 | — | — | |
| Adv-MixoutModel=RoBERTa-base2021.12 | 79 | — | — | |
| RIFTModel=BERT-base-uncased2021.12 | 78.3 | — | — | |
| Adv-MixoutModel=BERT-base-uncased2021.12 | 77.8 | — | — | |
| Llama 3.2 (3B) Zero-shotAdded Params (M)=0, Extra LM call=true2026.01 | 77.59 | — | — | |
| WeSTClassMethod categorization=Baseline2020.10 | 77.4 | — | — | |
| Adv-PTWDModel=BERT-base-uncased2021.12 | 76.6 | — | — | |
| Llama 3.2 (3B) Few-shotAdded Params (M)=0, Extra LM call=true2026.01 | 76.06 | — | — | |
| Adv-BaseModel=BERT-base-uncased2021.12 | 74.6 | — | — | |
| SnorkelMethod categorization=Baseline2020.10 | 73.22 | — | — | |
| ROBERTa-WL+Training supervision type=weak labels, Method categorization=Baseline2020.10 | 72.6 | — | — | |
| ExMatchMethod categorization=Baseline2020.10 | 71.28 | — | — | |
| LRUDepth (r)=6, State dimension (d)=322026.05 | 67.03 | — | — | |
| LRUDepth (r)=3, State dimension (d)=322026.05 | 66.84 | — | — | |
| LRUDepth (r)=1, State dimension (d)=322026.05 | 66.54 | — | — | |
| MINGRUDepth (r)=3, State dimension (d)=322026.05 | 66.53 | — | — | |
| MINGRUDepth (r)=6, State dimension (d)=322026.05 | 66.21 | — | — | |
| MINGRUDepth (r)=1, State dimension (d)=322026.05 | 65.88 | — | — | |
| CMRUDepth (r)=3, State dimension (d)=322026.05 | 65.18 | — | — | |
| CMRUDepth (r)=6, State dimension (d)=322026.05 | 65.04 | — | — | |
| αCMRUDepth (r)=1, State dimension (d)=322026.05 | 64.9 | — | — | |
| CMRUDepth (r)=1, State dimension (d)=322026.05 | 64.84 | — | — | |
| αCMRUDepth (r)=3, State dimension (d)=322026.05 | 64.62 | — | — | |
| αCMRUDepth (r)=6, State dimension (d)=322026.05 | 64.13 | — | — | |
| ImplyLossMethod categorization=Baseline2020.10 | 63.85 | — | — | |
| no-stitchstitching_type=none, decoder=SVM (linear kernel)2023.11 | 61 | — | — | |
| affinestitching_type=latent translation (affine), zero_shot=true, decoder=SVM (linear kernel)2023.11 | 59 | — | — | |
| orthostitching_type=latent translation (ortho), zero_shot=true, decoder=SVM (linear kernel)2023.11 | 59 | — | — | |
| linearstitching_type=latent translation (linear), zero_shot=true, decoder=SVM (linear kernel)2023.11 | 57 | — | — | |
| 1-orthostitching_type=latent translation (1-ortho), zero_shot=true, decoder=SVM (linear kernel)2023.11 | 56 | — | — | |
| relativestitching_type=relative, decoder=SVM (linear kernel)2023.11 | 51 | — | — | |
| absolutestitching_type=absolute (probe), decoder=SVM (linear kernel)2023.11 | 50 | — | — | |
| Bert LargeParam=355M, Data=16G, FLOPS=9.07E192023.05 | — | — | 94.76 | |
| Bert-BaseParam=109M, Data=16G, FLOPS=2.79E192023.05 | — | — | 93.77 | |
| BERT-Large2020.06 | — | 4.51 | — | |
| bow-CNNvocabulary=30K most frequent words2014.12 | — | 8.66 | — | |
| FLOPBackbone=BLOOM1B7, Trainable params=1.1B2024.02 | — | — | 80.9 | |
| FLOPBackbone=BLOOM560k, Trainable params=408M2024.02 | — | — | 72.1 | |
| FLOPBackbone=DeBERTa, Trainable params=88M2024.02 | — | — | 81.1 | |
| FLOPBackbone=Bert, Trainable params=66M2024.02 | — | — | 80.5 | |
| FLOPBackbone=Ernie, Trainable params=67M2024.02 | — | — | 81.1 | |
| FLOPBackbone=DistilBERT, Trainable params=45M2024.02 | — | — | 81.2 | |
| FLOPBackbone=Electra, Trainable params=28M2024.02 | — | — | 81.2 | |
| Funnel-Transformer (B10-10-10H1024)Architecture=B10-10-10H10242020.06 | — | 3.36 | — | |
| Funnel-Transformer (B4-4-4H768)Architecture=B4-4-4H7682020.06 | — | 4.12 | — | |
| Funnel-Transformer (B6-3x2-3x2H768)Architecture=B6-3x2-3x2H7682020.06 | — | 3.82 | — | |
| Funnel-Transformer (B6-6-6H768)Architecture=B6-6-6H7682020.06 | — | 3.72 | — | |
| Funnel-Transformer (B8-8-8H1024)Architecture=B8-8-8H10242020.06 | — | 3.42 | — | |
| ISS(Large-scale)Param=109M, Data=0.72G, FLOPS=8.30E182023.05 | — | — | 94.57 | |
| ISS(Medium-scale)Param=109M, Data=0.18G, FLOPS=4.15E182023.05 | — | — | 93.61 | |
| ISS(Small-scale)Param=109M, Data=0.18G, FLOPS=1.82E182023.05 | — | — | 93.25 | |
| KENBackbone=BLOOM1B7, Trainable params=531M2024.02 | — | — | 84.2 | |
| KENBackbone=BLOOM560k, Trainable params=404M2024.02 | — | — | 81.3 | |
| KENBackbone=DeBERTa, Trainable params=84M2024.02 | — | — | 82.5 | |
| KENBackbone=Bert, Trainable params=57M2024.02 | — | — | 84.9 |