Sentiment Analysis on SST-2 (Accuracy)
97.48AccuracyUD+-XXL
Evaluation Results
| Method | Links | |
|---|---|---|
| UD+-XXLMode=Fully-supervised, Backbone=T5-XXL encoder, Model Parameters=11B2022.11 | 97.48 | |
| Tay et al. (2022)Mode=Fully-supervised, Model Parameters=4x of UD (~44B)2022.11 | 97.3 | |
| PRewrite-SBackbone=PaLM2-S2026.03 | 96.6 | |
| PRewrite-IBackbone=PaLM2-S2026.03 | 96.5 | |
| PRewrite (Initial Prompt)Backbone=PaLM2-S2026.03 | 96.3 | |
| FADS-ICLBackbone=Llama-1 13B, Shots=128-shots2024.05 | 95.7 | |
| ColD-Fusion2022.12 | 95.16 | |
| GENICLBackbone=LLaMA-7B, Evaluation Setting=Few-shot (ICL)2025.05 | 95 | |
| ICLBackbone=Llama-1 13B, Shots=128-shots2024.05 | 94.9 | |
| ICLBackbone=Llama-2 13B, Shots=128-shots2024.05 | 94.9 | |
| kNN-promptingBackbone=Llama-2 13B, Shots=128-shots2024.05 | 94.6 | |
| kNN-promptBackbone=Llama-2 7B, Shots=128-shots2024.05 | 94.5 | |
| Multitask2022.12 | 94.27 | |
| ICLBackbone=Llama-1 30B, Shots=128-shots2024.05 | 94 | |
| kNN-promptBackbone=Llama-1 13B, Shots=128-shots2024.05 | 93.9 | |
| kNN-promptBackbone=Llama-1 30B, Shots=128-shots2024.05 | 93.9 | |
| Finetune2022.12 | 93.85 | |
| ICRModel=Llama3.1-70B2025.09 | 93.8 | |
| ICLBackbone=Llama-2 7B, Shots=128-shots2024.05 | 93.6 | |
| kNN-promptingBackbone=Llama-1 13B, Shots=128-shots2024.05 | 93.5 | |
| LLM-RBackbone=LLaMA-7B, Evaluation Setting=Few-shot (ICL)2025.05 | 93.4 | |
| ICLBackbone=Llama-1 7B, Shots=128-shots2024.05 | 93.3 | |
| kNN-promptBackbone=Llama-1 7B, Shots=128-shots2024.05 | 93.3 | |
| ICLBackbone=Llama-2 70B, Shots=128-shots2024.05 | 93.2 | |
| Zero-shotModel=Llama3.1-70B2025.09 | 93.2 | |
| kNN-promptBackbone=Llama-2 13B, Shots=128-shots2024.05 | 93.1 | |
| FADS-ICLBackbone=Llama-2 13B, Shots=128-shots2024.05 | 93.1 | |
| TIACBMModel=GPT-2, Fine-Tuning Strategy=TIACBM (ours)2025.02 | 92.96 | |
| CAMEmode=fine-tuning, batch size=32k2023.07 | 92.9 | |
| Baselinemode=fine-tuning2023.07 | 92.8 | |
| CAMEmode=fine-tuning, batch size=8k2023.07 | 92.8 | |
| kNN-promptingBackbone=Llama-1 7B, Shots=128-shots2024.05 | 92.8 | |
| kNN-promptingBackbone=Llama-2 7B, Shots=128-shots2024.05 | 92.8 | |
| kNN-promptingBackbone=Llama-1 30B, Shots=128-shots2024.05 | 92.8 | |
| Cyclic DecayingModel=GPT-2, Fine-Tuning Strategy=Cyclic Decaying2025.02 | 92.74 | |
| DecTn=256, Update model parameters=false2022.12 | 92.7 | |
| FADS-ICLBackbone=Llama-1 30B, Shots=128-shots2024.05 | 92.7 | |
| FADS-ICLBackbone=Llama-2 70B, Shots=128-shots2024.05 | 92.6 | |
| Ankner et al. (2024)Model=GPT-2, Fine-Tuning Strategy=Ankner et al. (2024)2025.02 | 92.6 | |
| ConstantModel=GPT-2, Fine-Tuning Strategy=Constant2025.02 | 92.54 | |
| Fine-tuningn=64, Update model parameters=true2022.12 | 92.5 | |
| DecTn=64, Update model parameters=false2022.12 | 92.4 | |
| FADS-ICLBackbone=Llama-1 7B, Shots=128-shots2024.05 | 92.4 | |
| E5baseBackbone=LLaMA-7B, Evaluation Setting=Few-shot (ICL)2025.05 | 92.4 | |
| ConventionalModel=GPT-2, Fine-Tuning Strategy=Conventional2025.02 | 92.35 | |
| Poesina et al. (2024)Model=GPT-2, Fine-Tuning Strategy=Poesina et al. (2024)2025.02 | 92.27 | |
| FADS-ICLBackbone=Llama-2 7B, Shots=128-shots2024.05 | 92.2 | |
| Rewrite (Pc)Backbone=Gemma-2-9B-IT, Strategy=Score + context rewrite2026.03 | 92.2 | |
| Fine-tuningn=256, Update model parameters=true2022.12 | 92 | |
| kNN-promptBackbone=Llama-2 70B, Shots=128-shots2024.05 | 92 | |
| TEMPERA2026.03 | 92 | |
| kNN-promptingBackbone=Llama-2 70B, Shots=128-shots2024.05 | 91.9 | |
| DecTn (Shots)=16, Tr. Time (s)=3, # Query=1, # Param. (K)=1302022.12 | 91 | |
| Few-shot*Model=Llama3.1-70B2025.09 | 91 | |
| DecTn (Shots)=12022.12 | 90.8 | |
| GPT Large FTBackbone=GPT Large, Shots=N/A (Fine-tuned)2024.05 | 90.7 | |
| FADS-ICLcandidate pool size (m)=2562024.05 | 90.5 | |
| BBTv2n (Shots)=16, Tr. Time (s)=9856, # Query=8000, # Param. (K)=122022.12 | 90.3 | |
| RLPrompt2026.03 | 90.1 | |
| Few-shot*Model=Qwen3-32B2025.09 | 89.8 | |
| BBTn (Shots)=16, Tr. Time (s)=10512, # Query=8000, # Param. (K)=0.52022.12 | 89.6 | |
| PromptBoostingn (Shots)=42022.12 | 88.9 | |
| FADS-ICLNumber of training shots (m)=128, LLM scale=1.5B2024.05 | 88.9 | |
| FADS-ICLBackbone=GPT-2 1.5B, Shots=128-shots2024.05 | 88.9 | |
| FADS-ICLcandidate pool size (m)=1282024.05 | 88.9 | |
| Latent Concept LearningLLM=GPT3-c (6.7B)2023.01 | 88.8 | |
| I2CLModel=Llama3.1-70B2025.09 | 88.8 | |
| EPRBackbone=LLaMA-7B, Evaluation Setting=Few-shot (ICL)2025.05 | 88.7 | |
| SimilarLLM=GPT3-d (175B)2023.01 | 88.5 | |
| BERT Large FTBackbone=BERT Large, Shots=N/A (Fine-tuned)2024.05 | 88.3 | |
| FADS-ICLNumber of training shots (m)=64, LLM scale=1.5B2024.05 | 88.2 | |
| FADS-ICLcandidate pool size (m)=642024.05 | 88.2 | |
| FADS-ICLNumber of training shots (m)=32, LLM scale=1.5B2024.05 | 88 | |
| kNN-promptBackbone=GPT-2 0.8B, Shots=128-shots2024.05 | 88 | |
| FADS-ICLcandidate pool size (m)=322024.05 | 88 | |
| Latent Concept LearningLLM=GPT3-d (175B)2023.01 | 87.8 | |
| SBERTBackbone=LLaMA-7B, Evaluation Setting=Few-shot (ICL)2025.05 | 87.8 | |
| FADS-ICLBackbone=GPT-2 0.8B, Shots=128-shots2024.05 | 87.7 | |
| Rewrite (Ps)Backbone=Gemma-2-9B-IT, Strategy=Score-only rewrite2026.03 | 87.7 | |
| DecTn (Shots)=42022.12 | 87.6 | |
| PromptBoostingn (Shots)=16, Tr. Time (s)=644, # Query=10, # Param. (K)=0.42022.12 | 87.6 | |
| Latent Concept LearningLLM=GPT3-b (1.3B)2023.01 | 87.3 | |
| ProGenClassifier=DistillBERT2023.05 | 87.2 | |
| kNN-promptingNumber of training shots (m)=64, LLM scale=1.5B2024.05 | 87.2 | |
| RLPromptn (Shots)=16, Tr. Time (s)=65579, # Query=12000, # Param. (K)=31002022.12 | 87 | |
| SuperGenClassifier=DistillBERT2023.05 | 86.7 | |
| PromptBoostingn (Shots)=12022.12 | 86.7 | |
| BBTv2n (Shots)=42022.12 | 86.6 | |
| FADS-ICLNumber of training shots (m)=16, LLM scale=1.5B2024.05 | 86.6 | |
| UniformLLM=GPT3-d (175B)2023.01 | 86.5 | |
| ICRModel=Qwen3-32B2025.09 | 86.4 | |
| FVModel=Llama3.1-70B2025.09 | 86.4 | |
| Latent Concept LearningLLM=GPT2-l (774M)2023.01 | 86.2 | |
| kNN-promptNumber of training shots (m)=128, LLM scale=1.5B2024.05 | 86.1 | |
| kNN-promptBackbone=GPT-2 1.5B, Shots=128-shots2024.05 | 86.1 | |
| kNN-promptingNumber of training shots (m)=128, LLM scale=1.5B2024.05 | 86 | |
| kNN-promptingBackbone=GPT-2 1.5B, Shots=128-shots2024.05 | 86 | |
| SimilarLLM=GPT3-c (6.7B)2023.01 | 85.7 | |
| I2CLModel=Qwen3-32B2025.09 | 85.6 | |
| Latent Concept LearningLLM=GPT3-a (350M)2023.01 | 85.4 |