Sentiment Classification on MR (test)
98.6AccuracyFew-shot*
Evaluation Results
| Method | Links | |
|---|---|---|
| Few-shot*Backbone=Llama2-7B, Few-shot settings=3-shot OOD2025.09 | 98.6 | |
| Fine-tuning (full)Backbone=RoBERTa-large, Training Samples (K)=full, Evaluation Protocol=Full fine-tuning2021.08 | 90.8 | |
| UniFew_metaShots=02021.07 | 90.5 | |
| UniFew_metaShots=16/162021.07 | 90.2 | |
| ICRBackbone=Qwen2.5-7B2025.09 | 89.4 | |
| ICL-goldModel=GPT-NeoX, Prompting=Direct, Training data access=true2022.12 | 89 | |
| DARTBackbone=RoBERTa-large, Training Samples (K)=16, Evaluation Protocol=Few-shot fine-tuning2021.08 | 88.2 | |
| RETROPROMPTShot=16-shot2022.05 | 88 | |
| RETROPROMPTshot=16-shot, domain_role=Source2025.12 | 88 | |
| LM-BFFShots=16/162021.07 | 87.7 | |
| ICL-randomModel=GPT-J, Prompting=Direct, Training data access=true2022.12 | 87.3 | |
| UniFewShots=16/162021.07 | 87.2 | |
| LM-BFF (man)Shot=16-shot2022.05 | 87 | |
| LM-BFF (man)shot=16-shot, domain_role=Source2025.12 | 87 | |
| ICL-goldModel=GPT-J, Prompting=Channel, Training data access=true2022.12 | 86.9 | |
| KPTShot=16-shot2022.05 | 86.8 | |
| KPTshot=16-shot, domain_role=Source2025.12 | 86.8 | |
| P-TuningBackbone=RoBERTa-large, Training Samples (K)=16, Evaluation Protocol=Few-shot fine-tuning2021.08 | 86.7 | |
| ICL-randomModel=GPT-J, Prompting=Channel, Training data access=true2022.12 | 86.6 | |
| LM-BFF (D-demo)Shot=16-shot2022.05 | 86.6 | |
| LM-BFF (D-demo)shot=16-shot, domain_role=Source2025.12 | 86.6 | |
| ICL-randomModel=GPT-NeoX, Prompting=Channel, Training data access=true2022.12 | 86.3 | |
| ICL-goldModel=GPT-NeoX, Prompting=Channel, Training data access=true2022.12 | 86.2 | |
| LM-BFFBackbone=RoBERTa-large, Training Samples (K)=16, Evaluation Protocol=Few-shot fine-tuning2021.08 | 85.5 | |
| Deep-ThinkingIn-context learning setting=w/ dev set2023.05 | 84.8 | |
| MEANmodel=full model2018.07 | 84.5 | |
| ICL-goldModel=GPT-J, Prompting=Direct, Training data access=true2022.12 | 84 | |
| Z-ICLModel=GPT-NeoX, Prompting=Direct, Training data access=false2022.12 | 84 | |
| MEAN w/o intensity wordsintensity words=excluded2018.07 | 83.5 | |
| MEAN w/o CharCNNcharacter-level embedding=excluded2018.07 | 83.2 | |
| LENSIn-context learning setting=w/ dev set2023.05 | 83.1 | |
| NSCL2018.07 | 82.9 | |
| MEAN w/o negation wordsnegation words=excluded2018.07 | 82.9 | |
| LR-Bi-LSTM2018.07 | 82.1 | |
| MEAN w/o sentiment wordssentiment words=excluded2018.07 | 82.1 | |
| Z-ICLModel=GPT-J, Prompting=Channel, Training data access=false2022.12 | 81.9 | |
| RandomIn-context learning setting=w/ dev set2023.05 | 81.8 | |
| Self-attentionimplementation=re-implemented by current paper2018.07 | 81.7 | |
| HBMPEmbedding Dimension=1200D, Evaluation Protocol=SentEval2018.08 | 81.7 | |
| ID-LSTM2018.07 | 81.6 | |
| CNN2018.07 | 81.5 | |
| HBMPEmbedding Dimension=600D, Evaluation Protocol=SentEval2018.08 | 81.5 | |
| ICL-randomModel=GPT-NeoX, Prompting=Direct, Training data access=true2022.12 | 81.2 | |
| InferSentEvaluation Protocol=SentEval2018.08 | 81.1 | |
| Z-ICLModel=GPT-J, Prompting=Direct, Training data access=false2022.12 | 81 | |
| LM-BFF_manShots=02021.07 | 80.8 | |
| Prompt-based zero-shotBackbone=RoBERTa-large, Training Samples (K)=0, Evaluation Protocol=Zero-shot2021.08 | 80.8 | |
| Tree-LSTMsource=Qian et al., 20172018.07 | 80.7 | |
| GPT-3 in-context learningBackbone=RoBERTa-large, Training Samples (K)=16, Evaluation Protocol=In-context learning2021.08 | 80.5 | |
| ICRBackbone=Llama2-7B2025.09 | 79.8 | |
| SkipThoughtEvaluation Protocol=SentEval2018.08 | 79.4 | |
| BiLSTMsource=Qian et al., 20172018.07 | 79.3 | |
| LSTMsource=Qian et al., 20172018.07 | 77.4 | |
| Fine-tuningBackbone=RoBERTa-large, Training Samples (K)=16, Evaluation Protocol=Few-shot fine-tuning2021.08 | 76.9 | |
| FTShot=16-shot2022.05 | 76.9 | |
| FTshot=16-shot, domain_role=Source2025.12 | 76.9 | |
| Random inputsModel=GPT-J, Prompting=Channel, Training data access=false2022.12 | 76.2 | |
| RNTNsource=Qian et al., 20172018.07 | 75.9 | |
| Random inputsModel=GPT-NeoX, Prompting=Direct, Training data access=false2022.12 | 74.9 | |
| UniFewShots=02021.07 | 74.8 | |
| M²IVBackbone=Llama2-7B2025.09 | 74 | |
| LIVEBackbone=Llama2-7B2025.09 | 73.8 | |
| Z-ICLModel=GPT-NeoX, Prompting=Channel, Training data access=false2022.12 | 73.2 | |
| MentorNetnoise_rate (r)=0.02019.09 | 72.44 | |
| CEnoise_rate (r)=0.02019.09 | 72.35 | |
| FWnoise_rate (r)=0.02019.09 | 72.35 | |
| LCCNnoise_rate (r)=0.02019.09 | 72.35 | |
| GCEnoise_rate (r)=0.02019.09 | 72.24 | |
| Zero-shotBackbone=Llama2-7B2025.09 | 72.2 | |
| DMInoise_rate (r)=0.02019.09 | 72.07 | |
| Naive Z-ICLModel=GPT-NeoX, Prompting=Direct, Training data access=false2022.12 | 71.7 | |
| Deep-ThinkingIn-context learning setting=w/o dev set2023.05 | 71.6 | |
| I2CLBackbone=Llama2-7B2025.09 | 71.6 | |
| M²IVBackbone=Qwen2.5-7B2025.09 | 71.2 | |
| LCCNnoise_rate (r)=0.12019.09 | 70.72 | |
| GCEnoise_rate (r)=0.12019.09 | 70.58 | |
| CEnoise_rate (r)=0.12019.09 | 70.51 | |
| FWnoise_rate (r)=0.12019.09 | 70.49 | |
| DMInoise_rate (r)=0.12019.09 | 70.42 | |
| Few-shot*Backbone=Qwen2.5-7B, Few-shot settings=3-shot OOD2025.09 | 70.2 | |
| MentorNetnoise_rate (r)=0.12019.09 | 69.54 | |
| I2CLBackbone=Qwen2.5-7B2025.09 | 69 | |
| Naive Z-ICLModel=GPT-J, Prompting=Channel, Training data access=false2022.12 | 68.8 | |
| LIVEBackbone=Qwen2.5-7B2025.09 | 68.6 | |
| Random inputsModel=GPT-J, Prompting=Direct, Training data access=false2022.12 | 68.2 | |
| GCEnoise_rate (r)=0.22019.09 | 67.48 | |
| DMInoise_rate (r)=0.22019.09 | 67.44 | |
| LCCNnoise_rate (r)=0.22019.09 | 67.33 | |
| FWnoise_rate (r)=0.22019.09 | 67.14 | |
| CEnoise_rate (r)=0.22019.09 | 67.12 | |
| MentorNetnoise_rate (r)=0.22019.09 | 66.72 | |
| GraphCutIn-context learning setting=w/o dev set2023.05 | 66.3 | |
| CALIn-context learning setting=w/o dev set2023.05 | 66.2 | |
| No-demosModel=GPT-J, Prompting=Channel, Training data access=false2022.12 | 65.7 | |
| DMInoise_rate (r)=0.32019.09 | 65.62 | |
| GCEnoise_rate (r)=0.32019.09 | 65.19 | |
| MentorNetnoise_rate (r)=0.32019.09 | 65.13 | |
| Random inputsModel=GPT-NeoX, Prompting=Channel, Training data access=false2022.12 | 65 | |
| FWnoise_rate (r)=0.32019.09 | 64.92 | |
| CEnoise_rate (r)=0.32019.09 | 64.68 |