Text Classification on TREC
98AccuracyRoBERTa_BASE + UniDrop
Evaluation Results
| Method | Links | |
|---|---|---|
| RoBERTa_BASE + UniDropvariant=BASE, regularization=UniDrop2021.04 | 98 | |
| RoBERTa_BASEvariant=BASE2021.04 | 97.6 | |
| BERT_BASEvariant=BASE2021.04 | 97.2 | |
| FPFTBackbone=Llama2, k (training examples per class)=200, #Param=6.7B2024.02 | 96.76 | |
| AdapterBackbone=GPT2-XL, k (training examples per class)=200, #Param=15.4M2024.02 | 96.6 | |
| inversedMixupK=All2026.01 | 96.5 | |
| ULMFiT2021.04 | 96.4 | |
| LLM-MixK=All2026.01 | 96.4 | |
| LLM-RewK=All2026.01 | 96.2 | |
| BLSTM-2DCNN2019.02 | 96.1 | |
| TreeNet-GloVe2019.02 | 96.1 | |
| DARLM2019.02 | 96 | |
| Fixed ICLBackbone=Llama-3.1-8B, Context Length=90k2025.03 | 96 | |
| Fine-TuningBackbone=Llama-3.1-8B, Context Length=90k2025.03 | 96 | |
| BaseK=All2026.01 | 96 | |
| BTK=All2026.01 | 96 | |
| TextsmoothK=All2026.01 | 95.8 | |
| AWDK=All2026.01 | 95.8 | |
| MixupK=All2026.01 | 95.7 | |
| MPAD2019.08 | 95.6 | |
| MPAD-sentence-att2019.08 | 95.6 | |
| GNNAVI-GCNBackbone=Llama2, k (training examples per class)=200, #Param=16.8M2024.02 | 95.5 | |
| LLM-GenK=All2026.01 | 95.3 | |
| MPAD-clique2019.08 | 95.2 | |
| Fine-TuningBackbone=Llama-2-7B, Context Length=30k2025.03 | 95 | |
| Ret ICLBackbone=Llama-3.1-8B, Context Length=90k2025.03 | 95 | |
| DBSABackbone=Llama-3.1-8B, Context Length=90k2025.03 | 95 | |
| GNNAVI-SAGEBackbone=Llama2, k (training examples per class)=200, #Param=33.6M2024.02 | 94.76 | |
| EDAK=All2026.01 | 94.7 | |
| C-LSTM2019.08 | 94.6 | |
| VLAWE2019.02 | 94.2 | |
| DiSAN2019.08 | 94.2 | |
| BLSTM-Att2019.02 | 93.8 | |
| DC-TreeLSTM2019.02 | 93.8 | |
| MPAD-path2019.08 | 93.8 | |
| CNNmode=non-static2018.03 | 93.6 | |
| CNN2019.08 | 93.6 | |
| LORABackbone=Llama2, k (training examples per class)=200, #Param=4.2M2024.02 | 93.6 | |
| skip-thoughtsd*=48002018.05 | 93 | |
| BLSTM2019.02 | 93 | |
| Fixed ICLBackbone=Llama-2-7B, Context Length=30k2025.03 | 93 | |
| CNNmode=static2018.03 | 92.8 | |
| Capsule-B2018.03 | 92.8 | |
| CNN-LSTMd*=48002018.05 | 92.6 | |
| MC-QTd*=48002018.05 | 92.4 | |
| AdaSent2019.02 | 92.4 | |
| GNNAVI-SAGEBackbone=GPT2-XL, k (training examples per class)=200, #Param=5.1M2024.02 | 92.32 | |
| Combine-skip2019.02 | 92.2 | |
| SWEM-average2019.02 | 92.2 | |
| SDBN-LoRAData %=100%, Base Model=BERT-base2026.06 | 92.12 | |
| Ret ICLBackbone=Llama-2-7B, Context Length=30k2025.03 | 92 | |
| GNNAVI-GCNBackbone=GPT2-XL, k (training examples per class)=200, #Param=2.6M2024.02 | 91.88 | |
| SDBN-AdapterData %=100%, Base Model=BERT-base2026.06 | 91.88 | |
| Tree-LSTM2018.03 | 91.8 | |
| Capsule-A2018.03 | 91.8 | |
| Paragraph vectors2019.02 | 91.8 | |
| SWEM-concat2019.02 | 91.8 | |
| COV + BOW2019.02 | 91.8 | |
| TreeNet2019.02 | 91.6 | |
| COV + Mean + BOW2019.02 | 91.6 | |
| LoRAData %=100%, Base Model=BERT-base2026.06 | 91.44 | |
| CNNmode=rand2018.03 | 91.2 | |
| AdapterData %=100%, Base Model=BERT-base2026.06 | 91.16 | |
| DBSABackbone=Llama-2-7B, Context Length=30k2025.03 | 91 | |
| HN-ATT2019.08 | 90.8 | |
| LORABackbone=GPT2-XL, k (training examples per class)=200, #Param=2.5M2024.02 | 90.8 | |
| SDBN-AdapterData %=50%, Base Model=BERT-base2026.06 | 90.72 | |
| SPGK2019.08 | 90.69 | |
| Drop Clause Tsetlin MachineDrop Clause Probability=0.52021.05 | 90.5 | |
| byte mLSTMd*=40962018.05 | 90.4 | |
| COV + Mean2019.02 | 90.3 | |
| BitFitData %=100%, Base Model=BERT-base2026.06 | 90.2 | |
| LoRAData %=50%, Base Model=BERT-base2026.06 | 90.2 | |
| SDBN-LoRAData %=50%, Base Model=BERT-base2026.06 | 90.16 | |
| AdapterData %=50%, Base Model=BERT-base2026.06 | 90.12 | |
| BonGn=2, d*=V1 + V22018.05 | 90 | |
| DisCn=2-3, d*=3200-48002018.05 | 90 | |
| SDBN-BitFitData %=100%, Base Model=BERT-base2026.06 | 89.96 | |
| BonGn=3, d*=V1 + V2 + V32018.05 | 89.8 | |
| BILSTM2018.03 | 89.6 | |
| DAN2019.08 | 89.6 | |
| LSTM-GRNN2019.08 | 89.4 | |
| BOWbaseline=true2019.02 | 89.3 | |
| à la carten=2, d*=32002018.05 | 89 | |
| à la carten=3, d*=48002018.05 | 89 | |
| LORABackbone=Llama2, k (training examples per class)=5, #Param=4.2M2024.02 | 88.4 | |
| AdapterBackbone=Llama2, k (training examples per class)=200, #Param=198M2024.02 | 88.2 | |
| BitFitData %=50%, Base Model=BERT-base2026.06 | 88.08 | |
| SDBN-BitFitData %=50%, Base Model=BERT-base2026.06 | 88.08 | |
| Tsetlin MachineDrop Clause Probability=02021.05 | 88.05 | |
| ICLBackbone=Qwen-2.5-7B2026.05 | 87.32 | |
| DiSPModel=Qwen2.5-7B, Setting=Few-shot2026.05 | 87.3 | |
| BonGn=1, d*=V12018.05 | 86.8 | |
| LSTM2018.03 | 86.8 | |
| TATRADataset-free=true2026.02 | 86.15 | |
| Sent2Vecn=1-2, d*=7002018.05 | 85.8 | |
| CL-CNN2018.03 | 85.7 | |
| à la carten=1, d*=16002018.05 | 85.6 | |
| VD-CNN2018.03 | 85.4 | |
| ICLBackbone=Qwen-3-8B2026.05 | 85.24 |