Question Classification on TREC (test)
97.53AccuracyCKD
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| CKDStudent Model Architecture=6L-768D2026.02 | 97.53 | — | |
| Fine-tuning (full)Backbone=RoBERTa-large, Training Samples (K)=full, Evaluation Protocol=Full fine-tuning2021.08 | 97.4 | — | |
| FT TeacherModel Role=Teacher2026.02 | 97.4 | — | |
| CUDStudent Model Architecture=6L-768D2026.02 | 97.4 | — | |
| HyperPELTNumber of Samples=20002022.03 | 97.2 | — | |
| Adapters_BASEnumber of samples=20002021.06 | 97.03 | — | |
| HYPERFORMER++_BASEnumber of samples=20002021.06 | 96.92 | — | |
| T5_BASEnumber of samples=20002021.06 | 96.87 | — | |
| HYPERFORMER++_BASEnumber of samples=10002021.06 | 96.72 | — | |
| LoRA-SAMBackbone=RoBERTa-large, Number of Parameters=355M, Evaluation Protocol=few-shot2024.10 | 96.7 | — | |
| LoRA-oBARBackbone=RoBERTa-large, Number of Parameters=355M, Evaluation Protocol=few-shot2024.10 | 96.7 | — | |
| LoRA-nBARBackbone=RoBERTa-large, Number of Parameters=355M, Evaluation Protocol=few-shot2024.10 | 96.7 | — | |
| LoRABackbone=RoBERTa-large, Number of Parameters=355M, Evaluation Protocol=few-shot2024.10 | 96.6 | — | |
| CUDStudent Model Architecture=4L-256D2026.02 | 96.47 | — | |
| PKDStudent Model Architecture=6L-768D2026.02 | 96.4 | — | |
| TinyBERTStudent Model Architecture=6L-768D, trained with distillation at pretraining stage=true2026.02 | 96.13 | — | |
| LKDStudent Model Architecture=6L-768D2026.02 | 96.07 | — | |
| Adapters_BASEnumber of samples=10002021.06 | 96.06 | — | |
| AD-KDStudent Model Architecture=6L-768D2026.02 | 96 | — | |
| AdamWBackbone=RoBERTa-large, Optimization Method=AdamW, Finetuning Protocol=Few-shot, Iterations=10K2025.11 | 95.9 | — | |
| LKDStudent Model Architecture=4L-256D2026.02 | 95.8 | — | |
| PKDStudent Model Architecture=4L-256D2026.02 | 95.8 | — | |
| T5_BASEnumber of samples=10002021.06 | 95.5 | — | |
| CKDStudent Model Architecture=4L-256D2026.02 | 95.4 | — | |
| AD-KDStudent Model Architecture=4L-256D2026.02 | 95.06 | — | |
| SRU (4 layers)Size=303k, Optimizer=Adam, Word embeddings=fixed2017.09 | 94.8 | — | |
| HYPERFORMER++_BASEnumber of samples=5002021.06 | 94.78 | — | |
| SRU (8 layers)Size=502k, Optimizer=Adam, Word embeddings=fixed2017.09 | 94.7 | — | |
| SRU (2 layers)Size=204k, Optimizer=Adam, Word embeddings=fixed2017.09 | 94 | — | |
| Adapters_BASEnumber of samples=5002021.06 | 93.65 | — | |
| CNN2018.05 | 93.6 | — | |
| Kim (2014)2017.09 | 93.6 | — | |
| CNNtype=supervised compositional model2015.06 | 93.6 | — | |
| T5_BASEnumber of samples=5002021.06 | 93.57 | — | |
| LSTMSize=352k, Optimizer=Adam, Word embeddings=fixed2017.09 | 93.4 | — | |
| CNNSize=360k, Optimizer=Adam, Word embeddings=fixed2017.09 | 93.2 | — | |
| QRNN (k=1) + highwaySize=204k, Optimizer=Adam, Word embeddings=fixed2017.09 | 93.2 | — | |
| Dynamic CNN2018.05 | 93 | — | |
| Kalchbrenner et al. (2014)2017.09 | 93 | — | |
| QRNN (k=1)Size=165k, Optimizer=Adam, Word embeddings=fixed2017.09 | 92.5 | — | |
| Zhao et al. (2015)2017.09 | 92.4 | — | |
| AdaSenttype=supervised compositional model2015.06 | 92.4 | — | |
| SWEM-aver2018.05 | 92.2 | — | |
| combine-skipclassifier=logistic regression2015.06 | 92.2 | — | |
| SWEM-concat2018.05 | 91.8 | — | |
| Paragraph-vectortype=unsupervised learning2015.06 | 91.8 | — | |
| Zhang and Wallace (2017)2017.09 | 91.6 | — | |
| uni-skipclassifier=logistic regression2015.06 | 91.4 | — | |
| LLM2LLM% Data=2.2, # Seed Examples=120, # Augmented=442024.03 | 91.2 | — | |
| BRNNtype=supervised compositional model2015.06 | 91 | — | |
| RNN2018.05 | 90.2 | — | |
| LLM2LLM% Data=1.6, # Seed Examples=90, # Augmented=222024.03 | 90.2 | — | |
| RNNtype=supervised compositional model2015.06 | 90.2 | — | |
| ConMeZOBackbone=RoBERTa-large, Optimization Method=ConMeZO, Finetuning Protocol=Few-shot, Iterations=10K, Smoothing parameter=10^-32025.11 | 90 | — | |
| LM-BFFShots=16/162021.07 | 89.4 | — | |
| bi-skipclassifier=logistic regression2015.06 | 89.4 | — | |
| MeZO+MomentumBackbone=RoBERTa-large, Optimization Method=MeZO+Momentum (Mom.), Finetuning Protocol=Few-shot, Iterations=10K, Smoothing parameter=10^-32025.11 | 89.2 | — | |
| SWEM-max2018.05 | 89 | — | |
| Fine-tuningBackbone=RoBERTa-large, Training Samples (K)=16, Evaluation Protocol=Few-shot fine-tuning2021.08 | 88.8 | — | |
| HYPERFORMER++_BASEnumber of samples=1002021.06 | 88.42 | — | |
| Hyperformer++Number of Samples=1002022.03 | 88.42 | — | |
| GrConvtype=supervised compositional model2015.06 | 88.4 | — | |
| MeZOBackbone=RoBERTa-large, Optimization Method=MeZO, Finetuning Protocol=Few-shot, Iterations=10K, Smoothing parameter=10^-32025.11 | 88.4 | — | |
| LM-BFFBackbone=RoBERTa-large, Training Samples (K)=16, Evaluation Protocol=Few-shot fine-tuning2021.08 | 88.2 | — | |
| TextGrad-MLMforward=Gemini 2.5 flash-lite, LMbackward=Gemini 2.5 flash, Optimization Granularity=Blockwise2025.05 | 87.98 | 0.91 | |
| T5_BASEnumber of samples=1002021.06 | 87.79 | — | |
| MGSKDStudent Model Architecture=6L-768D, trained with distillation at pretraining stage=true2026.02 | 87.6 | — | |
| Prompt-tuningNumber of Samples=5002022.03 | 87.52 | — | |
| cBoWtype=bag-of-words model2015.06 | 87.3 | — | |
| DARTBackbone=RoBERTa-large, Training Samples (K)=16, Evaluation Protocol=Few-shot fine-tuning2021.08 | 87.1 | — | |
| GEPALMforward=Gemini 2.5 flash-lite, LMbackward=Gemini 2.5 flash2025.05 | 87 | 0.024 | |
| TextGrad-MLMforward=Gemini 2.5 flash-lite, LMbackward=Gemini 2.5 flash, Optimization Granularity=Promptwise2025.05 | 86.82 | 0.76 | |
| UniFewShots=16/162021.07 | 86.7 | — | |
| AdalFlowLMforward=Gemini 2.5 flash-lite, LMbackward=Gemini 2.5 flash2025.05 | 86.67 | 0.02 | |
| AdalFlow-MLMforward=Gemini 2.5 flash-lite, LMbackward=Gemini 2.5 flash, Optimization Granularity=Promptwise2025.05 | 86.6 | 0.018 | |
| P-TuningBackbone=RoBERTa-large, Training Samples (K)=16, Evaluation Protocol=Few-shot fine-tuning2021.08 | 86.3 | — | |
| UniFew_metaShots=16/162021.07 | 86.1 | — | |
| TextGradLMforward=Gemini 2.5 flash-lite, LMbackward=Gemini 2.5 flash, Validation Revert Strategy=w/o2025.05 | 84.86 | 1.32 | |
| Few-shot*Backbone=Llama2-7B, Few-shot settings=5-shot ID2025.09 | 84.6 | — | |
| ICRBackbone=Llama2-7B2025.09 | 83.8 | — | |
| COPRO-MLMforward=Gemini 2.5 flash-lite, LMbackward=Gemini 2.5 flash, Optimization Granularity=Promptwise2025.05 | 82.57 | 0.98 | |
| COPRO-MLMforward=Gemini 2.5 flash-lite, LMbackward=Gemini 2.5 flash, Optimization Granularity=Blockwise2025.05 | 82.21 | 1.21 | |
| M²IVBackbone=Llama2-7B2025.09 | 81.5 | — | |
| COPROLMforward=Gemini 2.5 flash-lite, LMbackward=Gemini 2.5 flash2025.05 | 81.45 | 1.25 | |
| Baseline% Data=2.2, # Seed Examples=120, # Augmented=442024.03 | 81.2 | — | |
| LIVEBackbone=Llama2-7B2025.09 | 81 | — | |
| Baseline% Data=1.6, # Seed Examples=90, # Augmented=222024.03 | 80.8 | — | |
| TextGradLMforward=Gemini 2.5 flash-lite, LMbackward=Gemini 2.5 flash, Validation Revert Strategy=w/2025.05 | 80.12 | 0 | |
| LLM2LLM% Data=1.1, # Seed Examples=60, # Augmented=1052024.03 | 78.8 | — | |
| I2CLBackbone=Llama2-7B2025.09 | 78.6 | — | |
| Adapters_BASEnumber of samples=1002021.06 | 78.07 | — | |
| LONGLLAMAModel size=7B, Context length=8K2023.07 | 75.9 | — | |
| LONGLLAMAModel size=7B, Context length=6K2023.07 | 74.9 | — | |
| LONGLLAMAModel size=3B, Context length=8K2023.07 | 73.3 | — | |
| LONGLLAMAModel size=3B, Context length=6K2023.07 | 72.9 | — | |
| LONGLLAMAModel size=7B, Context length=4K2023.07 | 72.7 | — | |
| LONGLLAMAModel size=3B, Context length=4K2023.07 | 71.6 | — | |
| M²IVBackbone=Qwen2.5-7B2025.09 | 70.8 | — | |
| ICRBackbone=Qwen2.5-7B2025.09 | 70.6 | — | |
| LIVEBackbone=Qwen2.5-7B2025.09 | 70.4 | — |