Intent Classification on Banking77 (test)
93.83AccuracyCUD
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| CUDStudent Model Architecture=6L-768D2026.02 | 93.83 | — | — | — | — | |
| SOLOISTNumber of training examples per intent=Full2020.05 | 93.8 | — | — | — | — | |
| FT TeacherModel Role=Teacher2026.02 | 93.7 | — | — | — | — | |
| BERT-TunedNumber of training examples per intent=Full2020.05 | 93.66 | — | — | — | — | |
| CKDStudent Model Architecture=6L-768D2026.02 | 93.54 | — | — | — | — | |
| CAPHybridTraining setting=full-train2025.12 | 93.51 | — | 68.21 | 23.98 | — | |
| CAPTraining setting=full-train2025.12 | 93.41 | — | 68 | 24.3 | — | |
| USE+ConveRTNumber of training examples per intent=Full2020.05 | 93.36 | — | — | — | — | |
| ClearAdapterNoise Type=None, Noise Level=0%2024.10 | 93.1 | 93.1 | — | — | — | |
| LKDStudent Model Architecture=6L-768D2026.02 | 93.03 | — | — | — | — | |
| IGTraining setting=full-train2025.12 | 93.02 | — | 61.75 | 32.81 | — | |
| ConveRTNumber of training examples per intent=Full2020.05 | 93.01 | — | — | — | — | |
| LoRANoise Type=None, Noise Level=0%2024.10 | 93 | 93 | — | — | — | |
| Full Fine-tuningNoise Type=None, Noise Level=0%2024.10 | 92.9 | 92.9 | — | — | — | |
| Cayley-Encoder (GCN)Base Model=Llama3-8B2026.03 | 92.85 | — | — | — | — | |
| LIMETraining setting=full-train2025.12 | 92.82 | — | 62.87 | 32.23 | — | |
| USENumber of training examples per intent=Full2020.05 | 92.81 | — | — | — | — | |
| CleaRLoRANoise Type=None, Noise Level=0%2024.10 | 92.8 | 92.8 | — | — | — | |
| AdapterNoise Type=None, Noise Level=0%2024.10 | 92.7 | 92.7 | — | — | — | |
| AD-KDStudent Model Architecture=6L-768D2026.02 | 92.66 | — | — | — | — | |
| BaseTraining setting=full-train2025.12 | 92.63 | — | 56.73 | 35.87 | — | |
| Cayley-Encoder (GIN)Base Model=Gemma2-2B2026.03 | 92.58 | — | — | — | — | |
| BitFitNoise Type=None, Noise Level=0%2024.10 | 92.5 | 92.5 | — | — | — | |
| Cayley-Encoder (GIN)Base Model=Llama3-8B2026.03 | 92.46 | — | — | — | — | |
| ClearBitFitNoise Type=None, Noise Level=0%2024.10 | 92.4 | 92.4 | — | — | — | |
| FC-Encoder (GIN)Base Model=Gemma2-2B2026.03 | 92.39 | — | — | — | — | |
| FC-Encoder (GCN)Base Model=Llama3-8B2026.03 | 92.38 | — | — | — | — | |
| Meta-SelModel=GPT-OSS-20B2026.02 | 92.3 | — | — | — | — | |
| CleaRPromptNoise Type=None, Noise Level=0%2024.10 | 92.1 | 92.1 | — | — | — | |
| FC-Encoder (GIN)Base Model=Llama3-8B2026.03 | 92.1 | — | — | — | — | |
| ReLUBackbone=Qwen3, MLP Width=256, Activation Scheme=ReLU2026.05 | 91.91 | — | — | — | — | |
| PreciseBackbone=Qwen3, MLP Width=256, Activation Scheme=Prec.2026.05 | 91.91 | — | — | — | — | |
| PromptNoise Type=None, Noise Level=0%2024.10 | 91.9 | 91.9 | — | — | — | |
| Cayley-Encoder (GCN)Base Model=Gemma2-2B2026.03 | 91.86 | — | — | — | — | |
| OLABackbone=Qwen3, MLP Width=256, Activation Scheme=OLA2026.05 | 91.84 | — | — | — | — | |
| RDESModel=GPT-OSS-20B2026.02 | 91.8 | — | — | — | — | |
| FC-Encoder (GCN)Base Model=Gemma2-2B2026.03 | 91.71 | — | — | — | — | |
| Remez-7Backbone=Qwen3, MLP Width=256, Activation Scheme=Rmz-72026.05 | 91.68 | — | — | — | — | |
| MGSKDStudent Model Architecture=6L-768D, trained with distillation at pretraining stage=true2026.02 | 91.58 | — | — | — | — | |
| QUAD4FHEBackbone=Qwen3, MLP Width=256, Activation Scheme=Quad.2026.05 | 91.55 | — | — | — | — | |
| InfluenceModel=GPT-OSS-20B2026.02 | 91.4 | — | — | — | — | |
| UncertaintyModel=GPT-OSS-20B2026.02 | 91.4 | — | — | — | — | |
| CUDStudent Model Architecture=4L-256D2026.02 | 91.22 | — | — | — | — | |
| FC-Encoder (GCN)Base Model=Pythia-410m2026.03 | 90.65 | — | — | — | — | |
| USE+ConveRTNumber of training examples per intent=302020.05 | 90.57 | — | — | — | — | |
| ClearBitFitNoise Type=Asymmetric, Noise Level=10%2024.10 | 90.4 | 90.7 | — | — | — | |
| ClearAdapterNoise Type=Asymmetric, Noise Level=10%2024.10 | 90.3 | 91.4 | — | — | — | |
| CleaRLoRANoise Type=Asymmetric, Noise Level=10%2024.10 | 90.3 | 91.3 | — | — | — | |
| Meta-SelModel=Qwen3-8B2026.02 | 90.1 | — | — | — | — | |
| FC-Encoder (GIN)Base Model=Pythia-410m2026.03 | 90.1 | — | — | — | — | |
| DWATTBase Model=Llama3-8B2026.03 | 90.04 | — | — | — | — | |
| BERT-TunedNumber of training examples per intent=302020.05 | 90.03 | — | — | — | — | |
| Set-EncoderBase Model=Gemma2-2B2026.03 | 90.03 | — | — | — | — | |
| PKDStudent Model Architecture=6L-768D2026.02 | 89.87 | — | — | — | — | |
| CKDStudent Model Architecture=4L-256D2026.02 | 89.86 | — | — | — | — | |
| BitFitNoise Type=Asymmetric, Noise Level=10%2024.10 | 89.8 | 90.2 | — | — | — | |
| CleaRLoRANoise Type=Symmetric, Noise Level=20%2024.10 | 89.8 | 90 | — | — | — | |
| USENumber of training examples per intent=302020.05 | 89.74 | — | — | — | — | |
| ClearAdapterNoise Type=Symmetric, Noise Level=20%2024.10 | 89.7 | 90.1 | — | — | — | |
| TinyBERTStudent Model Architecture=6L-768D, trained with distillation at pretraining stage=true2026.02 | 89.47 | — | — | — | — | |
| Cayley-Encoder (GCN)Base Model=Pythia-410m2026.03 | 89.43 | — | — | — | — | |
| ConveRTNumber of training examples per intent=302020.05 | 89.37 | — | — | — | — | |
| SOLOISTNumber of training examples per intent=302020.05 | 89.28 | — | — | — | — | |
| ClearBitFitNoise Type=Symmetric, Noise Level=20%2024.10 | 89.2 | 89.8 | — | — | — | |
| CleaRPromptNoise Type=Asymmetric, Noise Level=10%2024.10 | 89.2 | 89.9 | — | — | — | |
| Cayley-Encoder (GIN)Base Model=Pythia-410m2026.03 | 89.12 | — | — | — | — | |
| DWATTBase Model=Gemma2-2B2026.03 | 88.82 | — | — | — | — | |
| BitFitNoise Type=Symmetric, Noise Level=20%2024.10 | 88.7 | 88.9 | — | — | — | |
| BERT Seq-CLSTraining Strategy=individual2023.05 | 88.6 | — | — | — | — | |
| AdapterNoise Type=Asymmetric, Noise Level=10%2024.10 | 88.6 | 90.3 | — | — | — | |
| LoRANoise Type=Asymmetric, Noise Level=10%2024.10 | 88.6 | 90.1 | — | — | — | |
| PromptNoise Type=Asymmetric, Noise Level=10%2024.10 | 88.4 | 89.7 | — | — | — | |
| LoRANoise Type=Symmetric, Noise Level=20%2024.10 | 88.3 | 89.2 | — | — | — | |
| Meta-SelModel=DeepSeek-R1-14B2026.02 | 87.9 | — | — | — | — | |
| Set-EncoderBase Model=Llama3-8B2026.03 | 87.62 | — | — | — | — | |
| Full Fine-tuningNoise Type=Asymmetric, Noise Level=10%2024.10 | 87.6 | 90.8 | — | — | — | |
| CleaRPromptNoise Type=Symmetric, Noise Level=20%2024.10 | 87.6 | 88.1 | — | — | — | |
| MLP Last LayerBase Model=Gemma2-2B2026.03 | 87.47 | — | — | — | — | |
| PromptNoise Type=Symmetric, Noise Level=20%2024.10 | 87.4 | 87.8 | — | — | — | |
| ClearAdapterNoise Type=Symmetric, Noise Level=40%2024.10 | 87.3 | 88.2 | — | — | — | |
| BERT-FixedNumber of training examples per intent=Full2020.05 | 87.19 | — | — | — | — | |
| ClearBitFitNoise Type=Symmetric, Noise Level=40%2024.10 | 86.9 | 87.3 | — | — | — | |
| CleaRLoRANoise Type=Symmetric, Noise Level=40%2024.10 | 86.9 | 87.4 | — | — | — | |
| MLP Last LayerBase Model=Llama3-8B2026.03 | 86.7 | — | — | — | — | |
| PKDStudent Model Architecture=4L-256D2026.02 | 86.6 | — | — | — | — | |
| ClearAdapterNoise Type=Asymmetric, Noise Level=20%2024.10 | 86.1 | 87.6 | — | — | — | |
| ClearBitFitNoise Type=Asymmetric, Noise Level=20%2024.10 | 86.1 | 87.5 | — | — | — | |
| Switch-ACMoEPre-training dataset=WikiText-103, Task=Finetuning2025.02 | 86.01 | — | — | — | — | |
| BitFitNoise Type=Symmetric, Noise Level=40%2024.10 | 85.9 | 86.7 | — | — | — | |
| CleaRLoRANoise Type=Asymmetric, Noise Level=20%2024.10 | 85.9 | 87.2 | — | — | — | |
| LoRANoise Type=Symmetric, Noise Level=40%2024.10 | 85.8 | 86.8 | — | — | — | |
| AdapterNoise Type=Symmetric, Noise Level=20%2024.10 | 85.4 | 88.5 | — | — | — | |
| USE+ConveRTNumber of training examples per intent=102020.05 | 85.19 | — | — | — | — | |
| PromptNoise Type=Asymmetric, Noise Level=20%2024.10 | 84.9 | 85.4 | — | — | — | |
| CleaRPromptNoise Type=Symmetric, Noise Level=40%2024.10 | 84.9 | 85.8 | — | — | — | |
| CleaRPromptNoise Type=Asymmetric, Noise Level=20%2024.10 | 84.8 | 85.7 | — | — | — | |
| BERT Seq-CLSTraining Strategy=full2023.05 | 84.7 | — | — | — | — | |
| TS-BanditModel=GPT-OSS-20B2026.02 | 84.7 | — | — | — | — | |
| PromptNoise Type=Symmetric, Noise Level=40%2024.10 | 84.5 | 85.6 | — | — | — | |
| USENumber of training examples per intent=102020.05 | 84.23 | — | — | — | — |