Text Classification on MNLI
88.64AccuracyDECA
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| DECABackbone=Llama-3.1-8B2026.06 | 88.64 | — | 88.46 | |
| Res-TuningTrain Time=47.3, Train Param.=0.97 (0.77%), Test Param.=0.97 (0.77%), Mem.=19.3G2023.10 | 87.45 | — | — | |
| MAM AdapterTrain Time=41.4, Train Param.=46.78 (37.4%), Test Param.=0.61 (0.5%), Mem.=22.4G2023.10 | 87.4 | — | — | |
| theta_B fine-tune|Ds^c|=-2025.10 | 86.3 | — | — | |
| DECABackbone=Qwen2.5-3B2026.06 | 85.67 | — | 85.34 | |
| Dec-LoRABackbone=Llama-3.1-8B2026.06 | 85.15 | — | 84.68 | |
| F-PABEEParams [M]=258, Operating Point=22026.01 | 83.9 | -53 | — | |
| BERT baseParams [M]=345, Operating Point=12026.01 | 83.1 | 0 | — | |
| BERT baseParams [M]=345, Operating Point=22026.01 | 83.1 | 0 | — | |
| Res-Tuning-BypassTrain Time=24.2, Train Param.=0.98 (0.78%), Test Param.=0.98 (0.78%), Mem.=4.3G2023.10 | 82.01 | — | — | |
| Dec-AdapterBackbone=Qwen2.5-3B2026.06 | 80.57 | — | 78.65 | |
| ADEPTParams [M]=42, Operating Point=22026.01 | 80.5 | -52 | — | |
| DeeBERTParams [M]=258, Operating Point=22026.01 | 79.2 | -37 | — | |
| PABEEParams [M]=258, Operating Point=22026.01 | 78.9 | -52 | — | |
| DistilBERTParams [M]=258, Operating Point=22026.01 | 78.6 | -40 | — | |
| BranchyNetParams [M]=258, Operating Point=22026.01 | 78.3 | -53 | — | |
| Shallow-DeepParams [M]=258, Operating Point=22026.01 | 78.2 | -51 | — | |
| Dec-LoRABackbone=Qwen2.5-3B2026.06 | 77.63 | — | 76.01 | |
| DeCAFBackbone=Llama-3.1-8B2026.06 | 72.49 | — | 67.95 | |
| DeCAFBackbone=Qwen2.5-3B2026.06 | 72.09 | — | 67.62 | |
| ADEPTParams [M]=42, Operating Point=12026.01 | 71.8 | -75 | — | |
| theta_B + delta*|Ds^c|=-2025.10 | 69.97 | — | — | |
| F-PABEEParams [M]=258, Operating Point=12026.01 | 66.9 | -72 | — | |
| Shallow-DeepParams [M]=258, Operating Point=12026.01 | 64.1 | -77 | — | |
| PABEEParams [M]=258, Operating Point=12026.01 | 63.9 | -77 | — | |
| BranchyNetParams [M]=258, Operating Point=12026.01 | 63.8 | -76 | — | |
| GCPannotation budget=20%, setting=pool-based active learning, MLP hidden dimension=256, dropout=0.12026.02 | 62.47 | — | — | |
| BADGEannotation budget=20%, setting=pool-based active learning, MLP hidden dimension=256, dropout=0.12026.02 | 60.91 | — | — | |
| CALannotation budget=20%, setting=pool-based active learning, MLP hidden dimension=256, dropout=0.12026.02 | 60.83 | — | — | |
| ALPSannotation budget=20%, setting=pool-based active learning, MLP hidden dimension=256, dropout=0.12026.02 | 60.12 | — | — | |
| CoreSet (k-center)annotation budget=20%, setting=pool-based active learning, MLP hidden dimension=256, dropout=0.12026.02 | 59.6 | — | — | |
| Entropy Samplingannotation budget=20%, setting=pool-based active learning, MLP hidden dimension=256, dropout=0.12026.02 | 59.3 | — | — | |
| Disagreement (Vote Entropy)annotation budget=20%, setting=pool-based active learning, MLP hidden dimension=256, dropout=0.12026.02 | 58.91 | — | — | |
| Least Confidenceannotation budget=20%, setting=pool-based active learning, MLP hidden dimension=256, dropout=0.12026.02 | 57.42 | — | — | |
| Randomannotation budget=20%, setting=pool-based active learning, MLP hidden dimension=256, dropout=0.12026.02 | 51.21 | — | — | |
| GradFix (theta_B + delta^A)|Ds^c|=502025.10 | 49.68 | — | — | |
| theta_B zero-shot|Ds^c|=-2025.10 | 35.21 | — | — | |
| Dec-AdapterBackbone=Llama-3.1-8B2026.06 | 32.68 | — | 20.1 | |
| theta_B + tau_A|Ds^c|=-2025.10 | 30.75 | — | — | |
| theta_B^opt|Ds^c|=502025.10 | 26.05 | — | — | |
| DeeBERTParams [M]=258, Operating Point=12026.01 | — | -75 | — | |
| DistilBERTParams [M]=258, Operating Point=12026.01 | — | -75 | — |