Emotion Classification on GoEmotions
55.8AccuracyVariational Credal
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| Variational Credal2026.02 | 55.8 | 0.06 | 0.68 | 0.75 | 0.33 | — | — | |
| CBDL2026.02 | 55.2 | 0.6 | 0.33 | 0.7 | 0.3 | — | — | |
| Deep Ensembles2026.02 | 55.1 | 0.8 | 0.28 | 0.67 | 0.26 | — | — | |
| CreINNs2026.02 | 55 | 0.63 | 0.31 | 0.69 | 0.29 | — | — | |
| MC Dropout2026.02 | 54.8 | 0.78 | 0.26 | 0.66 | 0.25 | — | — | |
| Semantic Entropy2026.02 | 54.2 | 0.83 | 0.25 | 0.65 | 0.24 | — | — | |
| P(True)2026.02 | 53.8 | 0.76 | 0.22 | 0.63 | 0.22 | — | — | |
| CREDENCE-DeBERTaCategory=Encoder, Backbone=DeBERTa, Parameters=66M–184M2026.04 | 51.2 | — | — | — | — | 0.095 | 0.198 | |
| CREDENCE-MistralCategory=LLM, Backbone=Mistral-7B, Parameters=3.8B–8B, Adapter=LoRA2026.04 | 49.4 | — | — | — | — | 0.103 | 0.185 | |
| CREDENCE-Llama-3Category=LLM, Backbone=Llama-3, Parameters=3.8B–8B, Adapter=LoRA2026.04 | 48.9 | — | — | — | — | 0.091 | 0.182 | |
| CREDENCE-RoBERTaCategory=Encoder, Backbone=RoBERTa, Parameters=66M–184M2026.04 | 48.7 | — | — | — | — | 0.071 | 0.198 | |
| Deep EnsembleCategory=General2026.04 | 48.2 | — | — | — | — | 0.056 | 0.112 | |
| CBM + EnsembleCategory=CBM2026.04 | 47.9 | — | — | — | — | 0.052 | 0.118 | |
| Temp. ScalingCategory=General2026.04 | 47.5 | — | — | — | — | 0.029 | 0.076 | |
| P-CBMCategory=CBM2026.04 | 47.2 | — | — | — | — | 0.048 | 0.108 | |
| MC DropoutCategory=General2026.04 | 47.1 | — | — | — | — | 0.042 | 0.089 | |
| Standard CBMCategory=CBM2026.04 | 46.8 | — | — | — | — | 0.038 | 0.094 | |
| CBM + MC DropCategory=CBM2026.04 | 46.5 | — | — | — | — | 0.044 | 0.102 | |
| CREDENCE-Phi-3Category=LLM, Backbone=Phi-3, Parameters=3.8B–8B, Adapter=LoRA2026.04 | 46.1 | — | — | — | — | 0.082 | 0.176 | |
| Evidential DLCategory=General2026.04 | 45.8 | — | — | — | — | 0.031 | 0.078 | |
| CREDENCE-DistilBERTCategory=Encoder, Backbone=DistilBERT, Parameters=66M–184M2026.04 | 45.3 | — | — | — | — | 0.089 | 0.183 | |
| gpt-oss-120BN (#rows)=1,000, K (#candidate labels)=28, Decoding Strategy=prefill-only logit scoring2026.05 | 36.3 | — | — | — | — | — | — | |
| Llama-3.3-70B-Instruct-FP8N (#rows)=1,000, K (#candidate labels)=28, Decoding Strategy=greedy2026.05 | 35.6 | — | — | — | — | — | — | |
| Credal CBM2026.02 | — | 0.06 | 0.65 | 0.75 | — | — | — | |
| Deep Ensembles2026.02 | — | 0.8 | 0.28 | 0.68 | — | — | — | |
| P(True)2026.02 | — | 0.78 | 0.21 | 0.62 | — | — | — | |
| Rep. Probes2026.02 | — | 0.76 | 0.27 | 0.65 | — | — | — | |
| Semantic Entropy2026.02 | — | 0.83 | 0.25 | 0.65 | — | — | — |