Natural Language Inference on SNLI Slang variant (test)
92.8AccuracyRoBERTa
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| RoBERTaMitigation Approach=None2026.04 | 92.8 | — | |
| ELECTRA (Hybrid)Training Configuration=Hybrid, Evaluation Protocol=Fine-tuned2026.04 | 89.13 | 0.0001 | |
| Hybrid MitigationMitigation Approach=Hybrid, Base Model=ELECTRA-small2026.04 | 89.13 | — | |
| AugmentationMitigation Approach=Augmentation, Base Model=ELECTRA-small2026.04 | 89.05 | — | |
| PreprocessingMitigation Approach=Preprocessing, Base Model=ELECTRA-small2026.04 | 88.93 | — | |
| ELECTRA (Baseline)Training Configuration=Baseline, Evaluation Protocol=Fine-tuned2026.04 | 87.97 | 0.1247 | |
| ELECTRA-smallMitigation Approach=None2026.04 | 87.97 | — | |
| GPT-4o-miniEvaluation Protocol=Zero-shot2026.04 | 87.33 | — | |
| GPT-4o-miniMitigation Approach=None2026.04 | 87.33 | — | |
| GPT-3.5Mitigation Approach=None2026.04 | 66.22 | — |