Binary Intent Classification on MALINT
70.2UCPI F1 ScoreGPT 4.1 Mini
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| GPT 4.1 MiniModel Type=Large Language Model, Evaluation Protocol=Zero-shot2026.03 | 70.2 | 46.9 | 71.7 | 47.9 | 37.1 | |
| DeBERTa V3 LargeModel Type=Small Language Model, Evaluation Protocol=Fine-tuned2026.03 | 69.6 | 64.9 | 68.3 | 54.7 | 46 | |
| RoBERTa BaseModel Type=Small Language Model, Evaluation Protocol=Fine-tuned2026.03 | 69.3 | 54.7 | 67.4 | 51.5 | 48.6 | |
| RoBERTa LargeModel Type=Small Language Model, Evaluation Protocol=Fine-tuned2026.03 | 68.2 | 63 | 68 | 50.5 | 44.4 | |
| Gemma 3 27B itModel Type=Large Language Model, Evaluation Protocol=Zero-shot2026.03 | 68.2 | 39.5 | 66.7 | 42.4 | 40.7 | |
| DeBERTa V3 BaseModel Type=Small Language Model, Evaluation Protocol=Fine-tuned2026.03 | 67.5 | 50.5 | 58 | 52.3 | 40 | |
| Gemini 2.0 FlashModel Type=Large Language Model, Evaluation Protocol=Zero-shot2026.03 | 63.9 | 60.4 | 72.2 | 45.2 | 44.4 | |
| DistilBERT BaseModel Type=Small Language Model, Evaluation Protocol=Fine-tuned2026.03 | 59.9 | 54.7 | 56.4 | 45 | 40 | |
| LR with BoWModel Type=Baseline2026.03 | 58.1 | 47.7 | 59.5 | 42.4 | 37.6 | |
| Llama 3.3 70BModel Type=Large Language Model, Evaluation Protocol=Zero-shot2026.03 | 56.9 | 42.7 | 73.8 | 41.5 | 49.6 | |
| BERT BaseModel Type=Small Language Model, Evaluation Protocol=Fine-tuned2026.03 | 56.2 | 48.4 | 50 | 61.4 | 29.3 | |
| GPT 4o MiniModel Type=Large Language Model, Evaluation Protocol=Zero-shot2026.03 | 54.3 | 54.7 | 63.2 | 45.8 | 32.4 | |
| BERT LargeModel Type=Small Language Model, Evaluation Protocol=Fine-tuned2026.03 | 52.8 | 43.7 | 54.3 | 52.9 | 30.6 | |
| RandomModel Type=Baseline2026.03 | 27.9 | 20.5 | 12.2 | 17.9 | 16.2 |