Hate speech classification on Latent Hate (test)
69Macro F1 ScoreSMARTER (Llama_DPO-256)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SMARTER (Llama_DPO-256)Training Paradigm=DPO, K shots=256, Model Backbone=Llama, Data%=7%2025.09 | 69 | 100 | |
| SMARTER (T5_DPO-256)Training Paradigm=DPO, K shots=256, Model Backbone=T5, Data%=7%2025.09 | 65 | 94 | |
| GPT-4.1Training Paradigm=16-shot ICL, K shots=16, Model Backbone=GPT-4.12025.09 | 63 | 91 | |
| Llama_FullTraining Paradigm=Full, Model Backbone=Llama, Data%=100%2025.09 | 62 | 90 | |
| T5_FullTraining Paradigm=Full, Model Backbone=T5, Data%=100%2025.09 | 61 | 88 | |
| ModernBERTTraining Paradigm=Full, Model Backbone=ModernBERT, Data%=100%2025.09 | 61 | 88 | |
| GPT-4.1Training Paradigm=Zero-shot, K shots=0, Model Backbone=GPT-4.12025.09 | 60 | 87 | |
| GPT-5-chatTraining Paradigm=16-shot ICL, K shots=16, Model Backbone=GPT-5-chat2025.09 | 60 | 87 | |
| Qwen-32BTraining Paradigm=16-shot ICL, K shots=16, Model Backbone=Qwen-32B2025.09 | 57 | 83 | |
| GPT-4o-miniTraining Paradigm=Zero-shot, K shots=0, Model Backbone=GPT-4o-mini2025.09 | 54 | 78 | |
| GPT-5-chatTraining Paradigm=Zero-shot, K shots=0, Model Backbone=GPT-5-chat2025.09 | 51 | 74 | |
| Qwen-32BTraining Paradigm=Zero-shot, K shots=0, Model Backbone=Qwen-32B2025.09 | 47 | 68 | |
| GPT-4o-miniTraining Paradigm=16-shot ICL, K shots=16, Model Backbone=GPT-4o-mini2025.09 | 25 | 36 |