Abusive language detection on AbusEval (test)
76.5Macro F1HateBERT
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| HateBERTMode=Fine-tuned, Evaluation=In-dataset, Aggregation=Average of 5 runs2020.10 | 76.5 | 62.3 | |
| BERTMode=Fine-tuned, Evaluation=In-dataset, Aggregation=Average of 5 runs2020.10 | 72.7 | 55.2 | |
| Caselli et al. (2020)Mode=Standard evaluation2020.10 | 71.6 | 53.1 |