Abusive language detection on HatEval (test)
0.651Macro F1Best system (original shared task)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Best system (original shared task)Mode=Original shared task best system2020.10 | 0.651 | — | |
| HateBERTMode=Fine-tuned, Evaluation=In-dataset, Aggregation=Average of 5 runs2020.10 | 0.516 | 0.645 | |
| BERTMode=Fine-tuned, Evaluation=In-dataset, Aggregation=Average of 5 runs2020.10 | 0.48 | 0.633 |