Toxicity Classification on ImplicitHate
83.97AccuracyLlama-3.2-3B-Instruct (SRD)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Llama-3.2-3B-Instruct (SRD)Training Mode=SRD (Self-Reflective Detoxification), Base Model=Llama-3.2-3B-Instruct, Contrastive Dataset Size=20K prompts, Signal List Length=502026.01 | 83.97 | 70.89 | |
| Llama-3.2-3B-Instruct (Vanilla)Training Mode=Vanilla, Base Model=Llama-3.2-3B-Instruct2026.01 | 75.71 | 68.92 |