Toxicity Mitigation on ATTAQ
0.122Average Max ToxicityM+
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| M+Backbone=Aya-23-8B, Training Dataset=RTP-Challenging, Generations per prompt=252025.11 | 0.122 | 0 | 10.81 | 23.7 | |
| M+Backbone=Llama-2-7B, Training Dataset=RTP-Challenging, Generations per prompt=252025.11 | 0.207 | 5.8 | 9.29 | 11.8 | |
| M+Backbone=Llama-3-8B, Training Dataset=RTP-Challenging, Generations per prompt=252025.11 | 0.331 | 17.5 | 11.26 | 16 | |
| M'Backbone=Aya-23-8B, Training Dataset=RTP-Challenging, Generations per prompt=252025.11 | 0.364 | 20 | 8.35 | 12.9 | |
| M'Backbone=Llama-3-8B, Training Dataset=RTP-Challenging, Generations per prompt=252025.11 | 0.426 | 35.8 | 8.36 | 12.7 | |
| M'Backbone=Llama-2-7B, Training Dataset=RTP-Challenging, Generations per prompt=252025.11 | 0.468 | 41.7 | 7.63 | 10 | |
| MBackbone=Llama-3-8B, Training Dataset=RTP-Challenging, Generations per prompt=252025.11 | 0.643 | 80.8 | 7.48 | 15 | |
| MBackbone=Aya-23-8B, Training Dataset=RTP-Challenging, Generations per prompt=252025.11 | 0.661 | 75 | 7.34 | 13.7 | |
| MBackbone=Llama-2-7B, Training Dataset=RTP-Challenging, Generations per prompt=252025.11 | 0.682 | 84.2 | 7.21 | 12.3 |