Toxicity Evaluation on Toxicity Evaluation Dataset
5Avg. Toxicity ScoreAnthropic Claude Haiku
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Anthropic Claude HaikuType=Closed-source, Refused all prompts?=Yes, Refusal style=Refused immediately and firmly, Safety fine-tuning level=Very strong2026.06 | 5 | 1 | |
| Meta Llama 3 8BType=Open-source, Refused all prompts?=Yes, Refusal style=Constructively redirected and Declined, Safety fine-tuning level=Strong2026.06 | 6 | 2 | |
| Google Gemma 2 9BType=Open-source, Refused all prompts?=Yes, Refusal style=Declined firmly and briefly, Safety fine-tuning level=Strong2026.06 | 6 | 2 | |
| Mistral AI Mistral 7BType=Open-source, Refused all prompts?=Yes, Refusal style=Apologised and declined, Safety fine-tuning level=Moderate2026.06 | 8.5 | 4 |