Text Moderation on HarmBench n = 400
42Flagged CountGemma 3-1B
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Gemma 3-1BBackbone=Gemma 3-1B2026.04 | 42 | 10.2 | 3.1 | 84.6 | |
| RESTABackbone=Llama 3.1-8B2026.04 | 25 | 7.2 | 0.6 | 85.2 | |
| AdaSteerBackbone=Llama 3.1-8B2026.04 | 22 | 6.5 | 1.7 | 94.3 | |
| Llama 3.1-8BBackbone=Llama 3.1-8B2026.04 | 17 | 5.5 | 0.4 | 70.6 | |
| GPT-OSS-20BBackbone=GPT-OSS-20B2026.04 | 16 | 4.6 | 0.4 | 94.7 | |
| TSABackbone=Gemma 3-1B, # Iterations=6.3, Time (s)=19.5, Net (s)=11.32026.04 | 14 | 6.2 | 2.9 | 62.3 | |
| SmoothLLMBackbone=Llama 3.1-8B2026.04 | 13 | 3.5 | 0.1 | 81.1 | |
| Qwen3-14BBackbone=Qwen3-14B2026.04 | 8 | 5.3 | 2 | 44.8 | |
| Phi-3.5-4BBackbone=Phi-3.5-4B2026.04 | 8 | 5 | 1.7 | 52.3 | |
| TSABackbone=GPT-OSS-20B, # Iterations=1.3, Time (s)=9.5, Net (s)=7.52026.04 | 0 | 1.5 | 0.3 | 9.8 | |
| TSABackbone=Qwen3-14B, # Iterations=1.9, Time (s)=20.5, Net (s)=17.92026.04 | 0 | 2.5 | 1 | 9.6 | |
| TSABackbone=Llama 3.1-8B, # Iterations=1.1, Time (s)=10.7, Net (s)=9.02026.04 | 0 | 0.9 | 0.1 | 10 | |
| TSABackbone=Phi-3.5-4B, # Iterations=1.8, Time (s)=12.0, Net (s)=9.42026.04 | 0 | 2 | 0.6 | 9.7 |