Prompt Moderation on FlexBench
83.99Strict AccuracyFlexGuard
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| FlexGuardMode=Continuous-score, Evaluation Strategy=Calibrated thresholding2026.02 | 83.99 | 83.08 | 78.26 | 81.78 | 78.26 | |
| FlexGuardMethod=Calibrated thresholding, Score Type=continuous-score2026.02 | 83.99 | 83.08 | 78.26 | 81.78 | 78.26 | |
| Qwen3Guard-8B-GenCategory=Logit-thresholded moderator2026.02 | 83.01 | 75.23 | 67.06 | 75.1 | 67.06 | |
| Qwen3Guard-Gen-8BRegime=strict2026.02 | 82.55 | 71.36 | 54.75 | 69.55 | 54.75 | |
| BingoGuard-8BCategory=Level-thresholded moderator2026.02 | 81.83 | 72.53 | 68.31 | 74.22 | 68.31 | |
| BingoGuard-8B2026.02 | 80.98 | 79.07 | 67.13 | 75.73 | 67.13 | |
| FlexGuardMode=Continuous-score, Evaluation Strategy=Rubric thresholding2026.02 | 80.63 | 83.6 | 76.63 | 80.29 | 76.63 | |
| FlexGuardMethod=Rubric thresholding, Score Type=continuous-score2026.02 | 80.63 | 83.6 | 76.63 | 80.29 | 76.63 | |
| WildGuard-7B2026.02 | 79.02 | 74.8 | 59.97 | 71.26 | 59.97 | |
| WildGuard-7BCategory=Logit-thresholded moderator2026.02 | 78.76 | 74.41 | 59.2 | 70.79 | 59.2 | |
| Doubao-1.8Category=Rubric-prompted LLM2026.02 | 78.07 | 79.9 | 73.8 | 77.26 | 73.8 | |
| GPT-5Category=Rubric-prompted LLM2026.02 | 70.95 | 77.56 | 71.29 | 73.26 | 70.95 | |
| DeepSeek-R1Category=Rubric-prompted LLM2026.02 | 70.75 | 67.97 | 66.07 | 68.26 | 66.07 | |
| LlamaGuard3-8BCategory=Logit-thresholded moderator2026.02 | 66.67 | 54 | 56.63 | 59.1 | 54 | |
| Qwen3Guard-Gen-8BRegime=loose2026.02 | 58.02 | 64.65 | 66.74 | 63.14 | 58.02 | |
| LlamaGuard3-8B2026.02 | 48.28 | 54 | 56.63 | 52.97 | 48.28 |