Trajectory-level safety evaluation on R-judge (test)
95.2AccuracyGemini-3-Flash
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Gemini-3-FlashModel Category=Closed-Source2026.01 | 95.2 | 98.6 | 92.3 | 95.3 | |
| Gemini-3-ProModel Category=Closed-Source2026.01 | 94.3 | 94 | 95.3 | 94.7 | |
| SFTBackbone=Qwen2.5-3B-Instruct2026.05 | 93.1 | 88.24 | 97.16 | 93.75 | |
| MAGEBackbone=Qwen2.5-3B-Instruct2026.05 | 91.95 | 88.07 | 97.78 | 92.63 | |
| AgentDoG-Qwen3-4BModel Category=Our Models, Backbone=Qwen3-4B2026.01 | 91.8 | 88 | 98 | 92.7 | |
| AgentDoG-Qwen2.5-7BModel Category=Our Models, Backbone=Qwen2.5-7B2026.01 | 91.7 | 88.2 | 97.3 | 92.5 | |
| GPT-5.2Model Category=Closed-Source2026.01 | 90.8 | 87.1 | 97 | 91.8 | |
| SFTBackbone=Qwen2.5-7B-Instruct2026.05 | 90.8 | 89.36 | 93.34 | 91.31 | |
| QwQ-32BModel Category=Open-Source2026.01 | 89.5 | 95.8 | 83.9 | 89.5 | |
| TRACEBackbone=Qwen3-4B-Instruct-25072026.05 | 87.36 | 96.96 | 88.89 | 87.91 | |
| SFTBackbone=Qwen3-4B-Instruct-25072026.05 | 86.21 | 78.95 | 96.04 | 88.24 | |
| TRACEBackbone=Qwen2.5-3B-Instruct2026.05 | 86.05 | 82.14 | 93.18 | 87.23 | |
| Qwen3-235B-A22B-Instruct-2507Model Category=Open-Source2026.01 | 85.1 | 80.6 | 94.6 | 87 | |
| MAGEBackbone=Qwen3-4B-Instruct-25072026.05 | 85.06 | 88.1 | 82.42 | 85.06 | |
| ShieldAgentModel Category=Guard Models2026.01 | 81 | 74.1 | 98.7 | 84.6 | |
| TRACEBackbone=Qwen2.5-7B-Instruct2026.05 | 80.46 | 75.93 | 92.01 | 82.83 | |
| AgentDoG-Llama3.1-8BModel Category=Our Models, Backbone=Llama3.1-8B2026.01 | 78.2 | 71.6 | 97.3 | 82.5 | |
| MAGEBackbone=Qwen2.5-7B-Instruct2026.05 | 72.41 | 67.21 | 91.26 | 77.36 | |
| Qwen3-4B-Instruct-2507Model Category=Open-Source2026.01 | 68.4 | 73.8 | 62.4 | 67.6 | |
| Qwen2.5-7B-InstructModel Category=Open-Source2026.01 | 68.4 | 77.5 | 56.7 | 65.5 | |
| LlamaGuard4-12BModel Category=Guard Models, Note=Shows high precision but low recall on ATBench2026.01 | 63.8 | 71.9 | 56.4 | 63.2 | |
| LlamaGuard3-8BModel Category=Guard Models, Note=Shows high precision but low recall on ATBench2026.01 | 61.2 | 73 | 46.4 | 56.7 | |
| BaseBackbone=Qwen2.5-7B-Instruct2026.05 | 57.12 | 56.7 | 77.45 | 65.47 | |
| BaseBackbone=Qwen2.5-3B-Instruct2026.05 | 56.76 | 62.93 | 42.16 | 50.49 | |
| BaseBackbone=Qwen3-4B-Instruct-25072026.05 | 55.17 | 62.99 | 26.14 | 36.95 | |
| NemoGuardModel Category=Guard Models2026.01 | 54.4 | 60.2 | 40.6 | 48.5 | |
| PolyGuardModel Category=Guard Models2026.01 | 54.3 | 54.1 | 87.9 | 67 | |
| Llama3.1-8B-InstructModel Category=Open-Source2026.01 | 53.7 | 53.3 | 100 | 69.5 | |
| JoySafetyModel Category=Guard Models2026.01 | 52.5 | 57.1 | 40.3 | 47.2 | |
| ShieldGemma-27BModel Category=Guard Models, Evaluation Protocol=Different method used due to context length constraints2026.01 | 47.7 | 100 | 1 | 2 | |
| ShieldGemma-9BModel Category=Guard Models, Evaluation Protocol=Different method used due to context length constraints2026.01 | 47.7 | 100 | 1 | 2 | |
| Qwen3-GuardModel Category=Guard Models, Note=Shows high precision but low recall on ATBench2026.01 | 40.6 | 25.4 | 5.5 | 9 |