Policy Following on POLICYSHIFTBENCH Shift
90AccuracyHuman Performance
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Human Performance2026.07 | 90 | 91.1 | — | |
| Gemini-3-Flash-PreviewTime (ms)=5963.52026.07 | 78.3 | 74.2 | 55.9 | |
| PolicyShiftGuard-7BTime (ms)=163.92026.07 | 69.9 | 67 | 70.4 | |
| GPT-5.4Time (ms)=1037.52026.07 | 68.9 | 57.8 | 48.4 | |
| Qwen3.5-35B-A3BThinking mode=true, Time (ms)=13099.42026.07 | 63.2 | 41.3 | 23.3 | |
| Qwen3.5-4BThinking mode=true, Time (ms)=20292.22026.07 | 61.4 | 25.3 | 9.3 | |
| PolicyShiftGuard-3BTime (ms)=128.52026.07 | 61.1 | 47.8 | 50.5 | |
| Qwen2.5-VL-72BTime (ms)=540.32026.07 | 58 | 37.1 | 22.6 | |
| Qwen3.5-35B-A3BThinking mode=false, Time (ms)=118.52026.07 | 54.7 | 32.5 | 11.8 | |
| Claude-Sonnet-4.6Time (ms)=1065.72026.07 | 53.5 | 27 | 11.8 | |
| Qwen2.5-VL-32BTime (ms)=358.42026.07 | 53.1 | 35.1 | 2.7 | |
| GuardReasoner-VL-3BTime (ms)=2084.52026.07 | 52.3 | 54.7 | 2.2 | |
| Qwen3.5-4BThinking mode=false, Time (ms)=89.22026.07 | 51.8 | 35.7 | 5.8 | |
| Qwen2.5-VL-3BTime (ms)=231.12026.07 | 51.7 | 56.1 | 2.7 | |
| GuardReasoner-VL-7BTime (ms)=3300.32026.07 | 51.6 | 54.2 | 10.2 | |
| SafeGuard-VL-RL-7BTime (ms)=154.42026.07 | 51 | 39.5 | 1.6 | |
| Qwen3.5-0.8BTime (ms)=42.92026.07 | 50.7 | 4.3 | 1.1 | |
| Qwen2.5-VL-7BTime (ms)=273.32026.07 | 50.7 | 12.1 | 0 | |
| Llama Guard-4-12BTime (ms)=277.92026.07 | 50.7 | 9.9 | 1.1 | |
| ShieldGemma2-4BTime (ms)=134.02026.07 | 50.2 | 45.4 | 4.3 | |
| QwenGuard-7BTime (ms)=211.62026.07 | 49.9 | 39.5 | 9.1 | |
| Qwen3.5-2BTime (ms)=57.32026.07 | 49 | 0 | 0 |