Overrefusal Evaluation on OrBench-H
99.85RRDb as Alpaca
Evaluation Results
| Method | Links | |
|---|---|---|
| Db as AlpacaSetting / Model=P-SFT, Llama22026.03 | 99.85 | |
| Db as AlpacaSetting / Model=RLVR, Llama3-U2026.03 | 98.48 | |
| AutoSteerBackbone=LLaVA-1.5-7B2025.10 | 89.3 | |
| DEEPSEEK R1 DISTILL LLAMABackbone=DEEPSEEK R1 DISTILL LLAMA2026.03 | 84.61 | |
| Low-Rank CombinationBackbone=DEEPSEEK R1 DISTILL LLAMA, Method=Low-Rank Combination2026.03 | 77.1 | |
| FigStepBackbone=LLaVA-1.5-7B2025.10 | 72.6 | |
| LLAMA 3 8B INSTRUCTBackbone=LLAMA 3 8B INSTRUCT2026.03 | 60.58 | |
| Low-Rank CombinationBackbone=LLAMA 3 8B INSTRUCT, Method=Low-Rank Combination2026.03 | 57.92 | |
| Db as Our DataSetting / Model=RLVR, Llama3-U2026.03 | 57.09 | |
| Db as Our DataSetting / Model=P-SFT, Llama22026.03 | 55.34 | |
| Nemotron-4BGuardrail Type=Generative, Protected LLM=GPT-4o2026.06 | 53 | |
| w/o Future-Aware ReasoningGuardrail Type=Stream, Protected LLM=GPT-4o, Future-Aware Reasoning=false, Safety-Aligned Optimization=true2026.06 | 50.42 | |
| Qwen-Stream-0.6BGuardrail Type=Stream, Protected LLM=GPT-4o2026.06 | 49.86 | |
| Qwen-Gen-0.6B Prompt OnlyGuardrail Type=Generative, Input Context=Prompt Only, Protected LLM=GPT-4o2026.06 | 47.77 | |
| CoCABackbone=LLaVA-1.5-7B2025.10 | 37.5 | |
| ETABackbone=LLaVA-1.5-7B2025.10 | 35.9 | |
| Qwen-Gen-0.6B Prompt+ResponseGuardrail Type=Generative, Input Context=Prompt+Response, Protected LLM=GPT-4o2026.06 | 31.37 | |
| MoRASBackbone=LLaVA-1.5-7B2025.10 | 27.7 | |
| ECSOBackbone=LLaVA-1.5-7B2025.10 | 25.2 | |
| ASTRABackbone=LLaVA-1.5-7B2025.10 | 24.2 | |
| REFUSE-LLAMABackbone=REFUSE-LLAMA2026.03 | 23.88 | |
| VanillaBackbone=LLaVA-1.5-7B2025.10 | 23.4 | |
| YuFeng-XGuard-0.6B Prompt OnlyGuardrail Type=Generative, Input Context=Prompt Only, Protected LLM=GPT-4o2026.06 | 22.81 | |
| Qwen-Stream-8BGuardrail Type=Stream, Protected LLM=GPT-4o2026.06 | 22.66 | |
| BaselineSetting / Model=RLVR, Llama3-U2026.03 | 16.83 | |
| YuFeng-XGuard-0.6B Prompt+ResponseGuardrail Type=Generative, Input Context=Prompt+Response, Protected LLM=GPT-4o2026.06 | 15.74 | |
| Low-Rank CombinationBackbone=REFUSE-LLAMA, Method=Low-Rank Combination2026.03 | 12.89 | |
| w/o Safety-Aligned OptimizationGuardrail Type=Stream, Protected LLM=GPT-4o, Future-Aware Reasoning=true, Safety-Aligned Optimization=false2026.06 | 11.66 | |
| FreoStreamGuardrail Type=Stream, Protected LLM=GPT-4o, Future-Aware Reasoning=true, Safety-Aligned Optimization=true2026.06 | 11.56 | |
| GuardReasoner-1BGuardrail Type=Generative, Protected LLM=GPT-4o2026.06 | 10.17 | |
| Categorical SteeringBackbone=REFUSE-LLAMA, Method=Categorical Steering2026.03 | 5.84 | |
| BaselineSetting / Model=P-SFT, Llama22026.03 | 4.93 |