Reasoning-Aware Threat Detection on Reasoning-Aware Security Evaluation Post-Hoc Rationalizations
0.99Det-F1GPT-5.2 + Instructions
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPT-5.2 + InstructionsModel Class=Upper Bound, Detector Variant=GPT-5.2 + Instructions [28]2026.03 | 0.99 | 0.995 | |
| Qwen3-4B-GuardModel Class=TRACEGUARD (Ours), Detector Variant=Qwen3-4B-Guard2026.03 | 0.948 | 0.97 | |
| GPT-OSS-20B-GuardModel Class=TRACEGUARD (Ours), Detector Variant=GPT-OSS-20B-Guard2026.03 | 0.922 | 0.929 | |
| DeepSeek-R1-7B-GuardModel Class=TRACEGUARD (Ours), Detector Variant=DeepSeek-R1-7B-Guard2026.03 | 0.78 | 0.832 | |
| GPT-OSS-20BModel Class=Zero-Shot Baselines, Detector Variant=GPT-OSS-20B [29]2026.03 | 0.56 | 0.595 | |
| Qwen3-4BModel Class=Zero-Shot Baselines, Detector Variant=Qwen3-4B [14]2026.03 | 0.538 | 0.566 | |
| Qwen3-4BModel Class=CoS [22], Detector Variant=Qwen3-4B [14]2026.03 | 0.511 | 0.44 | |
| DeepSeek-R1-7BModel Class=CoS [22], Detector Variant=DeepSeek-R1-7B [3]2026.03 | 0.232 | 0.275 | |
| GPT-OSS-20BModel Class=CoS [22], Detector Variant=GPT-OSS-20B [29]2026.03 | 0.221 | 0.325 | |
| DeepSeek-R1-7BModel Class=Zero-Shot Baselines, Detector Variant=DeepSeek-R1-7B [3]2026.03 | 0.136 | 0.492 | |
| Llama-Guard-3-8BModel Class=Llama-Guard [26], Detector Variant=Llama-Guard-3-8B [26]2026.03 | 0 | — |