Overall Cybersecurity Performance on Cybersecurity Multi-Benchmark Suite
86.29Overall Mean ScoreGPT-5
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-5evaluation_context=Larger Instruct and Proprietary Model Evaluation, few_shot_setting=0-shot2026.01 | 86.29 | |
| RedSage-8B-Baseevaluation_context=Base Model Evaluation, few_shot_setting=5-shot2026.01 | 84.56 | |
| RedSage-8B-Seedevaluation_context=Base Model Evaluation, few_shot_setting=5-shot2026.01 | 84.45 | |
| RedSage-8B-CFWevaluation_context=Base Model Evaluation, few_shot_setting=5-shot2026.01 | 82.66 | |
| Qwen3-32Bevaluation_context=Larger Instruct and Proprietary Model Evaluation, few_shot_setting=0-shot2026.01 | 82.31 | |
| RedSage-8B-Insevaluation_context=Instruct Model Evaluation, few_shot_setting=0-shot2026.01 | 81.3 | |
| RedSage-8B-DPOevaluation_context=Instruct Model Evaluation, few_shot_setting=0-shot2026.01 | 81.1 | |
| Qwen3-8B-Baseevaluation_context=Base Model Evaluation, few_shot_setting=5-shot2026.01 | 80.81 | |
| Foundation-Sec-8Bevaluation_context=Base Model Evaluation, few_shot_setting=5-shot2026.01 | 76.9 | |
| Qwen3-8Bevaluation_context=Instruct Model Evaluation, few_shot_setting=0-shot2026.01 | 75.71 | |
| Llama-3.1-8Bevaluation_context=Base Model Evaluation, few_shot_setting=5-shot2026.01 | 75.44 | |
| DeepHat-V1-7Bevaluation_context=Instruct Model Evaluation, few_shot_setting=0-shot2026.01 | 75.44 | |
| Foundation-Sec-8B-Instructevaluation_context=Instruct Model Evaluation, few_shot_setting=0-shot2026.01 | 75.44 | |
| Llama-Primus-Baseevaluation_context=Instruct Model Evaluation, few_shot_setting=0-shot2026.01 | 71.69 | |
| Llama-Primus-Mergedevaluation_context=Instruct Model Evaluation, few_shot_setting=0-shot2026.01 | 71.23 | |
| Llama-3.1-8B-Instructevaluation_context=Instruct Model Evaluation, few_shot_setting=0-shot2026.01 | 68.52 | |
| Lily-Cybersecurity-7B-v0.2evaluation_context=Instruct Model Evaluation, few_shot_setting=0-shot2026.01 | 55.74 |