Cybersecurity Benchmarking on ScBen En
87.48EnGPT-5
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-5evaluation_context=Larger Instruct and Proprietary Model Evaluation, few_shot_setting=0-shot2026.01 | 87.48 | |
| Qwen3-32Bevaluation_context=Larger Instruct and Proprietary Model Evaluation, few_shot_setting=0-shot2026.01 | 84.23 | |
| RedSage-8B-CFWevaluation_context=Base Model Evaluation, few_shot_setting=5-shot2026.01 | 83.62 | |
| Qwen3-8B-Baseevaluation_context=Base Model Evaluation, few_shot_setting=5-shot2026.01 | 82.84 | |
| RedSage-8B-Baseevaluation_context=Base Model Evaluation, few_shot_setting=5-shot2026.01 | 81.76 | |
| RedSage-8B-Seedevaluation_context=Base Model Evaluation, few_shot_setting=5-shot2026.01 | 81.61 | |
| RedSage-8B-DPOevaluation_context=Instruct Model Evaluation, few_shot_setting=0-shot2026.01 | 80.06 | |
| RedSage-8B-Insevaluation_context=Instruct Model Evaluation, few_shot_setting=0-shot2026.01 | 79.91 | |
| Qwen3-8Bevaluation_context=Instruct Model Evaluation, few_shot_setting=0-shot2026.01 | 73.26 | |
| Llama-3.1-8Bevaluation_context=Base Model Evaluation, few_shot_setting=5-shot2026.01 | 72.8 | |
| DeepHat-V1-7Bevaluation_context=Instruct Model Evaluation, few_shot_setting=0-shot2026.01 | 70.63 | |
| Foundation-Sec-8Bevaluation_context=Base Model Evaluation, few_shot_setting=5-shot2026.01 | 69.86 | |
| Foundation-Sec-8B-Instructevaluation_context=Instruct Model Evaluation, few_shot_setting=0-shot2026.01 | 68.78 | |
| Llama-Primus-Mergedevaluation_context=Instruct Model Evaluation, few_shot_setting=0-shot2026.01 | 64.91 | |
| Llama-Primus-Baseevaluation_context=Instruct Model Evaluation, few_shot_setting=0-shot2026.01 | 63.68 | |
| Llama-3.1-8B-Instructevaluation_context=Instruct Model Evaluation, few_shot_setting=0-shot2026.01 | 59.66 | |
| Lily-Cybersecurity-7B-v0.2evaluation_context=Instruct Model Evaluation, few_shot_setting=0-shot2026.01 | 57.65 |