Cybersecurity Knowledge Question Answering on MMLU CSec
88CSec ScoreRedSage-8B-Seed
Evaluation Results
| Method | Links | |
|---|---|---|
| RedSage-8B-Seedevaluation_context=Base Model Evaluation, few_shot_setting=5-shot2026.01 | 88 | |
| RedSage-8B-Baseevaluation_context=Base Model Evaluation, few_shot_setting=5-shot2026.01 | 87 | |
| RedSage-8B-CFWevaluation_context=Base Model Evaluation, few_shot_setting=5-shot2026.01 | 86 | |
| GPT-5evaluation_context=Larger Instruct and Proprietary Model Evaluation, few_shot_setting=0-shot2026.01 | 86 | |
| LLM-based QA agent with external memoryModel Backend=GPT 5.4 Mini, Access Type=Closed-source2026.06 | 85.34 | |
| LLM-based QA agent with external memoryModel Backend=GPT-4o Mini, Access Type=Closed-source2026.06 | 85.34 | |
| Qwen3-32Bevaluation_context=Larger Instruct and Proprietary Model Evaluation, few_shot_setting=0-shot2026.01 | 84 | |
| Llama-3.1-8Bevaluation_context=Base Model Evaluation, few_shot_setting=5-shot2026.01 | 83 | |
| Qwen3-8B-Baseevaluation_context=Base Model Evaluation, few_shot_setting=5-shot2026.01 | 83 | |
| Foundation-Sec-8Bevaluation_context=Base Model Evaluation, few_shot_setting=5-shot2026.01 | 80 | |
| Llama-Primus-Baseevaluation_context=Instruct Model Evaluation, few_shot_setting=0-shot2026.01 | 79 | |
| RedSage-8B-DPOevaluation_context=Instruct Model Evaluation, few_shot_setting=0-shot2026.01 | 79 | |
| RedSage-8B-Insevaluation_context=Instruct Model Evaluation, few_shot_setting=0-shot2026.01 | 78 | |
| LLM-based QA agent with external memoryModel Backend=Gemma2-9B, Access Type=Open-source2026.06 | 77.59 | |
| Llama-Primus-Mergedevaluation_context=Instruct Model Evaluation, few_shot_setting=0-shot2026.01 | 76 | |
| Foundation-Sec-8B-Instructevaluation_context=Instruct Model Evaluation, few_shot_setting=0-shot2026.01 | 76 | |
| Qwen3-8Bevaluation_context=Instruct Model Evaluation, few_shot_setting=0-shot2026.01 | 76 | |
| DeepHat-V1-7Bevaluation_context=Instruct Model Evaluation, few_shot_setting=0-shot2026.01 | 74 | |
| Llama-3.1-8B-Instructevaluation_context=Instruct Model Evaluation, few_shot_setting=0-shot2026.01 | 72 | |
| Lily-Cybersecurity-7B-v0.2evaluation_context=Instruct Model Evaluation, few_shot_setting=0-shot2026.01 | 68 | |
| LLM-based QA agent with external memoryModel Backend=Phi3-14B, Access Type=Open-source2026.06 | 64.66 |