Cybersecurity Knowledge Evaluation on CyMtc (500)
95.6CyMtc (500) ScoreGPT-5
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-5evaluation_context=Larger Instruct and Proprietary Model Evaluation, few_shot_setting=0-shot2026.01 | 95.6 | |
| RedSage-8B-CFWevaluation_context=Base Model Evaluation, few_shot_setting=5-shot2026.01 | 93.8 | |
| RedSage-8B-Baseevaluation_context=Base Model Evaluation, few_shot_setting=5-shot2026.01 | 92.6 | |
| RedSage-8B-Seedevaluation_context=Base Model Evaluation, few_shot_setting=5-shot2026.01 | 92.2 | |
| Qwen3-8B-Baseevaluation_context=Base Model Evaluation, few_shot_setting=5-shot2026.01 | 92 | |
| Qwen3-32Bevaluation_context=Larger Instruct and Proprietary Model Evaluation, few_shot_setting=0-shot2026.01 | 91.8 | |
| RedSage-8B-DPOevaluation_context=Instruct Model Evaluation, few_shot_setting=0-shot2026.01 | 90 | |
| RedSage-8B-Insevaluation_context=Instruct Model Evaluation, few_shot_setting=0-shot2026.01 | 89.8 | |
| Qwen3-8Bevaluation_context=Instruct Model Evaluation, few_shot_setting=0-shot2026.01 | 88.6 | |
| Foundation-Sec-8Bevaluation_context=Base Model Evaluation, few_shot_setting=5-shot2026.01 | 86.6 | |
| DeepHat-V1-7Bevaluation_context=Instruct Model Evaluation, few_shot_setting=0-shot2026.01 | 86 | |
| Llama-3.1-8Bevaluation_context=Base Model Evaluation, few_shot_setting=5-shot2026.01 | 84.2 | |
| Llama-Primus-Mergedevaluation_context=Instruct Model Evaluation, few_shot_setting=0-shot2026.01 | 83.8 | |
| Llama-Primus-Baseevaluation_context=Instruct Model Evaluation, few_shot_setting=0-shot2026.01 | 83.8 | |
| Foundation-Sec-8B-Instructevaluation_context=Instruct Model Evaluation, few_shot_setting=0-shot2026.01 | 83 | |
| Llama-3.1-8B-Instructevaluation_context=Instruct Model Evaluation, few_shot_setting=0-shot2026.01 | 82.8 | |
| Lily-Cybersecurity-7B-v0.2evaluation_context=Instruct Model Evaluation, few_shot_setting=0-shot2026.01 | 65.2 |