Cyber Threat Intelligence Reasoning on CTI-Reasoning
64.3ScoreGPT-5
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-5Model Group=frontier OpenAI API models, Trials=52026.01 | 64.3 | |
| GPT-4.1Model Group=frontier OpenAI API models, Trials=52026.01 | 59.6 | |
| GPT-5-MiniModel Group=frontier OpenAI API models, Trials=52026.01 | 57.8 | |
| Llama-3.3-70B-InstructModel Group=Llama-family and cybersecurity-specialized, Number of Parameters=70B, Trials=52026.01 | 50.7 | |
| GPT-OSS-120BModel Group=GPT-OSS models, Number of Parameters=120B, Trials=52026.01 | 49.6 | |
| Llama-Primus-Nemotron-70B-InstructModel Group=Llama-family and cybersecurity-specialized, Number of Parameters=70B, Trials=52026.01 | 48.5 | |
| o3-MiniModel Group=frontier OpenAI API models, Trials=52026.01 | 47.9 | |
| GPT-OSS-20BModel Group=GPT-OSS models, Number of Parameters=20B, Trials=52026.01 | 46 | |
| Phi-4Model Group=smaller specialized models, Trials=52026.01 | 44.4 | |
| Qwen-3-14BModel Group=smaller specialized models, Number of Parameters=14B, Trials=52026.01 | 44.1 | |
| GPT-5-NanoModel Group=frontier OpenAI API models, Trials=52026.01 | 43.1 | |
| Foundation-Sec-8B-ReasoningModel Group=our reasoning model, Number of Parameters=8B, Trials=52026.01 | 41.1 | |
| Qwen-3-8BModel Group=smaller specialized models, Number of Parameters=8B, Trials=52026.01 | 39.5 | |
| Foundation-Sec-8B-InstructModel Group=Llama-family and cybersecurity-specialized, Number of Parameters=8B, Trials=52026.01 | 36.4 | |
| Llama-Primus-MergedModel Group=Llama-family and cybersecurity-specialized, Trials=52026.01 | 34.8 | |
| Llama-3.1-8B-InstructModel Group=Llama-family and cybersecurity-specialized, Number of Parameters=8B, Trials=52026.01 | 33.5 | |
| DeepHat-V1-7BModel Group=smaller specialized models, Number of Parameters=7B, Trials=52026.01 | 32.3 |