Attack Technique Extraction on CTIBench ATE
69.6Micro F1GPT-4.1
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-4.1Model Group=frontier OpenAI API models, Trials=52026.01 | 69.6 | |
| GPT-5-MiniModel Group=frontier OpenAI API models, Trials=52026.01 | 68.1 | |
| o3-MiniModel Group=frontier OpenAI API models, Trials=52026.01 | 59.9 | |
| GPT-5Model Group=frontier OpenAI API models, Trials=52026.01 | 57.8 | |
| Llama-3.3-70B-InstructModel Group=Llama-family and cybersecurity-specialized, Number of Parameters=70B, Trials=52026.01 | 51.9 | |
| Qwen-3-14BModel Group=smaller specialized models, Number of Parameters=14B, Trials=52026.01 | 50.2 | |
| Foundation-Sec-8B-ReasoningModel Group=our reasoning model, Number of Parameters=8B, Trials=52026.01 | 49.1 | |
| GPT-OSS-20BModel Group=GPT-OSS models, Number of Parameters=20B, Trials=52026.01 | 47.8 | |
| GPT-5-NanoModel Group=frontier OpenAI API models, Trials=52026.01 | 45.3 | |
| Phi-4Model Group=smaller specialized models, Trials=52026.01 | 43.5 | |
| Qwen-3-8BModel Group=smaller specialized models, Number of Parameters=8B, Trials=52026.01 | 40.8 | |
| Foundation-Sec-8B-InstructModel Group=Llama-family and cybersecurity-specialized, Number of Parameters=8B, Trials=52026.01 | 35.8 | |
| GPT-OSS-120BModel Group=GPT-OSS models, Number of Parameters=120B, Trials=52026.01 | 28.2 | |
| Llama-Primus-Nemotron-70B-InstructModel Group=Llama-family and cybersecurity-specialized, Number of Parameters=70B, Trials=52026.01 | 26.8 | |
| Llama-3.1-8B-InstructModel Group=Llama-family and cybersecurity-specialized, Number of Parameters=8B, Trials=52026.01 | 13.2 | |
| Llama-Primus-MergedModel Group=Llama-family and cybersecurity-specialized, Trials=52026.01 | 5.8 | |
| DeepHat-V1-7BModel Group=smaller specialized models, Number of Parameters=7B, Trials=52026.01 | 0.4 |