Root Cause Mapping on CTIBench RCM
0.753ScoreFoundation-Sec-8B-Reasoning
Evaluation Results
| Method | Links | |
|---|---|---|
| Foundation-Sec-8B-ReasoningModel Group=our reasoning model, Number of Parameters=8B, Trials=52026.01 | 0.753 | |
| GPT-4.1Model Group=frontier OpenAI API models, Trials=52026.01 | 0.73 | |
| GPT-5Model Group=frontier OpenAI API models, Trials=52026.01 | 0.728 | |
| GPT-5-MiniModel Group=frontier OpenAI API models, Trials=52026.01 | 0.723 | |
| GPT-OSS-120BModel Group=GPT-OSS models, Number of Parameters=120B, Trials=52026.01 | 0.712 | |
| o3-MiniModel Group=frontier OpenAI API models, Trials=52026.01 | 0.708 | |
| Foundation-Sec-8B-InstructModel Group=Llama-family and cybersecurity-specialized, Number of Parameters=8B, Trials=52026.01 | 0.704 | |
| Llama-3.3-70B-InstructModel Group=Llama-family and cybersecurity-specialized, Number of Parameters=70B, Trials=52026.01 | 0.684 | |
| GPT-5-NanoModel Group=frontier OpenAI API models, Trials=52026.01 | 0.672 | |
| Llama-Primus-MergedModel Group=Llama-family and cybersecurity-specialized, Trials=52026.01 | 0.665 | |
| Llama-Primus-Nemotron-70B-InstructModel Group=Llama-family and cybersecurity-specialized, Number of Parameters=70B, Trials=52026.01 | 0.664 | |
| Phi-4Model Group=smaller specialized models, Trials=52026.01 | 0.629 | |
| Qwen-3-14BModel Group=smaller specialized models, Number of Parameters=14B, Trials=52026.01 | 0.612 | |
| GPT-OSS-20BModel Group=GPT-OSS models, Number of Parameters=20B, Trials=52026.01 | 0.61 | |
| Qwen-3-8BModel Group=smaller specialized models, Number of Parameters=8B, Trials=52026.01 | 0.542 | |
| Llama-3.1-8B-InstructModel Group=Llama-family and cybersecurity-specialized, Number of Parameters=8B, Trials=52026.01 | 0.531 | |
| DeepHat-V1-7BModel Group=smaller specialized models, Number of Parameters=7B, Trials=52026.01 | 0.434 |