Multiple Choice Question Answering on CTIBench MCQA
0.819ScoreGPT-5
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-5Model Group=frontier OpenAI API models, Trials=52026.01 | 0.819 | |
| GPT-4.1Model Group=frontier OpenAI API models, Trials=52026.01 | 0.76 | |
| GPT-5-MiniModel Group=frontier OpenAI API models, Trials=52026.01 | 0.753 | |
| o3-MiniModel Group=frontier OpenAI API models, Trials=52026.01 | 0.716 | |
| GPT-OSS-120BModel Group=GPT-OSS models, Number of Parameters=120B, Trials=52026.01 | 0.714 | |
| Llama-Primus-Nemotron-70B-InstructModel Group=Llama-family and cybersecurity-specialized, Number of Parameters=70B, Trials=52026.01 | 0.705 | |
| Llama-3.3-70B-InstructModel Group=Llama-family and cybersecurity-specialized, Number of Parameters=70B, Trials=52026.01 | 0.692 | |
| Foundation-Sec-8B-ReasoningModel Group=our reasoning model, Number of Parameters=8B, Trials=52026.01 | 0.691 | |
| GPT-5-NanoModel Group=frontier OpenAI API models, Trials=52026.01 | 0.688 | |
| Qwen-3-14BModel Group=smaller specialized models, Number of Parameters=14B, Trials=52026.01 | 0.664 | |
| Phi-4Model Group=smaller specialized models, Trials=52026.01 | 0.658 | |
| GPT-OSS-20BModel Group=GPT-OSS models, Number of Parameters=20B, Trials=52026.01 | 0.655 | |
| Foundation-Sec-8B-InstructModel Group=Llama-family and cybersecurity-specialized, Number of Parameters=8B, Trials=52026.01 | 0.65 | |
| Qwen-3-8BModel Group=smaller specialized models, Number of Parameters=8B, Trials=52026.01 | 0.649 | |
| Llama-3.1-8B-InstructModel Group=Llama-family and cybersecurity-specialized, Number of Parameters=8B, Trials=52026.01 | 0.607 | |
| Llama-Primus-MergedModel Group=Llama-family and cybersecurity-specialized, Trials=52026.01 | 0.604 | |
| DeepHat-V1-7BModel Group=smaller specialized models, Number of Parameters=7B, Trials=52026.01 | 0.493 |