ResearchBenchmarksSecurity-related Issue Classification on In-the-Wild GitHub issue reports 1,000 (test)Follow85.83PrecisionGPT-4.160.048466.741773.43580.1283Dec 17, 2025Evaluation ResultsMethodMethodLinksPrecisionRecallF1-ScoreGPT-4.1Type=LLM-basedType=LLM-based2025.1285.8382.5882.18GPT-4oType=LLM-basedType=LLM-based2025.1285.368180.4SEBERTISType=In-the-Wild, Clas...Type=In-the-Wild, Classifier=BERT MLM2025.1271.2368.667.6LlamaType=LLM-basedType=LLM-based2025.1261.0454.3657.51