ResearchBenchmarksCode Generation on HumanEval (Final Acc., Correction Uplift)Follow51.22Final AccuracyROSA44.87646.52348.1749.817Sep 27, 2025Evaluation ResultsMethodMethodLinksFinal AccuracyCorrection UpliftROSAModel=DeepSeek-R1-Dist...Model=DeepSeek-R1-Distill-Qwen-8B2025.0951.2233.75BaselineModel=DeepSeek-R1-Dist...Model=DeepSeek-R1-Distill-Qwen-8B2025.0945.1217.05