Coding on HumanEval, MBPP, BigCodeBench Aggregate
69.9Aggregate Average ScoreHASP-Evolve + RS
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| HASP-Evolve + RSSize=7B-Inst, Training Setting=Training-based, Reference Method (for Delta)=HASP-Evolve + RS, Backbone=Qwen2.5-7B-Instruct2026.05 | 69.9 | — | |
| P-GRPO (Code+RM)Size=7B-Inst, Training Setting=Training-based, Reference Method (for Delta)=HASP-Evolve + RS2026.05 | 69.8 | 0.1 | |
| HASP-Intervention (w. Teacher)Size=7B-Inst, Training Setting=Training-free / inference-time, Reference Method (for Delta)=HASP-Intervention (Infer.), Backbone=Qwen2.5-7B-Instruct2026.05 | 68.7 | — | |
| GRPO (Code)Size=7B-Inst, Training Setting=Training-based, Reference Method (for Delta)=HASP-Evolve + RS2026.05 | 68.4 | 1.5 | |
| KodCode-RL Step 256Size=7B-Inst, Training Setting=Training-based, Reference Method (for Delta)=HASP-Evolve + RS2026.05 | 66.8 | 3.1 | |
| KodCode-RL Step 128Size=7B-Inst, Training Setting=Training-based, Reference Method (for Delta)=HASP-Evolve + RS2026.05 | 65.4 | 4.5 | |
| GPT-4oSize=~200B, Training Setting=Training-free / inference-time, Reference Method (for Delta)=HASP-Intervention (Infer.)2026.05 | 63.8 | 4.8 | |
| HASP-Intervention (PF-only)Size=7B-Inst, Training Setting=Training-free / inference-time, Reference Method (for Delta)=HASP-Intervention (Infer.), Backbone=Qwen2.5-7B-Instruct2026.05 | 63.4 | 5.3 | |
| AceCoderRMSize=7B-Inst, Training Setting=Training-based, Reference Method (for Delta)=HASP-Evolve + RS2026.05 | 63.2 | 6.7 | |
| AceCoderRuleSize=7B-Inst, Training Setting=Training-based, Reference Method (for Delta)=HASP-Evolve + RS2026.05 | 62.1 | 7.8 | |
| Prompt-Only SkillsSize=7B-Inst, Training Setting=Training-free / inference-time, Reference Method (for Delta)=HASP-Intervention (Infer.)2026.05 | 61.2 | 7.5 | |
| GPT-4o-miniSize=~8B, Training Setting=Training-free / inference-time, Reference Method (for Delta)=HASP-Intervention (Infer.)2026.05 | 61 | 7.7 | |
| Qwen2.5-7B-InstructSize=7B-Inst, Training Setting=Training-free / inference-time, Reference Method (for Delta)=HASP-Intervention (Infer.)2026.05 | 60.8 | 7.9 | |
| SFT (vanilla)Size=7B-Inst, Training Setting=Training-based, Reference Method (for Delta)=HASP-Evolve + RS2026.05 | 57.5 | 12.4 | |
| RA-Agent (multi-loop)Size=7B-Inst, Training Setting=Training-free / inference-time, Reference Method (for Delta)=HASP-Intervention (Infer.)2026.05 | 54.5 | 14.2 |