Code Generation on HumanEval+ (Accuracy, Follow Rate)
77.44AccuracyQwen-2.5-7B-Instruct
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qwen-2.5-7B-InstructStage=Base2026.06 | 77.44 | — | |
| Qwen-2.5-7B-InstructStage=Post-trained2026.06 | 76.22 | — | |
| A.X-4.0-LightStage=Base2026.06 | 75 | — | |
| EXAONE-3.5-7.8B-InstructStage=Base2026.06 | 75 | — | |
| Kanana-1.5-8B-InstructStage=Base2026.06 | 75 | — | |
| A.X-4.0-LightStage=Post-trained2026.06 | 75 | — | |
| Kanana-1.5-8B-InstructStage=Post-trained2026.06 | 75 | — | |
| EXAONE-3.5-7.8B-InstructStage=Post-trained2026.06 | 73.78 | — | |
| Gemma-3-4B-ITStage=Base2026.06 | 61.59 | — | |
| Gemma-3-4B-ITStage=Post-trained2026.06 | 61.59 | — | |
| Llama-3.1-8B-InstructStage=Post-trained2026.06 | 59.15 | — | |
| Llama-3.1-8B-InstructStage=Base2026.06 | 58.54 | — |