Code Generation on HumanEval+ (Pass@1 accuracy)
88.9Pass@1 AccuracyThinking onRL
Evaluation Results
| Method | Links | |
|---|---|---|
| Thinking onRLModel Size=Qwen3-32B, Condition=native RL-trained reasoning mode2026.06 | 88.9 | |
| Thinking onRLModel Size=Qwen3-8B, Condition=native RL-trained reasoning mode2026.06 | 87.8 | |
| NeuReasonerModel Size=Qwen3-32B, Condition=our cognitive scaffold, thinking off2026.06 | 87.6 | |
| Thinking offModel Size=Qwen3-32B, Condition=vanilla chain-of-thought, no scaffold2026.06 | 80.3 | |
| NeuReasonerModel Size=Qwen3-8B, Condition=our cognitive scaffold, thinking off2026.06 | 79.7 | |
| Thinking offModel Size=Qwen3-8B, Condition=vanilla chain-of-thought, no scaffold2026.06 | 78.5 |