Code Generation on LeetCode Contest Benchmark
89.7Easy AccuracyGRPO
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| GRPOModel=DS 8B, Training Context=16K, Testing Context=16K2026.03 | 89.7 | 26 | 1.3 | 35.3 | |
| MicroCoder-GRPOModel=DS 8B, Training Context=16K, Testing Context=16K2026.03 | 89.7 | 32.7 | 8.7 | 40.5 | |
| MicroCoder-GRPOModel=Qwen3 4B, Training Context=4K, Testing Context=4K2026.03 | 88.2 | 22.1 | 3.7 | 34.1 | |
| DAPOModel=DS 8B, Training Context=16K, Testing Context=16K2026.03 | 88.2 | 25 | 2.5 | 34.9 | |
| MicroCoder-GRPOModel=Qwen3 4B, Training Context=4K, Testing Context=8K2026.03 | 83.8 | 24 | 5 | 34.1 | |
| Qwen3-4B-InstructTesting Context=8K, RL Algorithm=None2026.03 | 76.5 | 18.3 | 2.5 | 29 | |
| DAPOModel=Qwen3 4B, Training Context=4K, Testing Context=8K2026.03 | 76.5 | 23.1 | 2.5 | 31 | |
| DAPOModel=Qwen3 4B, Training Context=4K, Testing Context=4K2026.03 | 75 | 22.1 | 1.3 | 29.8 | |
| Qwen3-4B-InstructTesting Context=4K, RL Algorithm=None2026.03 | 73.5 | 19.2 | 2.5 | 28.6 | |
| GPT-4-TurboChain of Thought (CoT)=false2024.01 | 73.3 | 31.9 | 25 | 40.6 | |
| GRPOModel=Qwen3 4B, Training Context=4K, Testing Context=4K2026.03 | 72.1 | 21.2 | 0 | 28.2 | |
| GRPOModel=Qwen3 4B, Training Context=4K, Testing Context=8K2026.03 | 72.1 | 22.1 | 0 | 28.6 | |
| GPT-4-TurboChain of Thought (CoT)=true2024.01 | 71.1 | 35.2 | 25 | 41.8 | |
| CoT (Greedy)Backbone=Qwen3-8B, decoding=Greedy2025.10 | 64.44 | 58.24 | 47.73 | — | |
| SwiRBackbone=Qwen3-8B2025.10 | 64.44 | 69.23 | 61.36 | — | |
| DeepSeek-Coder-InstructSize=33B, Chain of Thought (CoT)=false2024.01 | 57.8 | 22 | 9.1 | 27.8 | |
| CoTBackbone=Qwen3-8B2025.10 | 57.78 | 68.13 | 43.18 | — | |
| MicroCoder-GRPOModel=Qwen3 1.7B, Training Context=4K, Testing Context=4K2026.03 | 57.4 | 4.8 | 0 | 17.5 | |
| DAPOModel=Qwen3 1.7B, Training Context=4K, Testing Context=4K2026.03 | 55.9 | 1.9 | 0 | 15.9 | |
| DAPOModel=Qwen3 1.7B, Training Context=4K, Testing Context=8K2026.03 | 55.9 | 1.9 | 0 | 15.9 | |
| Soft ThinkingBackbone=Qwen3-8B2025.10 | 55.56 | 61.54 | 38.64 | — | |
| DeepSeek-Coder-InstructSize=33B, Chain of Thought (CoT)=true2024.01 | 53.3 | 25.3 | 11.4 | 28.9 | |
| MicroCoder-GRPOModel=Qwen3 1.7B, Training Context=4K, Testing Context=8K2026.03 | 52.9 | 8.7 | 0 | 17.9 | |
| GRPOModel=Qwen3 1.7B, Training Context=4K, Testing Context=4K2026.03 | 47.1 | 4.8 | 0 | 14.7 | |
| GRPOModel=Qwen3 1.7B, Training Context=4K, Testing Context=8K2026.03 | 47.1 | 4.8 | 0 | 14.7 | |
| GPT-3.5-TurboChain of Thought (CoT)=false2024.01 | 46.7 | 15.4 | 15.9 | 23.3 | |
| DeepSeek-Coder-InstructSize=6.7B, Chain of Thought (CoT)=false2024.01 | 44.4 | 12.1 | 9.1 | 19.4 | |
| DeepSeek-Coder-InstructSize=6.7B, Chain of Thought (CoT)=true2024.01 | 44.4 | 17.6 | 4.5 | 21.1 | |
| GPT-3.5-TurboChain of Thought (CoT)=true2024.01 | 42.2 | 15.4 | 20.5 | 23.3 | |
| Qwen3-1.7B InstructTesting Context=4K, RL Algorithm=None2026.03 | 33.8 | 3.8 | 0 | 10.7 | |
| Qwen3-1.7B InstructTesting Context=8K, RL Algorithm=None2026.03 | 33.8 | 3.8 | 0 | 10.7 | |
| Phind-CodeLlama-V2Size=34B2024.01 | 26.7 | 8.8 | 9.1 | 13.3 | |
| CodeLlama-InstructSize=34B2024.01 | 24.4 | 4.4 | 4.5 | 9.4 | |
| DeepSeek-Coder-InstructSize=1.3B, Chain of Thought (CoT)=false2024.01 | 22.2 | 1.1 | 4.5 | 7.2 | |
| DeepSeek-Coder-InstructSize=1.3B, Chain of Thought (CoT)=true2024.01 | 22.2 | 2.2 | 2.3 | 7.2 | |
| WizardCoder-V1.0Size=15B2024.01 | 17.8 | 1.1 | 0 | 5 |