Code Generation on MBPP (Acc. %, RSR %)
96.6AccuracyEG-CFG
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| EG-CFGModel=DeepSeek-V3-03242025.06 | 96.6 | 80.23 | |
| QualityFlowModel=Claude-Sonnet-3.52025.06 | 94.2 | — | |
| Baseline LLMModel=Claude-Sonnet-3.52025.06 | 88.7 | — | |
| MetaGPTModel=GPT-42025.06 | 87.7 | — | |
| MapCoderModel=DeepSeek-V3-03242025.06 | 87.2 | 25.58 | |
| MGDebuggerModel=DeepSeek-V3-03242025.06 | 86.8 | 23.25 | |
| LPWModel=GPT-4o2025.06 | 84.8 | — | |
| LPWModel=DeepSeek-V3-03242025.06 | 84 | 6.97 | |
| EG-CFGModel=DeepSeek-Coder 1.3B2025.06 | 83.2 | 66.79 | |
| MapCoderModel=GPT-42025.06 | 83.1 | — | |
| Baseline LLMModel=DeepSeek-V3-03242025.06 | 82.8 | 0 | |
| MGDebuggerModel=CodeQwen1.52025.06 | 80.8 | — | |
| Self-DebuggingModel=GPT-42025.06 | 80.6 | — | |
| MGDebuggerModel=DeepSeek-Coder-V2-Lite2025.06 | 80 | — | |
| Self-CollaborationModel=GPT-42025.06 | 78.9 | — | |
| DenserBackbone=Qwen3-32B-think2025.12 | 74.3 | — | |
| Process SupervisionBackbone=Qwen3-32B-think2025.12 | 73.8 | — | |
| Reflection-CoTBackbone=Qwen3-32B-think2025.12 | 73.2 | — | |
| Self-VerificationBackbone=Qwen3-32B-think2025.12 | 72.6 | — | |
| DenserBackbone=Qwen3-32B2025.12 | 72.5 | — | |
| Tree-of-ThoughtBackbone=Qwen3-32B-think2025.12 | 71.9 | — | |
| Process SupervisionBackbone=Qwen3-32B2025.12 | 71.7 | — | |
| Self-ConsistencyBackbone=Qwen3-32B-think2025.12 | 71.2 | — | |
| Reflection-CoTBackbone=Qwen3-32B2025.12 | 71.1 | — | |
| Self-VerificationBackbone=Qwen3-32B2025.12 | 70.8 | — | |
| Think-to-ThinkBackbone=Qwen3-32B-think2025.12 | 70.5 | — | |
| MGDebuggerModel=DeepSeek-Coder 1.3B2025.06 | 70.4 | 41.5 | |
| Tree-of-ThoughtBackbone=Qwen3-32B2025.12 | 69.8 | — | |
| Self-ConsistencyBackbone=Qwen3-32B2025.12 | 69.1 | — | |
| Think-to-ThinkBackbone=Qwen3-32B2025.12 | 68.4 | — | |
| Chain-of-ThoughtBackbone=Qwen3-32B-think2025.12 | 68.4 | — | |
| Baseline LLMModel=GPT-42025.06 | 68.3 | — | |
| DenserBackbone=Qwen3-8B-think2025.12 | 66.8 | — | |
| Chain-of-ThoughtBackbone=Qwen3-32B2025.12 | 66.2 | — | |
| DenserBackbone=Qwen3-14B-no-think2025.12 | 65.9 | — | |
| Process SupervisionBackbone=Qwen3-8B-think2025.12 | 65.6 | — | |
| Reflection-CoTBackbone=Qwen3-8B-think2025.12 | 65 | — | |
| Process SupervisionBackbone=Qwen3-14B-no-think2025.12 | 64.9 | — | |
| DenserBackbone=Qwen3-8B2025.12 | 64.7 | — | |
| Reflection-CoTBackbone=Qwen3-14B-no-think2025.12 | 64.3 | — | |
| Self-VerificationBackbone=Qwen3-8B-think2025.12 | 64.2 | — | |
| Process SupervisionBackbone=Qwen3-8B2025.12 | 63.7 | — | |
| Tree-of-ThoughtBackbone=Qwen3-8B-think2025.12 | 63.7 | — | |
| Self-VerificationBackbone=Qwen3-14B-no-think2025.12 | 63.7 | — | |
| Reflection-CoTBackbone=Qwen3-8B2025.12 | 63.1 | — | |
| Self-ConsistencyBackbone=Qwen3-8B-think2025.12 | 63.1 | — | |
| Tree-of-ThoughtBackbone=Qwen3-14B-no-think2025.12 | 63 | — | |
| DenserBackbone=Qwen3-4B-think2025.12 | 62.5 | — | |
| Self-VerificationBackbone=Qwen3-8B2025.12 | 62.4 | — | |
| Think-to-ThinkBackbone=Qwen3-8B-think2025.12 | 62.4 | — | |
| Self-ConsistencyBackbone=Qwen3-14B-no-think2025.12 | 62.4 | — | |
| Tree-of-ThoughtBackbone=Qwen3-8B2025.12 | 61.8 | — | |
| Think-to-ThinkBackbone=Qwen3-14B-no-think2025.12 | 61.7 | — | |
| Process SupervisionBackbone=Qwen3-4B-think2025.12 | 61.5 | — | |
| Self-ConsistencyBackbone=Qwen3-8B2025.12 | 61.2 | — | |
| Reflection-CoTBackbone=Qwen3-4B-think2025.12 | 60.9 | — | |
| Think-to-ThinkBackbone=Qwen3-8B2025.12 | 60.5 | — | |
| Chain-of-ThoughtBackbone=Qwen3-8B-think2025.12 | 60.5 | — | |
| Self-VerificationBackbone=Qwen3-4B-think2025.12 | 60.3 | — | |
| Chain-of-ThoughtBackbone=Qwen3-14B-no-think2025.12 | 59.8 | — | |
| Tree-of-ThoughtBackbone=Qwen3-4B-think2025.12 | 59.7 | — | |
| DenserBackbone=Qwen3-4B2025.12 | 59.4 | — | |
| Process SupervisionBackbone=Qwen3-4B2025.12 | 59.1 | — | |
| Self-ConsistencyBackbone=Qwen3-4B-think2025.12 | 59.1 | — | |
| Chain-of-ThoughtBackbone=Qwen3-8B2025.12 | 58.7 | — | |
| Reflection-CoTBackbone=Qwen3-4B2025.12 | 58.4 | — | |
| Think-to-ThinkBackbone=Qwen3-4B-think2025.12 | 58.4 | — | |
| Self-VerificationBackbone=Qwen3-4B2025.12 | 57.8 | — | |
| Tree-of-ThoughtBackbone=Qwen3-4B2025.12 | 57.2 | — | |
| Chain-of-ThoughtBackbone=Qwen3-4B-think2025.12 | 56.8 | — | |
| Self-ConsistencyBackbone=Qwen3-4B2025.12 | 56.7 | — | |
| Think-to-ThinkBackbone=Qwen3-4B2025.12 | 55.8 | — | |
| MapCoderModel=DeepSeek-Coder 1.3B2025.06 | 55.2 | 11.46 | |
| Chain-of-ThoughtBackbone=Qwen3-4B2025.12 | 54.3 | — | |
| ReMiTModel Family=SmolLM3-3B, Mid-Training=True, Few-shot=True2026.02 | 49.6 | — | |
| Baseline LLMModel=DeepSeek-Coder 1.3B2025.06 | 49.4 | 0 | |
| MiniPLMModel Family=SmolLM3-3B, Mid-Training=True, Few-shot=True2026.02 | 48 | — | |
| Vanilla NTPModel Family=SmolLM3-3B, Mid-Training=True, Few-shot=True2026.02 | 47.6 | — | |
| MiniPLMModel Family=Youtu-LLM-2B, Mid-Training=True, Few-shot=True2026.02 | 47 | — | |
| ReMiTModel Family=Youtu-LLM-2B, Mid-Training=True, Few-shot=True2026.02 | 47 | — | |
| Vanilla NTPModel Family=Youtu-LLM-2B, Mid-Training=True, Few-shot=True2026.02 | 46.6 | — | |
| RHO-1Model Family=SmolLM3-3B, Mid-Training=True, Few-shot=True2026.02 | 43.8 | — | |
| RHO-1Model Family=Youtu-LLM-2B, Mid-Training=True, Few-shot=True2026.02 | 39.2 | — | |
| Pre-TrainedModel Family=Youtu-LLM-2B, Mid-Training=Standard Checkpoint, Few-shot=True2026.02 | 38.8 | — | |
| Pre-TrainedModel Family=SmolLM3-3B, Mid-Training=Standard Checkpoint, Few-shot=True2026.02 | 37.4 | — | |
| ReMiTModel Family=OLMo-1B, Mid-Training=True, Few-shot=True2026.02 | 9.2 | — | |
| MiniPLMModel Family=OLMo-1B, Mid-Training=True, Few-shot=True2026.02 | 6.8 | — | |
| RHO-1Model Family=OLMo-1B, Mid-Training=True, Few-shot=True2026.02 | 6.2 | — | |
| Pre-TrainedModel Family=OLMo-1B, Mid-Training=Standard Checkpoint, Few-shot=True2026.02 | 4.8 | — | |
| Vanilla NTPModel Family=OLMo-1B, Mid-Training=True, Few-shot=True2026.02 | 4.6 | — |