Code Generation on MBPP (Pass@1, Pass@10, Pass@100)
80.4Pass@1Ouro 2.6B
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Ouro 2.6BModel Category=Looped Latent Reasoning Models2025.10 | 80.4 | — | — | — | |
| OpenCoderModel Category=AR Coding Models2025.10 | 79.9 | — | — | — | |
| GAC + Token-φseeds=32026.05 | 78.8 | — | — | — | |
| EntropyModel=WeDLM-8B, Default decoding=true2026.04 | 78.4 | — | — | 91.6 | |
| GAC w/o φseeds=32026.05 | 78 | — | — | — | |
| IMHModel=WeDLM-8B2026.04 | 77.6 | — | — | 93.8 | |
| HPTseeds=32026.05 | 76 | — | — | — | |
| LUFFYseeds=32026.05 | 75.8 | — | — | — | |
| KL-ctrlseeds=32026.05 | 75.8 | — | — | — | |
| CHORDseeds=32026.05 | 75.4 | — | — | — | |
| Nash-MTLseeds=32026.05 | 75.3 | — | — | — | |
| Diffu-CoderModel Category=Masked Diffusion Models2025.10 | 75.1 | — | — | — | |
| SRFTseeds=32026.05 | 74.9 | — | — | — | |
| SFT-best + RLseeds=32026.05 | 73.8 | — | — | — | |
| SFT-bestseeds=32026.05 | 71.2 | — | — | — | |
| ConfidenceModel=WeDLM-8B2026.04 | 70.7 | — | — | 83.4 | |
| DreamModel Category=Masked Diffusion Models2025.10 | 68.7 | — | — | — | |
| Qwen2.5-7B-Inst.seeds=32026.05 | 68.4 | — | — | — | |
| LaDiRModel Category=Reasoning Methods2025.10 | 66.8 | — | — | — | |
| TaH+Model Category=Reasoning Methods2025.10 | 65.6 | — | — | — | |
| Soft ThinkingModel Category=Reasoning Methods2025.10 | 64.2 | — | — | — | |
| AR SFTModel Category=Reasoning Methods2025.10 | 63.2 | — | — | — | |
| Qwen 2.5 Coder 7BModel Category=AR Coding Models2025.10 | 61.6 | — | — | — | |
| Dream-7B-Base + DyStructBackbone=Dream-7B, Decoding method=DyStruct, Few-shot examples=3, Unmasking iterations=256, Generation limit=2562026.05 | 59.8 | — | — | — | |
| MarginModel=WeDLM-8B2026.04 | 58.9 | — | — | 76.5 | |
| Dream-7B-BaseBackbone=Dream-7B, Decoding method=Base, Few-shot examples=3, Unmasking iterations=256, Generation limit=2562026.05 | 57.2 | — | — | — | |
| DreamModel size category=Reference Models (≥7B), Number of parameters=7B2025.09 | 55.4 | 56.2 | — | — | |
| Dream-7B-Base + DAEDALBackbone=Dream-7B, Decoding method=DAEDAL, Few-shot examples=3, Unmasking iterations=256, Generation limit=2562026.05 | 54.4 | — | — | — | |
| LLaDAModel Category=Masked Diffusion Models2025.10 | 50.1 | — | — | — | |
| ConfidenceModel=LLaDA-8B, Default decoding=true2026.04 | 48.4 | — | — | 64.5 | |
| IMHModel=LLaDA-8B2026.04 | 46.9 | — | — | 77.8 | |
| MarginModel=LLaDA-8B2026.04 | 45.1 | — | — | 71.6 | |
| EntropyModel=LLaDA-8B2026.04 | 44.4 | — | — | 75.1 | |
| LLaDA-8B-Base + DyStructBackbone=LLaDA-8B, Decoding method=DyStruct, Few-shot examples=3, Unmasking iterations=256, Generation limit=2562026.05 | 41.4 | — | — | — | |
| LLaDA-8B-Base + DAEDALBackbone=LLaDA-8B, Decoding method=DAEDAL, Few-shot examples=3, Unmasking iterations=256, Generation limit=2562026.05 | 40.2 | — | — | — | |
| LLaDA-8B-BaseBackbone=LLaDA-8B, Decoding method=Base, Few-shot examples=3, Unmasking iterations=256, Generation limit=2562026.05 | 39.8 | — | — | — | |
| LLaDAModel size category=Reference Models (≥7B), Number of parameters=8B2025.09 | 38.8 | 53.4 | — | — | |
| RandomModel=LLaDA-8B2026.04 | 27.9 | — | — | 68.5 | |
| AutoregressiveModel size category=Compact Models (≤1B), Number of parameters=1.3B2025.09 | 25.6 | 45.4 | — | — | |
| PASERPruning Scheme=LLM-Pruner (25%)2025.02 | 23.1 | 42.6 | 62 | — | |
| PASERPruning Scheme=SliceGPT (25%)2025.02 | 22.3 | 41 | 63.7 | — | |
| NuggetsPruning Scheme=LLM-Pruner (25%)2025.02 | 18.9 | 35.8 | 52.4 | — | |
| IFDPruning Scheme=LLM-Pruner (25%)2025.02 | 18.2 | 34.5 | 50.7 | — | |
| NuggetsPruning Scheme=SliceGPT (25%)2025.02 | 17.6 | 32.6 | 51.9 | — | |
| DLMModel size category=Compact Models (≤1B), Number of parameters=0.5B2025.09 | 17.6 | 32.6 | — | — | |
| DLM + PAPLModel size category=Compact Models (≤1B), Number of parameters=0.5B2025.09 | 16.7 | 38.4 | — | — | |
| IFDPruning Scheme=SliceGPT (25%)2025.02 | 16.4 | 30.8 | 52.4 | — | |
| Instruction MiningPruning Scheme=LLM-Pruner (25%)2025.02 | 15.7 | 29.2 | 43.1 | — | |
| Full DataPruning Scheme=LLM-Pruner (25%)2025.02 | 15.2 | 26.3 | 39.5 | — | |
| RandomPruning Scheme=LLM-Pruner (25%)2025.02 | 12.8 | 23.8 | 35 | — | |
| Instruction MiningPruning Scheme=SliceGPT (25%)2025.02 | 12.8 | 24.5 | 45.2 | — | |
| Edit FlowModel size category=Compact Models (≤1B), Number of parameters=1.3B2025.09 | 10 | 36.4 | — | — | |
| Uniform X0 + Edit FlowModel size category=Compact Models (≤1B), Number of parameters=1.3B2025.09 | 9.4 | 33.4 | — | — | |
| Full DataPruning Scheme=SliceGPT (25%)2025.02 | 9.3 | 19.7 | 36.4 | — | |
| RandomPruning Scheme=SliceGPT (25%)2025.02 | 8.5 | 17.2 | 38.6 | — | |
| w/o TrainingPruning Scheme=LLM-Pruner (25%)2025.02 | 7.1 | 15.5 | 24.6 | — | |
| w/o TrainingPruning Scheme=SliceGPT (25%)2025.02 | 6.2 | 11.8 | 21.5 | — | |
| Mask DFMModel size category=Compact Models (≤1B), Number of parameters=1.3B2025.09 | 6.2 | 25 | — | — |