Code Generation on LiveCodeBench v5
72.1AccuracyBF16
Evaluation Results
| Method | Links | |
|---|---|---|
| BF16Precision=BF16, Backbone=Nemotron 3 Nano2026.01 | 72.1 | |
| NVFP4 PTQPrecision=NVFP4, Type=PTQ, Backbone=Nemotron 3 Nano2026.01 | 68.9 | |
| NVFP4 QADPrecision=NVFP4, Type=QAD, Backbone=Nemotron 3 Nano2026.01 | 68.9 | |
| NVFP4 QATPrecision=NVFP4, Type=QAT, Backbone=Nemotron 3 Nano2026.01 | 62 | |
| DRIFTBase Model=OpenR1-Distill-7B2026.06 | 50.21 | |
| IF (Impl. w. GraSS)Base Model=OpenR1-Distill-7B2026.06 | 48.18 | |
| QuratingBase Model=OpenR1-Distill-7B2026.06 | 47.63 | |
| RDSBase Model=OpenR1-Distill-7B2026.06 | 45.19 | |
| Random BaselineBase Model=OpenR1-Distill-7B2026.06 | 44.81 | |
| DSIRBase Model=OpenR1-Distill-7B2026.06 | 44.46 | |
| BM25Base Model=OpenR1-Distill-7B2026.06 | 44.19 | |
| LESSBase Model=OpenR1-Distill-7B2026.06 | 44.1 | |
| OpenR1-Distill-7BBase Model=OpenR1-Distill-7B2026.06 | 43.86 | |
| Self-DistillationBase Model=OpenR1-Distill-7B2026.06 | 20.31 | |
| DRIFTBase Model=Olmo3-7B-Instruct-SFT2026.06 | 16.78 | |
| QuratingBase Model=Olmo3-7B-Instruct-SFT2026.06 | 16.52 | |
| BM25Base Model=Olmo3-7B-Instruct-SFT2026.06 | 16.28 | |
| RDSBase Model=Olmo3-7B-Instruct-SFT2026.06 | 15.55 | |
| Olmo3-7B-Instruct-SFTBase Model=Olmo3-7B-Instruct-SFT2026.06 | 15.43 | |
| Random BaselineBase Model=Olmo3-7B-Instruct-SFT2026.06 | 15.07 | |
| IF (Impl. w. GraSS)Base Model=Olmo3-7B-Instruct-SFT2026.06 | 14.53 | |
| LESSBase Model=Olmo3-7B-Instruct-SFT2026.06 | 14.22 | |
| DSIRBase Model=Olmo3-7B-Instruct-SFT2026.06 | 13.51 | |
| Self-DistillationBase Model=Olmo3-7B-Instruct-SFT2026.06 | 9.48 | |
| PromptCoT 2.0training_budget=11,000 mixed trajectories, recipe=GRPO, evaluation=Best-of-52026.02 | 0.742 | |
| Agentic Proposingtraining_budget=11,000 mixed trajectories, recipe=GRPO, evaluation=Best-of-5, proposer=Agentic-Proposer-30B2026.02 | 0.734 | |
| OpenCodeReasoningtraining_budget=11,000 mixed trajectories, recipe=GRPO, evaluation=Best-of-52026.02 | 0.708 | |
| OpenThoughts-S3training_budget=11,000 mixed trajectories, recipe=GRPO, evaluation=Best-of-52026.02 | 0.683 | |
| OpenR1training_budget=11,000 mixed trajectories, recipe=GRPO, evaluation=Best-of-52026.02 | 0.683 | |
| Qwen3-30B-A3B-Thinking-2507mode=zero-shot, evaluation=Best-of-52026.02 | 0.681 | |
| OpenMathReasoningtraining_budget=11,000 mixed trajectories, recipe=GRPO, evaluation=Best-of-52026.02 | 0.657 |