Code Generation on LiveCodeBench V2 (accuracy)
26.9AccuracyQwen2.5-7B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qwen2.5-7BModel Backbone=Qwen2.5-7B2026.06 | 26.9 | — | |
| Qwen2.5-7B-InstructArchitecture=Qwen2.5-7B, Training Stage=AR LLM2026.06 | 26.9 | 54.5 | |
| Llama3.1-8B-InstructArchitecture=Llama3.1-8B, Training Stage=AR LLM2026.06 | 20 | 42.7 | |
| AGDOArchitecture=Dream-v0-Instruct-7B, Training Stage=RL, Sampling Temperature=0.1, Decoding Strategy=Static decoding, Max Response Length=1,0242026.06 | 15.6 | 38.8 | |
| AGDO-RLArchitecture=Dream-v0-Instruct-7B, Training Stage=RL, Sampling Temperature=0.1, Decoding Strategy=Static decoding, Max Response Length=1,0242026.06 | 14.7 | 38.1 | |
| Diff-GRPOArchitecture=Dream-v0-Instruct-7B, Training Stage=RL, Sampling Temperature=0.1, Decoding Strategy=Static decoding, Max Response Length=1,0242026.06 | 13.9 | 35 | |
| AGDO-SFTArchitecture=Dream-v0-Instruct-7B, Training Stage=SFT, Sampling Temperature=0.1, Decoding Strategy=Static decoding, Max Response Length=1,0242026.06 | 13.1 | 36 | |
| TraceRLArchitecture=Dream-v0-Instruct-7B, Training Stage=RL, Sampling Temperature=0.1, Decoding Strategy=Static decoding, Max Response Length=1,0242026.06 | 13 | 36.5 | |
| blockwise SFTArchitecture=Dream-v0-Instruct-7B, Training Stage=SFT, Sampling Temperature=0.1, Decoding Strategy=Static decoding, Max Response Length=1,0242026.06 | 11.8 | 34.4 | |
| Coupled RLArchitecture=Dream-v0-Instruct-7B, Training Stage=RL, Sampling Temperature=0.1, Decoding Strategy=Static decoding, Max Response Length=1,0242026.06 | 11.8 | 34.9 | |
| SFTArchitecture=Dream-v0-Instruct-7B, Training Stage=SFT, Sampling Temperature=0.1, Decoding Strategy=Static decoding, Max Response Length=1,0242026.06 | 11.5 | 33.9 | |
| Uni-EnergyModel Backbone=Dream-7B, Decoding Strategy=Unified energy-based decoding2026.06 | 11.1 | — | |
| Dream-v0-Instruct-7BArchitecture=Dream-v0-Instruct-7B, Training Stage=Masked DLLM2026.06 | 10.7 | 28.3 | |
| Uni-EnergyModel Backbone=LLaDa-8B, Decoding Strategy=Unified energy-based decoding2026.06 | 10.5 | — | |
| Qwen2.5-0.5BModel Backbone=Qwen2.5-0.5B2026.06 | 10.4 | — | |
| Ind-EnergyModel Backbone=Dream-7B, Decoding Strategy=Independent energy-based decoding2026.06 | 9.8 | — | |
| Ind-EnergyModel Backbone=LLaDa-8B, Decoding Strategy=Independent energy-based decoding2026.06 | 9.2 | — | |
| APDModel Backbone=Dream-7B, Decoding Strategy=Speculative decoding2026.06 | 9 | — | |
| Inv-EnergyModel Backbone=Dream-7B, Decoding Strategy=Invariant energy-based decoding2026.06 | 8.1 | — | |
| Inv-EnergyModel Backbone=LLaDa-8B, Decoding Strategy=Invariant energy-based decoding2026.06 | 7.8 | — | |
| Dream-7BModel Backbone=Dream-7B, Decoding Strategy=Base2026.06 | 7.5 | — | |
| APDModel Backbone=LLaDa-8B, Decoding Strategy=Speculative decoding2026.06 | 7.1 | — | |
| COREModel Backbone=Dream-7B, Decoding Strategy=Remask decoding2026.06 | 6.3 | — | |
| LLaDA-8B-InstructArchitecture=LLaDA-8B, Training Stage=Masked DLLM2026.06 | 5.9 | 28.5 | |
| LLaDa-8BModel Backbone=LLaDa-8B, Decoding Strategy=Base2026.06 | 5.8 | — | |
| DAWNModel Backbone=LLaDa-8B, Decoding Strategy=Dependency decoding2026.06 | 5.5 | — | |
| COREModel Backbone=LLaDa-8B, Decoding Strategy=Remask decoding2026.06 | 4.9 | — |