Code generation on LiveBench
31.1AccuracyQwen2.5-7B-Instruct
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qwen2.5-7B-InstructArchitecture=Qwen2.5-7B, Training Stage=AR LLM2026.06 | 31.1 | 54.5 | |
| Llama3.1-8B-InstructArchitecture=Llama3.1-8B, Training Stage=AR LLM2026.06 | 19.7 | 42.7 | |
| AGDOArchitecture=Dream-v0-Instruct-7B, Training Stage=RL, Sampling Temperature=0.1, Decoding Strategy=Static decoding, Max Response Length=1,0242026.06 | 18.4 | 38.8 | |
| AGDO-RLArchitecture=Dream-v0-Instruct-7B, Training Stage=RL, Sampling Temperature=0.1, Decoding Strategy=Static decoding, Max Response Length=1,0242026.06 | 18.3 | 38.1 | |
| Diff-GRPOArchitecture=Dream-v0-Instruct-7B, Training Stage=RL, Sampling Temperature=0.1, Decoding Strategy=Static decoding, Max Response Length=1,0242026.06 | 15.2 | 35 | |
| TraceRLArchitecture=Dream-v0-Instruct-7B, Training Stage=RL, Sampling Temperature=0.1, Decoding Strategy=Static decoding, Max Response Length=1,0242026.06 | 14 | 36.5 | |
| Coupled RLArchitecture=Dream-v0-Instruct-7B, Training Stage=RL, Sampling Temperature=0.1, Decoding Strategy=Static decoding, Max Response Length=1,0242026.06 | 13.8 | 34.9 | |
| AGDO-SFTArchitecture=Dream-v0-Instruct-7B, Training Stage=SFT, Sampling Temperature=0.1, Decoding Strategy=Static decoding, Max Response Length=1,0242026.06 | 12.5 | 36 | |
| SFTArchitecture=Dream-v0-Instruct-7B, Training Stage=SFT, Sampling Temperature=0.1, Decoding Strategy=Static decoding, Max Response Length=1,0242026.06 | 11.3 | 33.9 | |
| Dream-v0-Instruct-7BArchitecture=Dream-v0-Instruct-7B, Training Stage=Masked DLLM2026.06 | 10.7 | 28.3 | |
| blockwise SFTArchitecture=Dream-v0-Instruct-7B, Training Stage=SFT, Sampling Temperature=0.1, Decoding Strategy=Static decoding, Max Response Length=1,0242026.06 | 10.2 | 34.4 | |
| LLaDA-8B-InstructArchitecture=LLaDA-8B, Training Stage=Masked DLLM2026.06 | 4.9 | 28.5 |