Code Generation on Codeforces (accuracy)
13.8AccuracyXRPO
Evaluation Results
| Method | Links | |
|---|---|---|
| XRPOModel=Qwen3-1.7B2025.10 | 13.8 | |
| GSPOModel=Qwen3-1.7B2025.10 | 13.72 | |
| TPO-SModel=Qwen3-1.7B2025.10 | 9.51 | |
| DAPOModel=Qwen3-1.7B2025.10 | 9.48 | |
| SDAR-8BScale=8B, Tokens=55B, FLOPs=2640, Decoding Strategy=greedy static decoding, Block Size=42026.06 | 5.8 | |
| OPDLM-4BScale=4B, Tokens=0.076B, FLOPs=2.4, Decoding Strategy=greedy static decoding, Block Size=42026.06 | 5 | |
| Fast-dLLM-v2-7BScale=8B, Tokens=1B, FLOPs=42, Decoding Strategy=greedy static decoding, Block Size=42026.06 | 5 | |
| SDAR-4BScale=4B, Tokens=55B, FLOPs=1320, Decoding Strategy=greedy static decoding, Block Size=42026.06 | 4 | |
| OPDLM-8BScale=8B, Tokens=0.066B, FLOPs=4.2, Decoding Strategy=greedy static decoding, Block Size=42026.06 | 3.5 |