Mathematical Reasoning on Countdown (test)
85AccuracyRandOpt
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| RandOptModel=OLMo3-7B-Inst, K=502026.05 | 85 | — | |
| CoRPModel=OLMo3-7B-Inst2026.05 | 78.9 | — | |
| RandOptModel=OLMo3-7B-Inst, K=12026.05 | 74.4 | — | |
| CCDDSize=6M, Depth T=32025.10 | 73.7 | — | |
| PPOModel=OLMo3-7B-Inst2026.05 | 69 | — | |
| GRPOModel=OLMo3-7B-Inst2026.05 | 68.5 | — | |
| LTSize=6M, Depth T=32025.10 | 68.2 | — | |
| CCDDSize=6M, Depth T=22025.10 | 67.8 | — | |
| BaseModel=OLMo3-7B-Inst2026.05 | 64.8 | — | |
| RandOptModel=Llama3.1-8B-Inst, K=502026.05 | 63.6 | — | |
| LTSize=6M, Depth T=22025.10 | 60.6 | — | |
| RandOptModel=Qwen2.5-3B-Inst, K=502026.05 | 58.4 | — | |
| RandOptModel=Qwen2.5-1.5B-Inst, K=502026.05 | 52.7 | — | |
| MDMSize=6M, Depth T=202025.10 | 52 | — | |
| wd1Inference Steps=2562025.07 | 51.2 | — | |
| LlamaSize=13B, Depth T=Seq. Len.2025.10 | 51.1 | — | |
| FDMBackbone=LLaDA-MoE-7B-Instruct, Search Width (K)=42025.12 | 46.88 | 3.7 | |
| FDMBackbone=LLaDA-MoE-7B-Instruct, Search Width (K)=32025.12 | 46.1 | 3.71 | |
| wd1Inference Steps=5122025.07 | 46.1 | — | |
| CoRPModel=Llama3.1-8B-Inst2026.05 | 42.3 | — | |
| d1Inference Steps=512, Evaluation Protocol=reported2025.07 | 42.2 | — | |
| GPT2 scratchSize=303M, Depth T=Seq. Len.2025.10 | 41.3 | — | |
| LlamaSize=7B, Depth T=Seq. Len.2025.10 | 41.1 | — | |
| ProbabilityBackbone=LLaDA-MoE-7B-Instruct, Step Number (T)=2562025.12 | 40.62 | 4.15 | |
| FDMBackbone=LLaDA-MoE-7B-Instruct, Search Width (K)=22025.12 | 40.62 | 3.73 | |
| MarginBackbone=LLaDA-MoE-7B-Instruct, Step Number (T)=2562025.12 | 40.23 | 4.03 | |
| EntropyBackbone=LLaDA-MoE-7B-Instruct, Step Number (T)=2562025.12 | 38.67 | 3.95 | |
| + diffu-GRPOInference Steps=512, Evaluation Protocol=reported2025.07 | 37.1 | — | |
| CoRPModel=Qwen2.5-3B-Inst2026.05 | 36.9 | — | |
| PPOModel=Qwen2.5-3B-Inst2026.05 | 35.3 | — | |
| d1Inference Steps=512, Evaluation Protocol=reproduced2025.07 | 35.2 | — | |
| + diffu-GRPOInference Steps=512, Evaluation Protocol=reproduced2025.07 | 34 | — | |
| GRPOModel=Qwen2.5-3B-Inst2026.05 | 32.6 | — | |
| d1Inference Steps=256, Evaluation Protocol=reported2025.07 | 32 | — | |
| GPT2 scratchSize=6M, Depth T=Seq. Len.2025.10 | 31.9 | — | |
| + diffu-GRPOInference Steps=256, Evaluation Protocol=reported2025.07 | 31.3 | — | |
| GRPOModel=Qwen2.5-1.5B-Inst2026.05 | 27.5 | — | |
| + diffu-GRPOInference Steps=256, Evaluation Protocol=reproduced2025.07 | 27 | — | |
| PPOModel=Qwen2.5-1.5B-Inst2026.05 | 27 | — | |
| CoRPModel=Qwen2.5-1.5B-Inst2026.05 | 26.2 | — | |
| d1Inference Steps=256, Evaluation Protocol=reproduced2025.07 | 25.8 | — | |
| FDMBackbone=LLaDA-8B-Instruct, Search Width (K)=42025.12 | 25 | 5.28 | |
| Self-SFTBackbone=Qwen2.5-3B-Instruct2026.05 | 23.7 | — | |
| FDMBackbone=LLaDA-1.5, Search Width (K)=42025.12 | 23.43 | 5.25 | |
| BaseBackbone=Qwen2.5-3B-Instruct2026.05 | 22.5 | — | |
| FDMBackbone=LLaDA-1.5, Search Width (K)=32025.12 | 21.88 | 6.48 | |
| FDMBackbone=LLaDA-8B-Instruct, Search Width (K)=32025.12 | 21.09 | 6.11 | |
| Iter-SFTBackbone=Qwen2.5-3B-Instruct2026.05 | 20.9 | — | |
| FDMBackbone=LLaDA-1.5, Search Width (K)=22025.12 | 20.7 | 8.62 | |
| MarginBackbone=LLaDA-1.5, Step Number (T)=2562025.12 | 20.31 | 10.92 | |
| EntropyBackbone=LLaDA-1.5, Step Number (T)=2562025.12 | 20.31 | 10.39 | |
| LLaDA-8B-InstructInference Steps=2562025.07 | 19.5 | — | |
| MarginBackbone=LLaDA-8B-Instruct, Step Number (T)=2562025.12 | 19.14 | 10.89 | |
| FDMBackbone=LLaDA-8B-Instruct, Search Width (K)=22025.12 | 19.14 | 8.21 | |
| ProbabilityBackbone=LLaDA-8B-Instruct, Step Number (T)=2562025.12 | 18.75 | 11.19 | |
| EntropyBackbone=LLaDA-8B-Instruct, Step Number (T)=2562025.12 | 18.44 | 10.39 | |
| FDMBackbone=MMaDA-8B-MixCoT, Search Width (K)=42025.12 | 17.58 | 4.97 | |
| Low-SFTBackbone=Qwen2.5-3B-Instruct2026.05 | 17.1 | — | |
| Anchored LearningBackbone=Qwen2.5-3B-Instruct2026.05 | 17 | — | |
| DFTBackbone=Qwen2.5-3B-Instruct2026.05 | 16.8 | — | |
| FDMBackbone=MMaDA-8B-MixCoT, Search Width (K)=32025.12 | 16.4 | 6.1 | |
| ProbabilityBackbone=LLaDA-1.5, Step Number (T)=2562025.12 | 16.02 | 11.2 | |
| LLaDA-8B-InstructInference Steps=5122025.07 | 16 | — | |
| PPOModel=Qwen2.5-0.5B-Inst2026.05 | 14.8 | — | |
| MarginBackbone=MMaDA-8B-MixCoT, Step Number (T)=2562025.12 | 14.61 | 10.84 | |
| CoRPModel=Qwen2.5-0.5B-Inst2026.05 | 14.5 | — | |
| FDMBackbone=MMaDA-8B-MixCoT, Search Width (K)=22025.12 | 14.45 | 8.13 | |
| SFTBackbone=Qwen2.5-3B-Instruct2026.05 | 13.8 | — | |
| GRPOModel=Qwen2.5-0.5B-Inst2026.05 | 13 | — | |
| STMBackbone=Qwen2.5-3B-Instruct2026.05 | 12.1 | — | |
| BaseModel=Llama3.1-8B-Inst2026.05 | 10.8 | — | |
| RandOptModel=Llama3.1-8B-Inst, K=12026.05 | 10.7 | — | |
| RandOptModel=Qwen2.5-3B-Inst, K=12026.05 | 10.4 | — | |
| BaseModel=Qwen2.5-3B-Inst2026.05 | 10 | — | |
| GRPOModel=Llama3.1-8B-Inst2026.05 | 10 | — | |
| PPOModel=Llama3.1-8B-Inst2026.05 | 9.9 | — | |
| RandOptModel=Qwen2.5-1.5B-Inst, K=12026.05 | 9.1 | — | |
| RandOptModel=Qwen2.5-0.5B-Inst, K=502026.05 | 8.4 | — | |
| EntropyBackbone=MMaDA-8B-MixCoT, Step Number (T)=2562025.12 | 7.03 | 10.31 | |
| BaseModel=Qwen2.5-1.5B-Inst2026.05 | 6.7 | — | |
| KL-SFTBackbone=Qwen2.5-3B-Instruct2026.05 | 5.6 | — | |
| RandOptModel=Qwen2.5-0.5B-Inst, K=12026.05 | 4.8 | — | |
| ProbabilityBackbone=MMaDA-8B-MixCoT, Step Number (T)=2562025.12 | 4.69 | 11.15 | |
| BaseModel=Qwen2.5-0.5B-Inst2026.05 | 0.1 | — |