Mathematical Reasoning on MATH 500 (acc@1, Overall)
94Overall ScoreICRL
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| ICRLBackbone=Qwen3-8B2026.05 | 94 | — | |
| Critique-GRPOBackbone=Qwen3-8B2026.05 | 92.6 | — | |
| GRPOBackbone=Qwen3-8B2026.05 | 91 | — | |
| SFTBackbone=Qwen3-8B2026.05 | 83.2 | — | |
| Qwen3-8BBackbone=Qwen3-8B2026.05 | 82 | — | |
| dInferBase Model=Dream-7B-Instruct, Fine-tuning Dataset=Mathematical Reasoning2026.05 | 41.6 | — | |
| TABOMBase Model=Dream-7B-Instruct, Fine-tuning Dataset=Mathematical Reasoning2026.05 | 41.1 | — | |
| T3DBase Model=Dream-7B-Instruct, Fine-tuning Dataset=Mathematical Reasoning2026.05 | 40.7 | — | |
| No-SFTBase Model=Dream-7B-Instruct, Fine-tuning Dataset=Mathematical Reasoning2026.05 | 39.8 | — | |
| SFT-SDBase Model=Dream-7B-Instruct, Fine-tuning Dataset=Mathematical Reasoning2026.05 | 39.8 | — | |
| SFT-GTBase Model=Dream-7B-Instruct, Fine-tuning Dataset=Mathematical Reasoning2026.05 | 37.4 | — | |
| TABOMBase Model=LLaDA-8B-Instruct, Fine-tuning Dataset=Mathematical Reasoning2026.05 | 36.8 | — | |
| dInferBase Model=LLaDA-8B-Instruct, Fine-tuning Dataset=Mathematical Reasoning2026.05 | 36.5 | — | |
| No-SFTBase Model=LLaDA-8B-Instruct, Fine-tuning Dataset=Mathematical Reasoning2026.05 | 36.2 | — | |
| T3DBase Model=LLaDA-8B-Instruct, Fine-tuning Dataset=Mathematical Reasoning2026.05 | 36.1 | — | |
| SFT-SDBase Model=LLaDA-8B-Instruct, Fine-tuning Dataset=Mathematical Reasoning2026.05 | 35.7 | — | |
| SFT-GTBase Model=LLaDA-8B-Instruct, Fine-tuning Dataset=Mathematical Reasoning2026.05 | 35.5 | — | |
| VM-AV-GRPOAlgorithm=VM-AV-GRPO, Backbone=Qwen2.5-Math-1.5B2026.03 | 0.525 | 0.736 | |
| AV-GRPOAlgorithm=AV-GRPO, Backbone=Qwen2.5-Math-1.5B2026.03 | 0.509 | 0.71 | |
| MSA-GRPOAlgorithm=MSA-GRPO, Backbone=Qwen2.5-Math-1.5B2026.03 | 0.507 | 0.694 | |
| FA-GRPOAlgorithm=FA-GRPO, Backbone=Qwen2.5-Math-1.5B2026.03 | 0.499 | 0.722 | |
| PR-GRPOAlgorithm=PR-GRPO, Backbone=Qwen2.5-Math-1.5B2026.03 | 0.494 | 0.712 | |
| CDA-GRPOAlgorithm=CDA-GRPO, Backbone=Qwen2.5-Math-1.5B2026.03 | 0.49 | 0.726 | |
| SVE-LNA-GRPOAlgorithm=SVE-LNA-GRPO, Backbone=Qwen2.5-Math-1.5B2026.03 | 0.487 | 0.73 | |
| GRPOAlgorithm=GRPO, Backbone=Qwen2.5-Math-1.5B2026.03 | 0.478 | 0.724 | |
| Qwen2.5-Math-7BModel=Qwen2.5-Math-7B, Optimization Method=Base2026.05 | — | 52.8 | |
| Qwen2.5-Math-7B + GRPOModel=Qwen2.5-Math-7B, Optimization Method=GRPO2026.05 | — | 78 | |
| Qwen2.5-Math-7B + Resample w/ LOPE (w/ Training Signal Shaping)Model=Qwen2.5-Math-7B, Optimization Method=GRPO + Resample, Prompt Strategy=LOPE, Training Signal Shaping=true2026.05 | — | 81.8 | |
| Qwen2.5-Math-7B + Resample w/ LOPE (w/o Training Signal Shaping)Model=Qwen2.5-Math-7B, Optimization Method=GRPO + Resample, Prompt Strategy=LOPE, Training Signal Shaping=false2026.05 | — | 77.4 | |
| Qwen2.5-Math-7B + Resample w/ Naive PromptModel=Qwen2.5-Math-7B, Optimization Method=GRPO + Resample, Prompt Strategy=Naive Prompt2026.05 | — | 78.2 | |
| Qwen3-1.7B-BaseModel=Qwen3-1.7B-Base, Optimization Method=Base2026.05 | — | 63.4 | |
| Qwen3-1.7B-Base + GRPOModel=Qwen3-1.7B-Base, Optimization Method=GRPO2026.05 | — | 64.2 | |
| Qwen3-1.7B-Base + Resample w/ LOPE (w/ Training Signal Shaping)Model=Qwen3-1.7B-Base, Optimization Method=GRPO + Resample, Prompt Strategy=LOPE, Training Signal Shaping=true2026.05 | — | 68.8 | |
| Qwen3-1.7B-Base + Resample w/ LOPE (w/o Training Signal Shaping)Model=Qwen3-1.7B-Base, Optimization Method=GRPO + Resample, Prompt Strategy=LOPE, Training Signal Shaping=false2026.05 | — | 68 | |
| Qwen3-1.7B-Base + Resample w/ Naive PromptModel=Qwen3-1.7B-Base, Optimization Method=GRPO + Resample, Prompt Strategy=Naive Prompt2026.05 | — | 67 | |
| Qwen3-4B-BaseModel=Qwen3-4B-Base, Optimization Method=Base2026.05 | — | 65.8 | |
| Qwen3-4B-Base + GRPOModel=Qwen3-4B-Base, Optimization Method=GRPO2026.05 | — | 77.8 | |
| Qwen3-4B-Base + Resample w/ LOPE (w/ Training Signal Shaping)Model=Qwen3-4B-Base, Optimization Method=GRPO + Resample, Prompt Strategy=LOPE, Training Signal Shaping=true2026.05 | — | 82.6 | |
| Qwen3-4B-Base + Resample w/ LOPE (w/o Training Signal Shaping)Model=Qwen3-4B-Base, Optimization Method=GRPO + Resample, Prompt Strategy=LOPE, Training Signal Shaping=false2026.05 | — | 85.4 | |
| Qwen3-4B-Base + Resample w/ Naive PromptModel=Qwen3-4B-Base, Optimization Method=GRPO + Resample, Prompt Strategy=Naive Prompt2026.05 | — | 79.8 |