Mathematical Reasoning on Omni-MATH (Domain Scores and Pass@1)
37Algebra AccuracyARISE + GRPO
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| ARISE + GRPOBase Model=Qwen3-4B-Instruct-2507, Training Dataset=DeepScaleR, G=82026.03 | 37 | 27.2 | 17.6 | 25.5 | 26.8 | |
| DAPOBase Model=Qwen3-4B-Instruct-2507, Training Dataset=DeepScaleR, G=82026.03 | 35.2 | 25.7 | 16.3 | 23.9 | 25.3 | |
| GSPOBase Model=Qwen3-4B-Instruct-2507, Training Dataset=DeepScaleR, G=82026.03 | 34.8 | 25.4 | 16.1 | 25.8 | 25.5 | |
| EvolveR + GRPOBase Model=Qwen3-4B-Instruct-2507, Training Dataset=DeepScaleR, G=82026.03 | 34.5 | 25.1 | 15.4 | 23.5 | 24.6 | |
| SimpleMem + GRPOBase Model=Qwen3-4B-Instruct-2507, Training Dataset=DeepScaleR, G=82026.03 | 33.9 | 24.6 | 16.5 | 23.1 | 24.5 | |
| Dr.GRPOBase Model=Qwen3-4B-Instruct-2507, Training Dataset=DeepScaleR, G=82026.03 | 33.8 | 24.1 | 15.3 | 22.6 | 23.9 | |
| GRPOBase Model=Qwen3-4B-Instruct-2507, Training Dataset=DeepScaleR, G=82026.03 | 33.5 | 24.3 | 15.1 | 22.8 | 23.9 | |
| BaselineBase Model=Qwen3-4B-Instruct-2507, Training Dataset=DeepScaleR, G=82026.03 | 30.1 | 21.2 | 13 | 20.5 | 21.2 | |
| ARISE + GRPOBase Model=Phi-4-mini-instruct-3.8B, Training Dataset=DeepScaleR, G=82026.03 | 15.9 | 11.5 | 7.7 | 10.6 | 11.4 | |
| DAPOBase Model=Phi-4-mini-instruct-3.8B, Training Dataset=DeepScaleR, G=82026.03 | 14.6 | 10.7 | 7 | 9.5 | 10.5 | |
| GSPOBase Model=Phi-4-mini-instruct-3.8B, Training Dataset=DeepScaleR, G=82026.03 | 14.2 | 10.3 | 6.7 | 9.8 | 10.2 | |
| EvolveR + GRPOBase Model=Phi-4-mini-instruct-3.8B, Training Dataset=DeepScaleR, G=82026.03 | 14 | 10.1 | 6.2 | 9.4 | 9.9 | |
| SimpleMem + GRPOBase Model=Phi-4-mini-instruct-3.8B, Training Dataset=DeepScaleR, G=82026.03 | 13.6 | 9.8 | 6.4 | 9 | 9.7 | |
| Dr.GRPOBase Model=Phi-4-mini-instruct-3.8B, Training Dataset=DeepScaleR, G=82026.03 | 13.4 | 9.3 | 6.1 | 8.6 | 9.4 | |
| GRPOBase Model=Phi-4-mini-instruct-3.8B, Training Dataset=DeepScaleR, G=82026.03 | 13.2 | 9.5 | 6 | 8.8 | 9.4 | |
| BaselineBase Model=Phi-4-mini-instruct-3.8B, Training Dataset=DeepScaleR, G=82026.03 | 10.3 | 7.1 | 4.4 | 6.5 | 7.1 |