Mathematical Reasoning Verification on D1-D5 (test)
81.8AccuracyVerify-RL
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Verify-RLModel=Qwen3-0.6B, RL Algorithm=GRPO, Training State=Trained2026.02 | 81.8 | 16.6 | 25 | — | |
| Verify-RLModel=Qwen3-3B, RL Algorithm=GRPO, Training State=Trained2026.02 | 81.8 | 10.2 | 14 | — | |
| Verify-RLModel=DS-R1-1.5B, RL Algorithm=GRPO, Training State=Trained2026.02 | 80 | 18.2 | 29 | — | |
| Verify-RLModel=Qwen3-1.5B, RL Algorithm=GRPO, Training State=Trained2026.02 | 79.6 | 22.6 | 40 | — | |
| Qwen3-3B (Base)Model=Qwen3-3B, Training State=Base (Pre-trained)2026.02 | 71.6 | — | — | — | |
| Qwen3-0.6B (Base)Model=Qwen3-0.6B, Training State=Base (Pre-trained)2026.02 | 65.2 | — | — | — | |
| DS-R1-1.5B (Base)Model=DS-R1-1.5B, Training State=Base (Pre-trained)2026.02 | 61.8 | — | — | — | |
| Qwen3-1.5B (Base)Model=Qwen3-1.5B, Training State=Base (Pre-trained)2026.02 | 57 | — | — | — |