Mathematical Reasoning on Minerva (pass@4)
44.84pass@4DrGRPO
Evaluation Results
| Method | Links | |
|---|---|---|
| DrGRPOBackbone=Qwen3-4B, Training-free=false2026.04 | 44.84 | |
| ThinkTwiceBackbone=OLMo3-7B, Training-free=false2026.04 | 44.33 | |
| ThinkTwiceBackbone=Qwen3-4B, Training-free=false2026.04 | 43.93 | |
| GRPOBackbone=Qwen3-4B, Training-free=false2026.04 | 42.9 | |
| DAPOBackbone=OLMo3-7B, Training-free=false2026.04 | 42.81 | |
| DrGRPOBackbone=OLMo3-7B, Training-free=false2026.04 | 42.75 | |
| Self-RefineBackbone=OLMo3-7B, Training-free=true2026.04 | 42.35 | |
| Base ModelBackbone=OLMo3-7B, Training-free=false2026.04 | 41.58 | |
| ReflexionBackbone=OLMo3-7B, Training-free=true2026.04 | 41.48 | |
| Self-RefineBackbone=Qwen3-4B, Training-free=true2026.04 | 41.19 | |
| GRPOBackbone=OLMo3-7B, Training-free=false2026.04 | 41.08 | |
| ReflexionBackbone=Qwen3-4B, Training-free=true2026.04 | 40.87 | |
| Base ModelBackbone=Qwen3-4B, Training-free=false2026.04 | 40.81 | |
| DAPOBackbone=Qwen3-4B, Training-free=false2026.04 | 40.09 |