Mathematical Reasoning on NuminaMath (val)
21.7AccuracyClaude 3.7 Sonnet
Evaluation Results
| Method | Links | |
|---|---|---|
| Claude 3.7 SonnetModel Size=~100B+, Training Strategy=Baseline Model2026.02 | 21.7 | |
| Claude 3.5 HaikuModel Size=~20B+, Training Strategy=Baseline Model2026.02 | 16.3 | |
| Gold Match (Verifiable Rewards)Model Size=4B, Training Strategy=RL with GT2026.02 | 16.1 | |
| DeepSeek-R1Model Size=37B, Training Strategy=Baseline Model2026.02 | 15.5 | |
| Rubric-Augmented ClassifierModel Size=4B, Training Strategy=RL w/o GT2026.02 | 15.2 | |
| Baseline ClassifierModel Size=4B, Training Strategy=RL w/o GT2026.02 | 10.5 | |
| Qwen3-4B (No RL)Model Size=4B, Training Strategy=Baseline (RL w/o GT)2026.02 | 5.5 | |
| Mistral-7BModel Size=7B, Training Strategy=Baseline Model2026.02 | 3.7 |