Mathematical Reasoning on ViRL39K (test)
18.06AccuracyBase model
Evaluation Results
| Method | Links | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Base modelGRPO-based RL=false, PromptTokLen (Overall)=377.10 ± 577.74, OutputTokLen (Overall)=1031.82 ± 1679.202026.04 | 18.06 | — | — | 26.92 | 14.29 | 0 | 0 | 0 | 16.67 | 16.67 | 18.06 | |
| GRPO-based RLGRPO-based RL=true, PromptTokLen (Overall)=856.10 ± 577.74, OutputTokLen (Overall)=835.92 ± 1435.582026.04 | 8.33 | 80.56 | 54.17 | 11.54 | 9.52 | 0 | 0 | 0 | 0 | 8.33 | 8.33 |