Mathematical Reasoning on MATH (Accuracy, Response Tokens, Length Reduction)
46.74AccuracyGRPO+FIRSTN
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GRPO+FIRSTNModel=Llama3.2-3B-Instruct, Rollout Group Size (G)=32, Training Time (s)=15519 ± 187, Speedup (times)=1.14 ± 0.122026.05 | 46.74 | 371.49 | 8 | |
| GRPOModel=Llama3.2-3B-Instruct, Rollout Group Size (G)=32, Training Time (s)=17763 ± 1776, Speedup (times)=1.002026.05 | 46.07 | 403.94 | — | |
| CPPOModel=Llama3.2-3B-Instruct, Rollout Group Size (G)=32, Training Time (s)=2553 ± 180, Speedup (times)=6.96 ± 0.492026.05 | 45.28 | 359.8 | 10.9 | |
| PairModel=Llama3.2-3B-Instruct, Rollout Group Size (G)=32, Training Time (s)=2382 ± 162, Speedup (times)=7.46 ± 0.512026.05 | 45.02 | 163.15 | 59.6 | |
| BPPOModel=Llama3.2-3B-Instruct, Rollout Group Size (G)=32, Training Time (s)=2041 ± 26, Speedup (times)=8.70 ± 0.112026.05 | 44.76 | 154.1 | 61.9 |