Mathematical Reasoning on In-Distribution Benchmarks Summary
65.7Average ScoreICPO
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| ICPOBase Model=8B, Training Method=ICPO2025.10 | 65.7 | 2.2 | |
| ICPO†Base Model=8B, Training Method=ICPO†2025.10 | 65 | 1.5 | |
| GRPOExpertDomainBase Model=8B, Training Method=GRPOExpertDomain2025.10 | 64.6 | 1.1 | |
| GRPOExtraRolloutsBase Model=8B, Training Method=GRPOExtraRollouts2025.10 | 64.3 | 0.8 | |
| GRPOBase Model=8B, Training Method=GRPO2025.10 | 63.5 | — | |
| ICPOBase Model=1.7B, Training Method=ICPO2025.10 | 52.5 | 4.1 | |
| ICPO†Base Model=1.7B, Training Method=ICPO†2025.10 | 51.4 | 3 | |
| GRPOExtraRolloutsBase Model=1.7B, Training Method=GRPOExtraRollouts2025.10 | 50.7 | 2.3 | |
| GRPOExpertDomainBase Model=1.7B, Training Method=GRPOExpertDomain2025.10 | 49.6 | 1.2 | |
| GRPOBase Model=1.7B, Training Method=GRPO2025.10 | 48.4 | — | |
| Qwen3Base Model=8B, Training Method=Base2025.10 | 48.4 | — | |
| Qwen3Base Model=1.7B, Training Method=Base2025.10 | 42.7 | — |