Mathematical Reasoning on AIME 25 (Accuracy, Avg)
12.3AccuracyGRPO + coupled rhythm credit
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GRPO + coupled rhythm creditBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 12.3 | 43.2 | |
| GRPO + global-anchor creditBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 11.5 | 42.1 | |
| GRPO + CAPOBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 10.9 | 41.4 | |
| GRPO + ThinkPRM-1.5BBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 10.6 | 41 | |
| GRPO + high-entropy selectionBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 10.5 | 40.9 | |
| GRPO + local-chunk creditBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 9.3 | 41.3 | |
| GRPO + gradient-based creditBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 9.2 | 40 | |
| GRPO + path-aggregation creditBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 9.1 | 40 | |
| GRPO + Reweight+LoptiBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 8.8 | 41.1 | |
| GRPO + AsyPPOBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 8.2 | 40 | |
| GRPO + random creditBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 8.1 | 39.6 | |
| GRPO + high-entropy creditBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 7.7 | 40.1 | |
| GRPOBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 7.3 | 39.4 | |
| GRPO + token-correlation creditBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 6.4 | 37.4 |