Mathematical Reasoning on MATH (Accuracy, Avg)
79.7AccuracyGRPO + coupled rhythm credit
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GRPO + coupled rhythm creditBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 79.7 | 43.2 | |
| GRPO + global-anchor creditBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 78.9 | 42.1 | |
| GRPO + Reweight+LoptiBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 78.6 | 41.1 | |
| GRPO + CAPOBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 78.4 | 41.4 | |
| GRPO + local-chunk creditBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 78.4 | 41.3 | |
| GRPO + high-entropy creditBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 78 | 40.1 | |
| GRPO + gradient-based creditBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 78 | 40 | |
| GRPO + AsyPPOBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 77.8 | 40 | |
| GRPO + high-entropy selectionBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 77.6 | 40.9 | |
| GRPO + ThinkPRM-1.5BBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 77.6 | 41 | |
| GRPO + random creditBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 77.4 | 39.6 | |
| GRPO + path-aggregation creditBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 77.3 | 40 | |
| GRPOBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 77.1 | 39.4 | |
| GRPO + token-correlation creditBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 76.4 | 37.4 |