Mathematical Reasoning on AMC 23 (Accuracy, Avg)
65.4AccuracyGRPO + coupled rhythm credit
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GRPO + coupled rhythm creditBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 65.4 | 43.2 | |
| GRPO + global-anchor creditBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 63.1 | 42.1 | |
| GRPO + local-chunk creditBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 62.9 | 41.3 | |
| GRPO + Reweight+LoptiBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 62.5 | 41.1 | |
| GRPO + CAPOBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 61.1 | 41.4 | |
| GRPO + ThinkPRM-1.5BBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 61 | 41 | |
| GRPO + high-entropy selectionBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 60.5 | 40.9 | |
| GRPO + high-entropy creditBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 60.4 | 40.1 | |
| GRPO + random creditBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 60.1 | 39.6 | |
| GRPO + AsyPPOBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 59.5 | 40 | |
| GRPO + path-aggregation creditBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 59.5 | 40 | |
| GRPOBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 59.1 | 39.4 | |
| GRPO + token-correlation creditBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 59 | 37.4 | |
| GRPO + gradient-based creditBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 58.4 | 40 |