Mathematical Reasoning on Olympiad (Accuracy, Avg)
47Accuracy (Olympiad)GRPO + coupled rhythm credit
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GRPO + coupled rhythm creditBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 47 | 43.2 | |
| GRPO + Reweight+LoptiBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 46.5 | 41.1 | |
| GRPO + global-anchor creditBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 46.1 | 42.1 | |
| GRPO + CAPOBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 45.9 | 41.4 | |
| GRPO + local-chunk creditBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 45.9 | 41.3 | |
| GRPO + high-entropy creditBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 45.6 | 40.1 | |
| GRPO + high-entropy selectionBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 45.6 | 40.9 | |
| GRPO + ThinkPRM-1.5BBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 45.4 | 41 | |
| GRPO + gradient-based creditBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 45.4 | 40 | |
| GRPO + token-correlation creditBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 45.3 | 37.4 | |
| GRPO + path-aggregation creditBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 45.1 | 40 | |
| GRPO + AsyPPOBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 44.7 | 40 | |
| GRPOBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 44.2 | 39.4 | |
| GRPO + random creditBase Model=Qwen3-8B-Base, Response Length=1K2025.10 | 43.3 | 39.6 |