Mathematical Reasoning on GSM8K (Acc. (%), ∆ (%))
90.4Accuracy (GSM8K)ED-GRPO
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| ED-GRPOBackbone=LLaMA3-8B-Instruct, Decoding Strategy=Self-Consistency2026.05 | 90.4 | 11.1 | |
| ETOBackbone=LLaMA3-8B-Instruct, Decoding Strategy=Self-Consistency2026.05 | 89.2 | 9.9 | |
| ED-iDPOBackbone=LLaMA3-8B-Instruct, Decoding Strategy=Self-Consistency2026.05 | 88.9 | 9.6 | |
| GRPOBackbone=LLaMA3-8B-Instruct, Decoding Strategy=Self-Consistency2026.05 | 88.7 | 9.4 | |
| DAPOBackbone=LLaMA3-8B-Instruct, Decoding Strategy=Self-Consistency2026.05 | 88 | 8.7 | |
| CoTBackbone=LLaMA3-8B-Instruct, Decoding Strategy=Self-Consistency2026.05 | 87.4 | 8.1 | |
| ED-GRPOBackbone=LLaMA3-8B-Instruct, Decoding Strategy=Greedy2026.05 | 86.3 | 7 | |
| GRPOBackbone=LLaMA3-8B-Instruct, Decoding Strategy=Greedy2026.05 | 83.6 | 4.3 | |
| DAPOBackbone=LLaMA3-8B-Instruct, Decoding Strategy=Greedy2026.05 | 82.1 | 2.8 | |
| ETOBackbone=LLaMA3-8B-Instruct, Decoding Strategy=Greedy2026.05 | 81.5 | 2.2 | |
| ED-iDPOBackbone=LLaMA3-8B-Instruct, Decoding Strategy=Greedy2026.05 | 81.5 | 2.2 | |
| CoTBackbone=LLaMA3-8B-Instruct, Decoding Strategy=Greedy2026.05 | 79.3 | — |