Mathematical Reasoning on s1K curated (eval)
48.9AccuracyED-iDPO
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| ED-iDPOBackbone=Qwen2.5-7B-Instruct-SFT, Decoding Strategy=Self-Consistency2026.05 | 48.9 | 14.3 | |
| ED-GRPOBackbone=Qwen2.5-7B-Instruct-SFT, Decoding Strategy=Self-Consistency2026.05 | 48.4 | 13.8 | |
| DAPOBackbone=Qwen2.5-7B-Instruct-SFT, Decoding Strategy=Self-Consistency2026.05 | 48.1 | 13.5 | |
| ETOBackbone=Qwen2.5-7B-Instruct-SFT, Decoding Strategy=Self-Consistency2026.05 | 47.3 | 12.7 | |
| GRPOBackbone=Qwen2.5-7B-Instruct-SFT, Decoding Strategy=Self-Consistency2026.05 | 46.7 | 12.1 | |
| CoTBackbone=Qwen2.5-7B-Instruct-SFT, Decoding Strategy=Self-Consistency2026.05 | 46.1 | 11.5 | |
| ED-GRPOBackbone=Qwen2.5-7B-Instruct-SFT, Decoding Strategy=Greedy2026.05 | 38.6 | 4 | |
| ETOBackbone=Qwen2.5-7B-Instruct-SFT, Decoding Strategy=Greedy2026.05 | 37 | 2.4 | |
| ED-iDPOBackbone=Qwen2.5-7B-Instruct-SFT, Decoding Strategy=Greedy2026.05 | 36.9 | 2.3 | |
| DAPOBackbone=Qwen2.5-7B-Instruct-SFT, Decoding Strategy=Greedy2026.05 | 36.1 | 1.5 | |
| GRPOBackbone=Qwen2.5-7B-Instruct-SFT, Decoding Strategy=Greedy2026.05 | 35.1 | 0.5 | |
| CoTBackbone=Qwen2.5-7B-Instruct-SFT, Decoding Strategy=Greedy2026.05 | 34.6 | — | |
| ED-iDPOBackbone=LLaMA3.1-8B-Instruct-SFT, Decoding Strategy=Self-Consistency2026.05 | 28.6 | 11.4 | |
| ED-GRPOBackbone=LLaMA3.1-8B-Instruct-SFT, Decoding Strategy=Self-Consistency2026.05 | 28.1 | 10.9 | |
| GRPOBackbone=LLaMA3.1-8B-Instruct-SFT, Decoding Strategy=Self-Consistency2026.05 | 27.3 | 10.1 | |
| CoTBackbone=LLaMA3.1-8B-Instruct-SFT, Decoding Strategy=Self-Consistency2026.05 | 27 | 9.8 | |
| ETOBackbone=LLaMA3.1-8B-Instruct-SFT, Decoding Strategy=Self-Consistency2026.05 | 26.7 | 9.5 | |
| DAPOBackbone=LLaMA3.1-8B-Instruct-SFT, Decoding Strategy=Self-Consistency2026.05 | 26.7 | 9.5 | |
| ED-GRPOBackbone=LLaMA3.1-8B-Instruct-SFT, Decoding Strategy=Greedy2026.05 | 21 | 3.8 | |
| GRPOBackbone=LLaMA3.1-8B-Instruct-SFT, Decoding Strategy=Greedy2026.05 | 18.9 | 1.7 | |
| ED-iDPOBackbone=LLaMA3.1-8B-Instruct-SFT, Decoding Strategy=Greedy2026.05 | 18.7 | 1.5 | |
| DAPOBackbone=LLaMA3.1-8B-Instruct-SFT, Decoding Strategy=Greedy2026.05 | 18.3 | 1.1 | |
| ETOBackbone=LLaMA3.1-8B-Instruct-SFT, Decoding Strategy=Greedy2026.05 | 17.7 | 0.5 | |
| CoTBackbone=LLaMA3.1-8B-Instruct-SFT, Decoding Strategy=Greedy2026.05 | 17.2 | — |