Math on GSM8K (Pass@1)
96.51Pass@1GT-Reward
Evaluation Results
| Method | Links | |
|---|---|---|
| GT-RewardBase Model=Qwen3-8B-Base2025.11 | 96.51 | |
| GT-RewardBase Model=Llama-3.1-8B-Instruct2025.11 | 95.6 | |
| Self-HarmonyBase Model=Qwen3-8B-Base2025.11 | 95.45 | |
| Co-RewardBase Model=Qwen3-8B-Base2025.11 | 94.8 | |
| GT-RewardBase Model=Qwen3-4B-Base2025.11 | 94.69 | |
| Self-HarmonyBase Model=Qwen3-4B-Base2025.11 | 94.31 | |
| Majority-VotingBase Model=Qwen3-8B-Base2025.11 | 94 | |
| GT-RewardBase Model=Llama-3.2-3B-Instruct2025.11 | 93.7 | |
| Co-RewardBase Model=Qwen3-4B-Base2025.11 | 93.47 | |
| Majority-VotingBase Model=Qwen3-4B-Base2025.11 | 93.44 | |
| IntuitorBase Model=Qwen3-8B-Base2025.11 | 92.28 | |
| Self-HarmonyBase Model=Llama-3.1-8B-Instruct2025.11 | 91.59 | |
| RentBase Model=Qwen3-8B-Base2025.11 | 91.2 | |
| Ling-flash base-2.0Shots=82026.04 | 90.75 | |
| N-3-Super 120B-A12B-BaseShots=82026.04 | 90.67 | |
| RentBase Model=Qwen3-4B-Base2025.11 | 90.49 | |
| Self-HarmonyBase Model=Llama-3.2-3B-Instruct2025.11 | 89.55 | |
| Co-RewardBase Model=Llama-3.1-8B-Instruct2025.11 | 89.48 | |
| Co-RewardBase Model=Llama-3.2-3B-Instruct2025.11 | 89.14 | |
| Majority-VotingBase Model=Llama-3.1-8B-Instruct2025.11 | 88.78 | |
| IntuitorBase Model=Qwen3-4B-Base2025.11 | 87.53 | |
| Self-HarmonyBase Model=Qwen3-1.7B-Base2025.11 | 87.47 | |
| Co-RewardBase Model=Qwen3-1.7B-Base2025.11 | 86.59 | |
| Majority-VotingBase Model=Llama-3.2-3B-Instruct2025.11 | 85.98 | |
| GT-RewardBase Model=Qwen3-1.7B-Base2025.11 | 85.97 | |
| Before RLBase Model=Qwen3-8B-Base2025.11 | 84.76 | |
| Majority-VotingBase Model=Qwen3-1.7B-Base2025.11 | 83.8 | |
| GLM-4.5 Air-BaseShots=82026.04 | 82.6 | |
| ProSeCo SamplingCorrector Sampling=true, Number of shots=52026.02 | 82.18 | |
| LLaDA1.5Corrector Sampling=false, Number of shots=52026.02 | 81.12 | |
| Vanilla SFT + ReMDMCorrector Sampling=true, Training strategy=SFT, Number of shots=52026.02 | 80.97 | |
| IntuitorBase Model=Qwen3-1.7B-Base2025.11 | 80.25 | |
| ProSeCo SFTCorrector Sampling=false, Number of shots=52026.02 | 79.45 | |
| LLaDA-Instruct + ReMDMCorrector Sampling=true, Number of shots=52026.02 | 79.08 | |
| LLaDA-InstructCorrector Sampling=false, Number of shots=52026.02 | 78.85 | |
| RentBase Model=Qwen3-1.7B-Base2025.11 | 78.64 | |
| Vanilla SFTCorrector Sampling=false, Training strategy=SFT, Number of shots=52026.02 | 77.48 | |
| Llama3.1-InstructCorrector Sampling=false, Number of shots=52026.02 | 76.88 | |
| RentBase Model=Llama-3.1-8B-Instruct2025.11 | 69.88 | |
| LLaDA-BaseCorrector Sampling=false, Number of shots=52026.02 | 66.72 | |
| IntuitorBase Model=Llama-3.1-8B-Instruct2025.11 | 66.25 | |
| Before RLBase Model=Qwen3-1.7B-Base2025.11 | 65.58 | |
| Before RLBase Model=Llama-3.1-8B-Instruct2025.11 | 60.48 | |
| Before RLBase Model=Qwen3-4B-Base2025.11 | 55.72 | |
| IntuitorBase Model=Llama-3.2-3B-Instruct2025.11 | 16.73 | |
| RentBase Model=Llama-3.2-3B-Instruct2025.11 | 16.67 | |
| Before RLBase Model=Llama-3.2-3B-Instruct2025.11 | 16.65 |