Mathematical Reasoning on HMMT (Pass@1)
73.3Pass@1Pairwise
Evaluation Results
| Method | Links | |
|---|---|---|
| PairwiseVerifier=V1, Verification Budget=3N budget, Sampling Strategy=Best-of-N2026.07 | 73.3 | |
| LLM-as-a-VerifierVerification Budget=2N budget, Sampling Strategy=Best-of-N2026.07 | 73.3 | |
| EvoEnvBase Model=Nemotron-Cascade-8B2026.05 | 64.2 | |
| PointwiseMechanism=LM judge, Verification Budget=N budget, Sampling Strategy=Best-of-N2026.07 | 63.5 | |
| UntrainedBase Model=Nemotron-Cascade-8B2026.05 | 61.5 | |
| EvoEnvBase Model=Qwen3-4B-Thinking-25072026.05 | 60 | |
| UntrainedBase Model=Qwen3-4B-Thinking-25072026.05 | 56.9 | |
| R-ZeroBase Model=Nemotron-Cascade-8B2026.05 | 55.8 | |
| RLVEBase Model=Nemotron-Cascade-8B2026.05 | 55.4 | |
| R-ZeroBase Model=Qwen3-4B-Thinking-25072026.05 | 53.5 | |
| Base modelSampling Strategy=Best-of-N2026.07 | 52 | |
| RLVEBase Model=Qwen3-4B-Thinking-25072026.05 | 51.5 | |
| DAPOBase Model=Qwen3-4B-Thinking-25072026.05 | 48.8 | |
| DAPOBase Model=Nemotron-Cascade-8B2026.05 | 40 | |
| UntrainedBase Model=Qwen3-4B-Instruct-25072026.05 | 30 | |
| EvoEnvBase Model=Qwen3-4B-Instruct-25072026.05 | 30 | |
| RLVEBase Model=Qwen3-4B-Instruct-25072026.05 | 27.7 | |
| R-ZeroBase Model=Qwen3-4B-Instruct-25072026.05 | 26.5 | |
| e3-1.7B + InT+ RLRL Data Size=64, Evaluation Protocol=8 rollouts avg.2026.01 | 24.58 | |
| e3-1.7B + RLRL Data Size=1216, Evaluation Protocol=8 rollouts avg.2026.01 | 22.5 | |
| e3-1.7B + Distill + RLRL Data Size=1216, Evaluation Protocol=8 rollouts avg.2026.01 | 22.5 | |
| DAPOBase Model=Qwen3-4B-Instruct-25072026.05 | 19.8 |