Mathematical Reasoning on HMMT (accuracy)
77.6HMMT AccuracyAgentic Proposing
Evaluation Results
| Method | Links | |
|---|---|---|
| Agentic Proposingtraining_budget=11,000 mixed trajectories, recipe=GRPO, evaluation=Mean@64, proposer=Agentic-Proposer-30B2026.02 | 77.6 | |
| PromptCoT 2.0training_budget=11,000 mixed trajectories, recipe=GRPO, evaluation=Mean@642026.02 | 76.7 | |
| OpenMathReasoningtraining_budget=11,000 mixed trajectories, recipe=GRPO, evaluation=Mean@642026.02 | 72.2 | |
| Qwen3-30B-A3B-Thinking-2507mode=zero-shot, evaluation=Mean@642026.02 | 71.4 | |
| OpenThoughts-S3training_budget=11,000 mixed trajectories, recipe=GRPO, evaluation=Mean@642026.02 | 70 | |
| OpenR1training_budget=11,000 mixed trajectories, recipe=GRPO, evaluation=Mean@642026.02 | 68.9 | |
| OpenCodeReasoningtraining_budget=11,000 mixed trajectories, recipe=GRPO, evaluation=Mean@642026.02 | 64.9 | |
| QuestABackbone=OpenMath-Nemotron-1.5B, Sampling Strategy (@k)=@32, Dynamic Sampling=true, Training Steps=2000, Train Batch Size=128, Rollout N=16, Max Context Length=32k, Token Budget=2.6x10^8k2025.12 | 40.94 | |
| JustRL-NemotronBackbone=OpenMath-Nemotron-1.5B, Sampling Strategy (@k)=@32, Dynamic Sampling=false, Training Steps=3440, Train Batch Size=256, Rollout N=8, Max Context Length=16k, Token Budget=1.1x10^8k2025.12 | 40.63 | |
| BackboneBackbone=OpenMath-Nemotron-1.5B, Sampling Strategy (@k)=@322025.12 | 30.1 | |
| JustRL-DeepSeekSampling strategy=@32, Backbone=DeepSeek-R1-Distill-Qwen-1.5B2025.12 | 21.98 | |
| ProRL-V2Sampling strategy=@32, Backbone=DeepSeek-R1-Distill-Qwen-1.5B2025.12 | 19.38 | |
| DeepScaleR-1.5BSampling strategy=@32, Backbone=DeepSeek-R1-Distill-Qwen-1.5B2025.12 | 18.96 | |
| Backbone (DeepSeek-R1-Distill-Qwen-1.5B)Sampling strategy=@32, Backbone=DeepSeek-R1-Distill-Qwen-1.5B2025.12 | 13.44 |