Mathematical Reasoning on BRUMO25
99.2AccuracyVibeThinker-3B + CLR
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| VibeThinker-3B + CLRParams=3B, Post-training optimization=CLR2026.06 | 99.2 | — | — | |
| VibeThinker-3BParams=3B, CLR reasoning enhancement=true2026.06 | 99.2 | — | — | |
| Kimi K2.5Params=1T2026.06 | 98.3 | — | — | |
| Gemini 3 ProParams=N/A2026.06 | 98.3 | — | — | |
| DeepSeek V3.2Params=671B2026.06 | 96.7 | — | — | |
| OpenAI o3 (high)Params=N/A2026.06 | 95.8 | — | — | |
| Grok 4Params=N/A2026.06 | 95 | — | — | |
| VibeThinker-3BParams=3B2026.06 | 93.8 | — | — | |
| VibeThinker-3BParams=3B2026.06 | 93.8 | — | — | |
| DeepConfModel=DeepSeek-8B, Compute Budget=@5122026.02 | 92.6 | 21.7 | — | |
| CoRefine TreeModel=DeepSeek-8B, Compute Budget=@5122026.02 | 92.6 | 0.44 | — | |
| MajorityModel=Qwen3-32B, Compute Budget=@5122026.02 | 92.6 | 21.7 | — | |
| DeepConfModel=DeepSeek-8B, Sample Count=@5122026.02 | 92.6 | — | — | |
| CoRefine TreeModel=DeepSeek-8B2026.02 | 92.6 | — | — | |
| MajorityModel=Qwen3-32B, Sampling Mode=Parallel, Sample Count=@5122026.02 | 92.6 | — | — | |
| DeepConfModel=PaCoRe-8B, Sample Count=@5122026.02 | 92.6 | — | — | |
| CoRefine TreeModel=PaCoRe-8B2026.02 | 92.6 | — | — | |
| GPT-OSS (high)Params=120B2026.06 | 92.5 | — | — | |
| DeepSeek R1 0528Params=671B2026.06 | 92.5 | — | — | |
| MajorityModel=DeepSeek-8B, Compute Budget=@5122026.02 | 92 | 35.6 | — | |
| DeepConfModel=Qwen3-32B, Compute Budget=@5122026.02 | 92 | 1.37 | — | |
| MajorityModel=DeepSeek-8B, Sampling Mode=Parallel, Sample Count=@5122026.02 | 92 | — | — | |
| DeepConfModel=DeepSeek-8B, Sample Count=@202026.02 | 92 | — | — | |
| DeepConfModel=Qwen3-32B, Sample Count=@5122026.02 | 92 | — | — | |
| GPT-5 (high)Params=N/A2026.06 | 91.7 | — | — | |
| MajorityModel=DeepSeek-8B, Sampling Mode=Parallel, Sample Count=@202026.02 | 91.3 | — | — | |
| MajorityModel=DeepSeek-8B, Sampling Mode=Sequential, Sample Count=@202026.02 | 90.7 | — | — | |
| MajorityModel=Qwen3-32B, Sampling Mode=Parallel, Sample Count=@202026.02 | 90.7 | — | — | |
| MajorityModel=Qwen3-32B, Sampling Mode=Sequential, Sample Count=@202026.02 | 90.7 | — | — | |
| DeepConfModel=Qwen3-32B, Sample Count=@202026.02 | 90.7 | — | — | |
| MajorityModel=PaCoRe-8B, Sampling Mode=Parallel, Sample Count=@5122026.02 | 90.7 | — | — | |
| MajorityModel=PaCoRe-8B, Sampling Mode=Parallel, Sample Count=@202026.02 | 90.7 | — | — | |
| MajorityModel=PaCoRe-8B, Sampling Mode=Sequential, Sample Count=@202026.02 | 90.7 | — | — | |
| DeepConfModel=PaCoRe-8B, Sample Count=@202026.02 | 90.7 | — | — | |
| CoRefine TreeModel=Qwen3-32B, Compute Budget=@5122026.02 | 90.6 | 0.28 | — | |
| CoRefine TreeModel=Qwen3-32B2026.02 | 90.6 | — | — | |
| Gemini 2.5 ProParams=N/A2026.06 | 90 | — | — | |
| GLM-4.5-AirParams=106B2026.06 | 90 | — | — | |
| CoRefineModel=DeepSeek-8B, Compute Budget=@5122026.02 | 86.7 | 0.4 | — | |
| CoRefineModel=Qwen3-32B, Compute Budget=@5122026.02 | 86.7 | 0.18 | — | |
| CoRefineModel=DeepSeek-8B2026.02 | 86.7 | — | — | |
| CoRefineModel=Qwen3-32B2026.02 | 86.7 | — | — | |
| CoRefineModel=PaCoRe-8B2026.02 | 86.7 | — | — | |
| Ministral-3-Reasoning-2512Params=14B2026.06 | 86.7 | — | — | |
| GPT-OSS-20B (high)Params=20B2026.06 | 86.7 | — | — | |
| Qwen3.5-4BParams=4B2026.06 | 83.5 | — | — | |
| Gemini 2.5 FlashParams=N/A2026.06 | 83.3 | — | — | |
| GPT-5 Nano (high)Params=N/A2026.06 | 80.8 | — | — | |
| PaCoRe-8BStrategy=Pass, Sample Count=@12026.02 | 80.7 | — | — | |
| Gemma-4-itParams=12B2026.06 | 80.4 | — | — | |
| DeepSeek-8BStrategy=Pass, Sample Count=@12026.02 | 80 | — | — | |
| Mimo7B-RL-0530Params=7B2026.06 | 79.8 | — | — | |
| OpenReasoning-NemotronParams=7B2026.06 | 78.8 | — | — | |
| Qwen3-32BStrategy=Pass, Sample Count=@12026.02 | 78.7 | — | — | |
| KnowRL-Nemotron-1.5BHint Setting=CBRS, Evaluation Protocol=mean@322026.04 | 78.33 | — | — | |
| KnowRL-Nemotron-1.5BHint Setting=CSS, Evaluation Protocol=mean@322026.04 | 78.12 | — | — | |
| Qwen3-4B-Thinking-2507Params=4B2026.06 | 77.7 | — | — | |
| QuestAHint Setting=CSS, Evaluation Protocol=mean@322026.04 | 73.75 | — | — | |
| QuestAHint Setting=CBRS, Evaluation Protocol=mean@322026.04 | 73.23 | — | — | |
| JustRLHint Setting=CSS, Evaluation Protocol=mean@322026.04 | 70.67 | — | — | |
| JustRLHint Setting=CBRS, Evaluation Protocol=mean@322026.04 | 70.49 | — | — | |
| KnowRL-Nemotron-1.5BHint Setting=w/o KP, Evaluation Protocol=mean@322026.04 | 69.48 | — | — | |
| Olmo-3-ThinkParams=7B2026.06 | 69 | — | — | |
| QuestAHint Setting=w/o KP, Evaluation Protocol=mean@322026.04 | 67.5 | — | — | |
| JustRLHint Setting=w/o KP, Evaluation Protocol=mean@322026.04 | 66.88 | — | — | |
| Phi4-Reasoning-PlusParams=14B2026.06 | 66.5 | — | — | |
| Nemotron-1.5BHint Setting=CSS, Evaluation Protocol=mean@322026.04 | 65.03 | — | — | |
| Nemotron-1.5BHint Setting=CBRS, Evaluation Protocol=mean@322026.04 | 64.17 | — | — | |
| Hunyuan-4B-InstructParams=4B2026.06 | 62.7 | — | — | |
| Nemotron-1.5BHint Setting=w/o KP, Evaluation Protocol=mean@322026.04 | 60.73 | — | — | |
| SmolLM3Params=3B2026.06 | 49.2 | — | — | |
| PPCVModel=Qwen3-32B, Mode=non-thinking mode2026.02 | 43.33 | — | — | |
| Phi-DecodingModel=Qwen3-32B, Mode=non-thinking mode2026.02 | 36.67 | — | — | |
| DyJRBackbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@256, alpha_JS=0.05, M=82026.03 | 33.4 | — | — | |
| Predictive DecodingModel=Qwen3-32B, Mode=non-thinking mode2026.02 | 33.33 | — | — | |
| Ex-GRPOBackbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@2562026.03 | 32.8 | — | — | |
| DyJR (Forward KL)Backbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@256, Divergence Type=Forward KL2026.03 | 32.6 | — | — | |
| DyJR (alpha_JS = 0.01)Backbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@256, alpha_JS=0.012026.03 | 31 | — | — | |
| DyJR (M = 16)Backbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@256, M=162026.03 | 31 | — | — | |
| DAPOBackbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@2562026.03 | 30.4 | — | — | |
| DyJR (M = 32)Backbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@256, M=322026.03 | 30.2 | — | — | |
| DPH-RLBackbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@2562026.03 | 30.1 | — | — | |
| Chain-of-ThoughtModel=Qwen3-32B, Mode=non-thinking mode2026.02 | 30 | — | — | |
| GRPOBackbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@2562026.03 | 29.6 | — | — | |
| DyJR (M = 64)Backbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@256, M=642026.03 | 29.3 | — | — | |
| DyJR (alpha_JS = 0.2)Backbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@256, alpha_JS=0.22026.03 | 29 | — | — | |
| RLEPBackbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@2562026.03 | 28.9 | — | — | |
| Guided DecodingModel=Qwen3-32B, Mode=non-thinking mode2026.02 | 28.67 | — | — | |
| Base ModelBackbone=Qwen3-4B-Base, Evaluation Protocol (Mean@k)=Mean@2562026.03 | 16.6 | — | — | |
| BaseModel Series=Qwen3-4B, Decoding Mode=Standard decoding2026.02 | — | — | 61.67 | |
| BaseModel Series=Qwen3-8B, Decoding Mode=Standard decoding2026.02 | — | — | 69.58 | |
| BaseModel Series=Qwen3-14B, Decoding Mode=Standard decoding2026.02 | — | — | 72.84 | |
| CCDModel Series=Qwen3-4B2026.02 | — | — | 65 | |
| CCDModel Series=Qwen3-8B2026.02 | — | — | 72.5 | |
| CCDModel Series=Qwen3-14B2026.02 | — | — | 75 | |
| MTIModel Series=Qwen3-4B2026.02 | — | — | 67.08 | |
| MTIModel Series=Qwen3-8B2026.02 | — | — | 70 | |
| MTIModel Series=Qwen3-14B2026.02 | — | — | 73.64 |