Mathematical Reasoning on AIME 2025 (Accuracy and Average Tokens)
96.7AccuracyPC-cubic
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| PC-cubicModel=Nemotron3-30B, Token Budget (B)=5M2026.05 | 96.7 | — | |
| GPT-5-High2026.05 | 94.6 | — | |
| SU-01Number of runs=82026.05 | 94.6 | — | |
| Nemotron-Cascade-2Number of runs=82026.05 | 94.2 | — | |
| PC-cubicModel=GPT-OSS-120B, Token Budget (B)=5M2026.05 | 94.1 | — | |
| Standard MVModel=Nemotron3-30B, Token Budget (B)=250k2026.05 | 93.5 | — | |
| DeepSeek-V3.2Active=37B, Total=671B2026.05 | 93.1 | — | |
| Qwen3.6-35B-A3BNumber of runs=82026.05 | 92.5 | — | |
| Qwen3-235B-A22B-Thinking-2507Active=22B, Total=235B2026.05 | 92.3 | — | |
| ZAYA1-8B + Markovian RSA (40K/4K)Active=0.7B, Total=8.0B, Configuration=40K/4K2026.05 | 91.9 | — | |
| GLM-4.7-FlashNumber of runs=82026.05 | 91.3 | — | |
| P1-30B-A3BNumber of runs=82026.05 | 90.4 | — | |
| Standard MVModel=GPT-OSS-120B, Token Budget (B)=250k2026.05 | 90.1 | — | |
| Gemma-4-31BNumber of runs=82026.05 | 88.8 | — | |
| ZAYA1-8B (single rollout)Active=0.7B, Total=8.0B, Configuration=single rollout2026.05 | 88.3 | — | |
| Gemini-2.5 Pro2026.05 | 88 | — | |
| DeepSeek-R1-0528Active=37B, Total=671B2026.05 | 87.5 | — | |
| RecursiveMASrecursion round=32026.04 | 86.7 | — | |
| Full AttnAttention Strategy=Full Attention2026.05 | 86.67 | — | |
| RTPurboselection_strategy=top-p2026.05 | 86.67 | — | |
| BCR-Qwen3-4B (Ours)Inference mode=N=1, Temperature=0.6, Top-p=0.9, Max generation length=32,7682026.04 | 83.3 | 17,498 | |
| Upper BoundModel=Qwen3-8B, k=62025.08 | 82.22 | — | |
| LaTER (training)Backbone=Qwen3-14B, Protocol=fully trained2026.05 | 80 | 10,575 | |
| RTPurboselection_strategy=top-k2026.05 | 80 | — | |
| Step-GRPOBackbone=Qwen3-8B2026.04 | 73.3 | 12,859 | |
| Single Agenttraining protocol=Full-SFT2026.04 | 73.3 | — | |
| TextGradrecursion round=32026.04 | 73.3 | — | |
| Recursive-TextMASrecursion round=32026.04 | 73.3 | — | |
| CoT-SFTBackbone=Qwen3-14B, Protocol=SFT with CE and KL objectives2026.05 | 73.3 | 12,687 | |
| LaTER (training-free)Backbone=Qwen3-14B, Protocol=training-free2026.05 | 73.3 | 10,661 | |
| FullModel=Qwen3-14B, KV Cache Budget=100%2026.06 | 70.21 | — | |
| Qwen3-4B-Thinking-2507Inference mode=N=1, Temperature=0.6, Top-p=0.9, Max generation length=32,7682026.04 | 70 | 20,773 | |
| Single Agenttraining protocol=LoRA2026.04 | 70 | — | |
| Upper BoundModel=DS-Distill-Qwen-2.5-7B, k=62025.08 | 70 | — | |
| CoT BaselineBackbone=Qwen3-14B, Protocol=standard explicit CoT prompting2026.05 | 70 | 15,730 | |
| PiCSARModel=Qwen3-8B, k=62025.08 | 68.89 | — | |
| AverageModel=Qwen3-8B, k=62025.08 | 67.04 | — | |
| GRPOBackbone=Qwen3-8B2026.04 | 66.7 | 15,407 | |
| GRPO+SOPBackbone=Qwen3-8B2026.04 | 66.7 | 11,844 | |
| LoopLMrecursion round=32026.04 | 66.7 | — | |
| FullModel=Qwen3-4B, KV Cache Budget=100%2026.06 | 66.41 | — | |
| Self-ConsistencyModel=Qwen3-8B, k=62025.08 | 65.56 | — | |
| SeerAttention-RModel=Qwen3-14B, KV Cache Budget=~25%2026.06 | 64.79 | — | |
| VASE-ATTNVModel=Qwen3-14B, KV Cache Budget=~25%2026.06 | 63.96 | — | |
| VASE-DKVModel=Qwen3-14B, KV Cache Budget=~25%2026.06 | 63.33 | — | |
| GRPO-λBackbone=Qwen3-8B2026.04 | 63.3 | 12,321 | |
| VanillaBackbone=Qwen3-4B2026.04 | 63.3 | 18,049 | |
| GRPOBackbone=Qwen3-4B2026.04 | 63.3 | 14,249 | |
| GRPO+LPBackbone=Qwen3-4B2026.04 | 63.3 | 10,783 | |
| PC-cubicModel=Ministral3-14B, Token Budget (B)=5M2026.05 | 62.1 | — | |
| Qwen2.5-32B-Instruct + Bootcamp-SFT-RLModel=Qwen2.5-32B-Instruct, Training Stage=Bootcamp-SFT-RL2025.08 | 60.5 | — | |
| Full AttentionBackbone=GPT-OSS2026.04 | 60 | — | |
| VanillaBackbone=Qwen3-8B2026.04 | 60 | 17,492 | |
| GRPO+LPBackbone=Qwen3-8B2026.04 | 60 | 11,932 | |
| GRPO-8kBackbone=Qwen3-4B2026.04 | 60 | 11,172 | |
| GRPO+SOPBackbone=Qwen3-4B2026.04 | 60 | 13,838 | |
| GRPO-λBackbone=Qwen3-4B2026.04 | 60 | 12,814 | |
| Step-GRPOBackbone=Qwen3-4B2026.04 | 60 | 13,356 | |
| Mixture-of-Agents (MoA)recursion round=32026.04 | 60 | — | |
| Qwen3-8B + ReTool-RLBase Model=Qwen3-8B, Training Stage=RL, Framework=ReTool, Augmentation=None2026.06 | 59.2 | — | |
| VASE-ATTNVModel=Qwen3-4B, KV Cache Budget=~25%2026.06 | 59.17 | — | |
| SeerAttention-RModel=Qwen3-4B, KV Cache Budget=~25%2026.06 | 58.59 | — | |
| VASE-DKVModel=Qwen3-4B, KV Cache Budget=~25%2026.06 | 57.29 | — | |
| DS-R1-Distilled-Qwen-32B + Bootcamp-RLModel=DS-R1-Distilled-Qwen-32B, Training Stage=Bootcamp-RL2025.08 | 56.8 | — | |
| DiffMASModel=Qwen3-8B, Communication Category=Trained Latent Communication2026.04 | 56.7 | — | |
| R-KVModel=Qwen3-14B, KV Cache Budget=~25%2026.06 | 54.58 | — | |
| CurDKVModel=Qwen3-14B, KV Cache Budget=~25%2026.06 | 53.75 | — | |
| GRPO-8kBackbone=Qwen3-8B2026.04 | 53.3 | 12,638 | |
| TextMASModel=Qwen3-8B, Communication Category=Text Communication2026.04 | 53.3 | — | |
| LatentMASModel=Qwen3-8B, Communication Category=Training-free Latent Communication2026.04 | 53.3 | — | |
| OPDTeacher-Student Pair=DeepSeek-R1-Distill-Qwen-7B / Skywork-OR1-7B, Time=17.5h, Red.=0%2026.05 | 52.5 | — | |
| DS-R1-Distilled-Qwen-32BModel=DS-R1-Distilled-Qwen-32B2025.08 | 52.5 | — | |
| PRUNE-OPD (overlap)Teacher-Student Pair=DeepSeek-R1-Distill-Qwen-7B / Skywork-OR1-7B, Time=17.0h, Red.=2.9%2026.05 | 52.1 | — | |
| Upper BoundModel=DS-Distill-llama-3-8B, k=62025.08 | 51.11 | — | |
| PiCSARModel=DS-Distill-Qwen-2.5-7B, k=62025.08 | 51.11 | — | |
| Standard MVModel=Ministral3-14B, Token Budget (B)=250k2026.05 | 50.8 | — | |
| DEER+SFTBackbone=Qwen3-8B2026.04 | 50 | 13,330 | |
| GSPOBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.04 | 50 | — | |
| LatentMASModel=Qwen3-4B, Communication Category=Training-free Latent Communication2026.04 | 50 | — | |
| DiffMASModel=Qwen3-4B, Communication Category=Trained Latent Communication2026.04 | 50 | — | |
| CIKAParameter Count=7B, Status=Frozen2026.05 | 50 | — | |
| EVOTDEvaluation Protocol=Pass@8, Training Iteration=Iter32026.05 | 49.9 | — | |
| R-KVModel=Qwen3-4B, KV Cache Budget=~25%2026.06 | 49.58 | — | |
| TriAttentionBackbone=GPT-OSS2026.04 | 49.2 | — | |
| GCPOBackbone=Qwen3-4B2026.05 | 49 | — | |
| OPD (Truncate 4k)Teacher-Student Pair=DeepSeek-R1-Distill-Qwen-7B / Skywork-OR1-7B, Time=11.3h, Red.=35.4%2026.05 | 48.8 | — | |
| CurDKVModel=Qwen3-4B, KV Cache Budget=~25%2026.06 | 48.75 | — | |
| SnapKVModel=Qwen3-4B, KV Cache Budget=~25%2026.06 | 47.92 | — | |
| EVOTDEvaluation Protocol=Pass@8, Training Iteration=Iter22026.05 | 47.6 | — | |
| EVOTDEvaluation Protocol=Pass@8, Training Iteration=Iter12026.05 | 47.3 | — | |
| e3-1.7BInference mode=N=1, Temperature=0.6, Top-p=0.9, Max generation length=32,768, Length Control=true2026.04 | 46.7 | 11,804 | |
| DEER+SFTBackbone=Qwen3-4B2026.04 | 46.7 | 12,503 | |
| SingleModel=Qwen3-8B, Communication Category=Text Communication2026.04 | 46.7 | — | |
| GRPOBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.04 | 46.67 | — | |
| SPSBackbone=DeepSeek-R1-Distill-Qwen-1.5B2026.04 | 46.67 | — | |
| QuestAttention Strategy=Sparse Attention2026.05 | 46.67 | — | |
| SnapKVAttention Strategy=Sparse Attention2026.05 | 46.67 | — | |
| DQOBackbone=Qwen3-4B2026.05 | 46.5 | — | |
| Qwen3-8B + ReTool-RL w/ CoBeBase Model=Qwen3-8B, Training Stage=RL, Framework=ReTool, Augmentation=CoBe2026.06 | 46.3 | — | |
| DIVERBackbone=Qwen3-4B2026.05 | 44.2 | — |