Mathematical Reasoning on MATH 500 (Language Accuracy Scores)
97.75AccuracyGemma-4-31B-it-NVFP4
Evaluation Results
| Method | Links | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Gemma-4-31B-it-NVFP4Quantization=NVFP42026.05 | 97.75 | — | — | — | — | — | — | — | — | — | — | 78.26 | |
| Gemma-4-31B-itQuantization=BF162026.05 | 97.33 | — | — | — | — | — | — | — | — | — | — | 82.36 | |
| Mix-QuantBase Model=Gemma-4-31B-it2026.05 | 97.2 | — | — | — | — | — | — | — | — | — | — | 81.39 | |
| Gemma-4-26B-A4B-itQuantization=BF162026.05 | 95.86 | — | — | — | — | — | — | — | — | — | — | 71.94 | |
| Mix-QuantBase Model=Gemma-4-26B-A4B-it2026.05 | 95.4 | — | — | — | — | — | — | — | — | — | — | 71.93 | |
| Qwen3-8B + ReTool-RLBase Model=Qwen3-8B, Training Stage=RL, Framework=ReTool, Augmentation=None2026.06 | 95.4 | — | — | — | — | — | — | — | — | — | — | — | |
| Gemma-4-26B-A4B-it-NVFP4Quantization=NVFP42026.05 | 95.2 | — | — | — | — | — | — | — | — | — | — | 66.31 | |
| Qwen3.5-9BQuantization=BF162026.05 | 94.85 | — | — | — | — | — | — | — | — | — | — | 72.04 | |
| Qwen-2.5-7B-REVESProtocol=O-322026.06 | 94.7 | — | — | — | — | — | — | — | — | — | — | — | |
| Mix-QuantBase Model=Qwen3-8B2026.05 | 94.4 | — | — | — | — | — | — | — | — | — | — | 61 | |
| Mix-QuantBase Model=Qwen3.5-9B2026.05 | 94.32 | — | — | — | — | — | — | — | — | — | — | 70.59 | |
| Qwen3-8B-NVFP4Quantization=NVFP42026.05 | 94.12 | — | — | — | — | — | — | — | — | — | — | 55.22 | |
| Qwen3-8BQuantization=BF162026.05 | 93.73 | — | — | — | — | — | — | — | — | — | — | 62.11 | |
| Qwen3.5-9B-NVFP4Quantization=NVFP42026.05 | 93.46 | — | — | — | — | — | — | — | — | — | — | 63.26 | |
| Qwen3-8B + ReTool-RL w/ CoBeBase Model=Qwen3-8B, Training Stage=RL, Framework=ReTool, Augmentation=CoBe2026.06 | 91.9 | — | — | — | — | — | — | — | — | — | — | — | |
| GSPO + REINFORCE++Objective=GSPO, G (Rollouts per prompt)=1, Steps=300, Time=13.0h2026.05 | 90.8 | — | — | — | — | — | — | — | — | — | — | — | |
| GSPO + TACBackbone=DeepSeek-R1-Distill-Qwen-7B, Training steps=30002025.10 | 90.4 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen-2.5-3B-REVESProtocol=O-322026.06 | 89.9 | — | — | — | — | — | — | — | — | — | — | — | |
| GSPO (G=8)Objective=GSPO, G (Rollouts per prompt)=8, Steps=300, Time=14.4h2026.05 | 89.6 | — | — | — | — | — | — | — | — | — | — | — | |
| GRPO + BASISObjective=GRPO, G (Rollouts per prompt)=1, Steps=150, Time=8.3h2026.05 | 89.2 | — | — | — | — | — | — | — | — | — | — | — | |
| GSPO + BASISObjective=GSPO, G (Rollouts per prompt)=1, Steps=150, Time=7.2h2026.05 | 88.8 | — | — | — | — | — | — | — | — | — | — | — | |
| GPG + BASISObjective=GPG, G (Rollouts per prompt)=1, Steps=150, Time=6.4h2026.05 | 88.4 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-8B + ReTool-SFTBase Model=Qwen3-8B, Training Stage=SFT, Framework=ReTool, Augmentation=None2026.06 | 88.2 | — | — | — | — | — | — | — | — | — | — | — | |
| GRPO (G=8)Objective=GRPO, G (Rollouts per prompt)=8, Steps=300, Time=15.5h2026.05 | 88 | — | — | — | — | — | — | — | — | — | — | — | |
| GPG (G=8)Objective=GPG, G (Rollouts per prompt)=8, Steps=300, Time=14.0h2026.05 | 88 | — | — | — | — | — | — | — | — | — | — | — | |
| GPG + REINFORCE++Objective=GPG, G (Rollouts per prompt)=1, Steps=300, Time=12.9h2026.05 | 87.4 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen-2.5-7B-Multi-turnProtocol=O-322026.06 | 87.1 | — | — | — | — | — | — | — | — | — | — | — | |
| SkyWorkBackbone=DeepSeek-R1-Distill-Qwen-7B, Training steps=30002025.10 | 87 | — | — | — | — | — | — | — | — | — | — | — | |
| GSPO VanillaObjective=GSPO, G (Rollouts per prompt)=1, Steps=300, Time=13.0h2026.05 | 86.8 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen-2.5-7B-PAGProtocol=O-322026.06 | 86.7 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen-2.5-7B-RLProtocol=O-322026.06 | 85.9 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen-2.5-7B-REVESProtocol=O-42026.06 | 85.6 | — | — | — | — | — | — | — | — | — | — | — | |
| GRPO-LEADBackbone=DeepSeek-R1-Distill-Qwen-7B, Training steps=30002025.10 | 84.6 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-Coder-7B-Instruct + ReTool-RL w/ CoBeBase Model=Qwen2.5-Coder-7B-Instruct, Training Stage=RL, Framework=ReTool, Augmentation=CoBe2026.06 | 84.5 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-Coder-7B-Instruct + ReTool-RLBase Model=Qwen2.5-Coder-7B-Instruct, Training Stage=RL, Framework=ReTool, Augmentation=None2026.06 | 84.4 | — | — | — | — | — | — | — | — | — | — | — | |
| DeepMathBackbone=DeepSeek-R1-Distill-Qwen-7B, Training steps=30002025.10 | 83.4 | — | — | — | — | — | — | — | — | — | — | — | |
| ExGRPOBackbone=DeepSeek-R1-Distill-Qwen-7B, Training steps=30002025.10 | 82.8 | — | — | — | — | — | — | — | — | — | — | — | |
| GRPO + REINFORCE++Objective=GRPO, G (Rollouts per prompt)=1, Steps=300, Time=17.7h2026.05 | 82.6 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-7B-Instruct + ReTool-RL w/ CoBeBase Model=Qwen2.5-7B-Instruct, Training Stage=RL, Framework=ReTool, Augmentation=CoBe2026.06 | 81.9 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-8B + ReTool-SFT w/ CoBeBase Model=Qwen3-8B, Training Stage=SFT, Framework=ReTool, Augmentation=CoBe2026.06 | 81.5 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen-2.5-3B-Multi-turnProtocol=O-322026.06 | 81.2 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen-2.5-7B-PAGProtocol=O-42026.06 | 80.8 | — | — | — | — | — | — | — | — | — | — | — | |
| BASELINEBackbone=DeepSeek-R1-Distill-Qwen-7B, Training steps=30002025.10 | 80.6 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen-2.5-3B-RLProtocol=O-322026.06 | 80.5 | — | — | — | — | — | — | — | — | — | — | — | |
| SFT→RLBackbone=Qwen3-8B-Base2025.09 | 80.4 | — | — | — | — | — | — | — | — | — | — | 45.5 | |
| DRGRPOBackbone=DeepSeek-R1-Distill-Qwen-7B, Training steps=30002025.10 | 80.2 | — | — | — | — | — | — | — | — | — | — | — | |
| VeriThinkerBackbone=DeepSeek-R1-Distill-Qwen-7B, Training steps=30002025.10 | 80.2 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen-2.5-7B-Multi-turnProtocol=O-42026.06 | 80.1 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen-2.5-7B-RLProtocol=O-42026.06 | 79.4 | — | — | — | — | — | — | — | — | — | — | — | |
| Eurus-2Backbone=DeepSeek-R1-Distill-Qwen-7B, Training steps=30002025.10 | 79.2 | — | — | — | — | — | — | — | — | — | — | — | |
| BRIDGEBackbone=Qwen3-8B-Base2025.09 | 79 | — | — | — | — | — | — | — | — | — | — | 49.9 | |
| VFModel Family=Qwen2.5-Instruct, Size (B)=142025.11 | 78.8 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen-2.5-3B-PAGProtocol=O-322026.06 | 78.8 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen-2.5-7B-RLProtocol=SC-42026.06 | 78.6 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen-2.5-7B-Multi-turnProtocol=SC-42026.06 | 78 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen-2.5-3B-REVESProtocol=O-42026.06 | 77.8 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen-2.5-7B-PAGProtocol=SC-42026.06 | 77.2 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-7B-Instruct + ReTool-RLBase Model=Qwen2.5-7B-Instruct, Training Stage=RL, Framework=ReTool, Augmentation=None2026.06 | 77.1 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen-2.5-7B-REVESProtocol=1-shot2026.06 | 76.8 | — | — | — | — | — | — | — | — | — | — | — | |
| CHORDBackbone=Qwen3-8B-Base2025.09 | 76.6 | — | — | — | — | — | — | — | — | — | — | 45.9 | |
| RLBackbone=Qwen3-8B-Base2025.09 | 76.2 | — | — | — | — | — | — | — | — | — | — | 42.9 | |
| Qwen-2.5-7B-REVESProtocol=SC-42026.06 | 76.2 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen-2.5-7B-RLProtocol=1-shot2026.06 | 76.1 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen-2.5-7B-PAGProtocol=1-shot2026.06 | 76.1 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen-2.5-7B-Multi-turnProtocol=1-shot2026.06 | 75.9 | — | — | — | — | — | — | — | — | — | — | — | |
| REPCategory=REP exposed trace, Victim / Teacher=Qwen3-14B, Student supervision=Exposed trace, answer-clean2026.05 | 75.8 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-7B-Instruct + ReTool-SFTBase Model=Qwen2.5-7B-Instruct, Training Stage=SFT, Framework=ReTool, Augmentation=None2026.06 | 75.7 | — | — | — | — | — | — | — | — | — | — | — | |
| CoTModel Family=Qwen2.5-Instruct, Size (B)=142025.11 | 75.6 | — | — | — | — | — | — | — | — | — | — | — | |
| LUFFYBackbone=Qwen3-8B-Base2025.09 | 75.4 | — | — | — | — | — | — | — | — | — | — | 44 | |
| UABBackbone Model=Cohere, Budget=N=42026.05 | 75 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-7B-Instruct + ReTool-SFT w/ CoBeBase Model=Qwen2.5-7B-Instruct, Training Stage=SFT, Framework=ReTool, Augmentation=CoBe2026.06 | 74.5 | — | — | — | — | — | — | — | — | — | — | — | |
| LLM-JudgeBackbone Model=Cohere, Budget=N=42026.05 | 74.1 | — | — | — | — | — | — | — | — | — | — | — | |
| GSPO + TACBackbone=DeepSeek-R1-Distill-Qwen-1.5B, Training steps=30002025.10 | 74 | — | — | — | — | — | — | — | — | — | — | — | |
| UniformBackbone Model=Cohere, Budget=N=42026.05 | 74 | — | — | — | — | — | — | — | — | — | — | — | |
| REPCategory=REP exposed trace, Victim / Teacher=Qwen3-32B, Student supervision=Exposed trace, all valid2026.05 | 73.9 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-Coder-7B-Instruct + ReTool-SFTBase Model=Qwen2.5-Coder-7B-Instruct, Training Stage=SFT, Framework=ReTool, Augmentation=None2026.06 | 73.9 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-8BBase Model=Qwen3-8B, Training Stage=None, Framework=None, Augmentation=None2026.06 | 73.7 | — | — | — | — | — | — | — | — | — | — | — | |
| LLM-JudgeBackbone Model=Gemma3-27B, Budget=N=42026.05 | 73.3 | — | — | — | — | — | — | — | — | — | — | — | |
| UABBackbone Model=Gemma3-27B, Budget=N=42026.05 | 73.2 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-Coder-7B-Instruct + ReTool-SFT w/ CoBeBase Model=Qwen2.5-Coder-7B-Instruct, Training Stage=SFT, Framework=ReTool, Augmentation=CoBe2026.06 | 73.2 | — | — | — | — | — | — | — | — | — | — | — | |
| UniformBackbone Model=Gemma3-27B, Budget=N=42026.05 | 73 | — | — | — | — | — | — | — | — | — | — | — | |
| REPCategory=REP exposed trace, Victim / Teacher=Qwen3-32B, Student supervision=Exposed trace, answer-clean2026.05 | 72.8 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen-2.5-3B-Multi-turnProtocol=O-42026.06 | 72.7 | — | — | — | — | — | — | — | — | — | — | — | |
| STILL-3Backbone=DeepSeek-R1-Distill-Qwen-1.5B, Training steps=30002025.10 | 72.4 | — | — | — | — | — | — | — | — | — | — | — | |
| REPCategory=REP exposed trace, Victim / Teacher=Qwen3-14B, Student supervision=Exposed trace, all valid2026.05 | 72.4 | — | — | — | — | — | — | — | — | — | — | — | |
| SFT+RLBackbone=Qwen3-8B-Base2025.09 | 72.2 | — | — | — | — | — | — | — | — | — | — | 40.1 | |
| SRFTBackbone=Qwen3-8B-Base2025.09 | 72.2 | — | — | — | — | — | — | — | — | — | — | 39.8 | |
| TIAVictim / Teacher=Qwen3-14B, Student supervision=Trace inversion attack2026.05 | 72 | — | — | — | — | — | — | — | — | — | — | — | |
| TIAVictim / Teacher=Qwen3-32B, Student supervision=Trace inversion attack2026.05 | 71.4 | — | — | — | — | — | — | — | — | — | — | — | |
| ExGRPOBackbone=DeepSeek-R1-Distill-Qwen-1.5B, Training steps=30002025.10 | 71.2 | — | — | — | — | — | — | — | — | — | — | — | |
| DRA-GRPOBackbone=DeepSeek-R1-Distill-Qwen-1.5B, Training steps=30002025.10 | 71 | — | — | — | — | — | — | — | — | — | — | — | |
| No distillationVictim / Teacher=–, Student supervision=–2026.05 | 71 | — | — | — | — | — | — | — | — | — | — | — | |
| Oracle internal traceVictim / Teacher=Qwen3-14B, Student supervision=Internal trace2026.05 | 70.3 | — | — | — | — | — | — | — | — | — | — | — | |
| Oracle internal traceVictim / Teacher=Qwen3-32B, Student supervision=Internal trace2026.05 | 70 | — | — | — | — | — | — | — | — | — | — | — | |
| UABModel=Qwen2.5-7B, N=42026.05 | 69.9 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen-2.5-3B-PAGProtocol=O-42026.06 | 69.9 | — | — | — | — | — | — | — | — | — | — | — | |
| Open-RS3Backbone=DeepSeek-R1-Distill-Qwen-1.5B, Training steps=30002025.10 | 69.8 | — | — | — | — | — | — | — | — | — | — | — | |
| Control supervisionVictim / Teacher=Qwen3-32B, Student supervision=Summary of reasoning trace2026.05 | 69.8 | — | — | — | — | — | — | — | — | — | — | — | |
| LengthModel=Qwen2.5-7B, N=42026.05 | 69.6 | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen-2.5-3B-RLProtocol=O-42026.06 | 69.6 | — | — | — | — | — | — | — | — | — | — | — |