Mathematical Reasoning on GPQA (Acc.@first, Acc.@final)
56.5AccuracyPOLIS
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| POLISModel=Gemma3-4b2025.07 | 56.5 | — | — | |
| POLISModel=Qwen3-1.7b2025.07 | 49.1 | — | — | |
| MAGPIEModel=Gemma3-4b2025.07 | 48.1 | — | — | |
| POLISModel=Phi-4-mini2025.07 | 47.4 | — | — | |
| SGDSModel=Gemma3-4b2025.07 | 46.5 | — | — | |
| MAGPIEModel=Phi-4-mini2025.07 | 45.5 | — | — | |
| DARTModel=Gemma3-4b2025.07 | 45 | — | — | |
| SGDSModel=Phi-4-mini2025.07 | 44 | — | — | |
| GRAModel=Gemma3-4b2025.07 | 44 | — | — | |
| MAGPIEModel=Qwen3-1.7b2025.07 | 43.2 | — | — | |
| MuggleMathModel=Gemma3-4b2025.07 | 42.8 | — | — | |
| DARTModel=Phi-4-mini2025.07 | 42.5 | — | — | |
| STaRModel=Gemma3-4b2025.07 | 42.4 | — | — | |
| BaselineModel=Gemma3-4b2025.07 | 42 | — | — | |
| SGDSModel=Qwen3-1.7b2025.07 | 41.5 | — | — | |
| GRAModel=Phi-4-mini2025.07 | 41 | — | — | |
| MAGPIEModel=BitCPM4-1b2025.07 | 40.1 | — | — | |
| MuggleMathModel=Phi-4-mini2025.07 | 40.1 | — | — | |
| STaRModel=Phi-4-mini2025.07 | 39 | — | — | |
| BaselineModel=Phi-4-mini2025.07 | 38.5 | — | — | |
| DARTModel=Qwen3-1.7b2025.07 | 38 | — | — | |
| SGDSModel=BitCPM4-1b2025.07 | 37.8 | — | — | |
| GRAModel=Qwen3-1.7b2025.07 | 36.5 | — | — | |
| DARTModel=BitCPM4-1b2025.07 | 34.5 | — | — | |
| MuggleMathModel=Qwen3-1.7b2025.07 | 34.5 | — | — | |
| POLISModel=BitCPM4-1b2025.07 | 33.7 | — | — | |
| STaRModel=Qwen3-1.7b2025.07 | 32.8 | — | — | |
| GRAModel=BitCPM4-1b2025.07 | 32.5 | — | — | |
| BaselineModel=Qwen3-1.7b2025.07 | 32 | — | — | |
| MuggleMathModel=BitCPM4-1b2025.07 | 30.2 | — | — | |
| STaRModel=BitCPM4-1b2025.07 | 29.1 | — | — | |
| BaselineModel=BitCPM4-1b2025.07 | 28.5 | — | — | |
| Sumi-7BParadigm=Uniform Diffusion, Training Tokens=1.5T, Training Data=Fully Released, Evaluation protocol=Evaluated under our protocol, Shots=52026.06 | 26.1 | — | — | |
| Llama 3-8BParadigm=AR, Training Tokens=15T, Training Data=Not Released, Evaluation protocol=Reported by prior work, Shots=52026.06 | 25.9 | — | — | |
| LLaDA-8BParadigm=Masked Diffusion, Training Tokens=2.3T, Training Data=Not Released, Evaluation protocol=Reported by prior work, Shots=52026.06 | 25.2 | — | — | |
| OLMo-7BParadigm=AR, Training Tokens=2.5T, Training Data=Fully Released, Evaluation protocol=Evaluated under our protocol, Shots=52026.06 | 24.8 | — | — | |
| Falcon-7BParadigm=AR, Training Tokens=1.5T, Training Data=Partially Released, Evaluation protocol=Evaluated under our protocol, Shots=52026.06 | 24.6 | — | — | |
| Llama 2-7BParadigm=AR, Training Tokens=2T, Training Data=Not Released, Evaluation protocol=Evaluated under our protocol, Shots=52026.06 | 24.3 | — | — | |
| Base ModelModel=Qwen3-1.7B-Base2026.03 | — | 27.5 | 27.3 | |
| Base ModelModel=Llama-3.2-3B-Instruct2026.03 | — | 26.9 | 26.1 | |
| Base ModelModel=Qwen2.5-7B2026.03 | — | 29.9 | 29.7 | |
| CoVerRLModel=Qwen3-1.7B-Base2026.03 | — | 32.9 | 33.6 | |
| CoVerRLModel=Llama-3.2-3B-Instruct2026.03 | — | 32.3 | 32.6 | |
| CoVerRLModel=Qwen2.5-7B2026.03 | — | 36.2 | 37.2 | |
| TTRLModel=Qwen3-1.7B-Base2026.03 | — | 30.9 | 30.7 | |
| TTRLModel=Llama-3.2-3B-Instruct2026.03 | — | 29.8 | 28.2 | |
| TTRLModel=Qwen2.5-7B2026.03 | — | 35.8 | 35.6 |