Mathematical Reasoning on MMLU Mathematics (test)
57.69Average AccuracySFT
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| SFTModel=Qwen3-8B2026.03 | 57.69 | — | — | — | — | — | — | 0.0769 | |
| VISAModel=Qwen3-8B2026.03 | 57.6 | — | — | — | — | — | — | 0.0437 | |
| BaseModel=Qwen3-8B2026.03 | 52.24 | — | — | — | — | — | — | 0 | |
| DR+DADepth=32, 0-SHOT=true2026.01 | 50 | — | — | — | — | — | 4 | — | |
| DRDepth=32, 0-SHOT=true2026.01 | 45.3 | — | — | — | — | — | 1.9 | — | |
| DRDepth=16, 0-SHOT=true2026.01 | 44.5 | — | — | — | — | — | 4.8 | — | |
| GAL 120B <work>Params (bn)=120, Shots=0-shot, Reasoning strategy=<work> token prompt2022.11 | 41.3 | 27 | 54.2 | 37 | 44 | 40.5 | — | — | |
| VISAModel=Qwen3-0.6B2026.03 | 38.55 | — | — | — | — | — | — | 0.0775 | |
| DR+DADepth=16, 0-SHOT=true2026.01 | 38.5 | — | — | — | — | — | 3.4 | — | |
| SFTModel=Qwen3-0.6B2026.03 | 38.07 | — | — | — | — | — | — | 0.1191 | |
| LADepth=32, 0-SHOT=true2026.01 | 37.2 | — | — | — | — | — | 1 | — | |
| GAL 30B <work>Params (bn)=30, Shots=0-shot, Reasoning strategy=<work> token prompt2022.11 | 37.1 | 33 | 41.5 | 33.3 | 39 | 37.3 | — | — | |
| BaseModel=Qwen3-0.6B2026.03 | 36.45 | — | — | — | — | — | — | 0 | |
| GAL 120BParams (bn)=120, Shots=0-shot, Reasoning strategy=Standard prompt2022.11 | 35.8 | 33 | 38.1 | 32.6 | 43 | 32.5 | — | — | |
| ChinchillaParams (bn)=70, Shots=5-shot2022.11 | 35.7 | 31 | 41.5 | 31.9 | 32 | 33.3 | — | — | |
| GopherParams (bn)=280, Shots=5-shot2022.11 | 30.6 | 25 | 33.6 | 23.7 | 37 | 35.7 | — | — | |
| GAL 30BParams (bn)=30, Shots=0-shot, Reasoning strategy=Standard prompt2022.11 | 29.9 | 30 | 30.2 | 26.3 | 36 | 31.7 | — | — | |
| GAL 6.7BParams (bn)=6.7, Shots=0-shot, Reasoning strategy=Standard prompt2022.11 | 29.2 | 28 | 28.9 | 26.7 | 36 | 31 | — | — | |
| GAL 6.7B <work>Params (bn)=6.7, Shots=0-shot, Reasoning strategy=<work> token prompt2022.11 | 28 | 33.3 | 30.7 | 25.2 | 26 | 33.3 | — | — | |
| GAL 1.3BParams (bn)=1.3, Shots=0-shot, Reasoning strategy=Standard prompt2022.11 | 27.1 | 28 | 27.2 | 26.7 | 30 | 24.6 | — | — | |
| OPTParams (bn)=175, Shots=5-shot2022.11 | 26.7 | 21 | 25.7 | 24.4 | 33 | 29.4 | — | — | |
| LADepth=16, 0-SHOT=true2026.01 | 26.6 | — | — | — | — | — | 1 | — | |
| BLOOMParams (bn)=176, Shots=5-shot2022.11 | 26.4 | 25 | 26.7 | 27 | 25 | 26.2 | — | — | |
| GAL 1.3B <work>Params (bn)=1.3, Shots=0-shot, Reasoning strategy=<work> token prompt2022.11 | 24.6 | 22 | 24.6 | 18.9 | 25 | 31 | — | — |