STEM Reasoning on MMLU STEM
73.7Accuracy (STEM)TaH+
Evaluation Results
| Method | Links | |
|---|---|---|
| TaH+Param=1.7B2025.11 | 73.7 | |
| StandardParam=1.7B2025.11 | 70.8 | |
| SoftThinkParam=1.7B2025.11 | 70.6 | |
| AlwaysThinkParam=1.7B2025.11 | 63.8 | |
| AGPOclipping=adaptive, ATS=true2026.05 | 58.1 | |
| AGPOclipping=adaptive, ATS=false2026.05 | 57.9 | |
| GRPOclipping=fixed ε2026.05 | 57.6 | |
| GRPOATS=true2026.05 | 57.1 | |
| TaH+Param=0.6B2025.11 | 56.3 | |
| Adaptive-KL PPO2026.05 | 56 | |
| StandardParam=0.6B2025.11 | 51.6 | |
| SoftThinkParam=0.6B2025.11 | 51.4 | |
| PPO2026.05 | 48.8 | |
| DPO2026.05 | 46.9 | |
| AlwaysThinkParam=0.6B2025.11 | 42.6 |