Mathematics on GSM8K
94.62GSM8K ScoreJT-Safe-V2-35B
Evaluation Results
| Method | Links | |
|---|---|---|
| JT-Safe-V2-35BParameters=35B2026.05 | 94.62 | |
| SOTA with Equivalent ParametersModel comparison=Equivalent Parameters2026.05 | 93.93 | |
| Qwen2-72B-InstructType=Instruction-tuned2024.07 | 93.2 | |
| Llama-3-70B-InstructType=Instruction-tuned2024.07 | 93 | |
| Mi:dm 2.0 Base-instType=Instruction-tuned2026.01 | 91.6 | |
| SKillWeave (Ours)#Params=10B2026.05 | 91 | |
| EMR-Merging#Params=16.7B2026.05 | 90.8 | |
| SkillWeave (Delta-Come replacement)#Params=10B2026.05 | 90.7 | |
| Qwen3-4BParameters=4B2026.01 | 90.4 | |
| Gemma2-27B-it#Params=27B2026.05 | 90.4 | |
| Qwen2.5-14B#Params=14B2026.05 | 90.2 | |
| TALL-Mask#Params=16.7B2026.05 | 90.1 | |
| SkillWeave (ASVD replacement)#Params=10B2026.05 | 89.7 | |
| Twin-Merging r1024#Params=21B2026.05 | 89.3 | |
| SkillWeave (Self-Rewarding replacement)#Params=10B2026.05 | 89.2 | |
| Mixtral-8x22B-InstructType=Instruction-tuned2024.07 | 89.1 | |
| PCB-Merge+DARE#Params=8B2026.05 | 88.9 | |
| Qwen3-14BParameters=14B2026.01 | 88 | |
| FuseChat3.0#Params=8B2026.05 | 88 | |
| Routed LoRA r1024#Params=21B2026.05 | 87.9 | |
| PCB-Merging#Params=8B2026.05 | 87.7 | |
| Twin-Merging r512#Params=14.1B2026.05 | 87.6 | |
| Jointly MTL#Params=8B2026.05 | 87.5 | |
| SkillWeave (Self-Specialize replacement)#Params=10B2026.05 | 87.3 | |
| Ties-Merging#Params=8B2026.05 | 87.2 | |
| Self-MoE#Params=9B2026.05 | 87 | |
| SkillWeave (PEFT replacement)#Params=10B2026.05 | 86.8 | |
| Sequentially MTL#Params=8B2026.05 | 86.7 | |
| Task Arithmetic#Params=8B2026.05 | 86.4 | |
| Routed LoRA r512#Params=14.1B2026.05 | 86.4 | |
| FuseLLM#Params=8B2026.05 | 85.6 | |
| Self-Align#Params=8B2026.05 | 85.5 | |
| Qwen2-57BA14B-it#Params=52B2026.05 | 85.3 | |
| Qwen1.5-110B-ChatType=Instruction-tuned2024.07 | 84.5 | |
| Llama3.1-8B-Instruct#Params=8B2026.05 | 84.5 | |
| Self-Rewarding#Params=8B2026.05 | 84.3 | |
| Mi:dm 2.0 Mini-instType=Instruction-tuned2026.01 | 83.1 | |
| Full AttentionSetting=CPT setting2026.04 | 82.75 | |
| Qwen1.5-72B-ChatType=Instruction-tuned2024.07 | 82.7 | |
| Qwen1.5-72B-Chat#Params=72B2026.05 | 82.7 | |
| Exaone-3.5-2.4B-instParameters=2.4B, Type=Instruction-tuned2026.01 | 82.5 | |
| Hybrid-SWASetting=CPT setting2026.04 | 81.92 | |
| Llama-3.1-8B-instParameters=8B, Type=Instruction-tuned2026.01 | 81.2 | |
| Exaone-3.5-7.8B-instParameters=7.8B, Type=Instruction-tuned2026.01 | 81.1 | |
| KSASetting=CPT setting2026.04 | 81.09 | |
| Hybrid-SCASetting=CPT setting2026.04 | 80.1 | |
| Hybrid-KSASetting=CPT setting2026.04 | 79.5 | |
| Hybrid-LinearSetting=CPT setting2026.04 | 72.44 | |
| UM-190kBackbone=Apertus-8B-SFT, Evaluation Protocol=5-shot2025.11 | 62.15 | |
| Qwen2-1.5Bnumber of parameters=1.5B2024.07 | 61.6 | |
| UM-187kBackbone=Apertus-8B-SFT, Evaluation Protocol=5-shot2025.11 | 61.29 | |
| UltraFBBackbone=Apertus-8B-SFT, Evaluation Protocol=5-shot2025.11 | 60.8 | |
| UM-170kBackbone=Apertus-8B-SFT, Evaluation Protocol=5-shot2025.11 | 60.1 | |
| TuluDPOBackbone=Apertus-8B-SFT, Evaluation Protocol=5-shot2025.11 | 59.97 | |
| Hybrid-KSAsetting=Train-from-scratch2026.04 | 59.14 | |
| Qwen2-1.5B# Non-Emb Params=1.2B2024.07 | 58.5 | |
| FuseChat3.0#Params=1.48B, Backbone=Llama-3.2-1B-Instruct2026.05 | 57.4 | |
| Phi-2# Non-Emb Params=2.5B2024.07 | 57.2 | |
| ORPOBackbone=Apertus-8B-SFT, Evaluation Protocol=5-shot2025.11 | 57.16 | |
| SkillWeave#Params=1.42B, Backbone=Llama-3.2-1B-Instruct2026.05 | 56.7 | |
| HelpSteerBackbone=Apertus-8B-SFT, Evaluation Protocol=5-shot2025.11 | 55.27 | |
| KSAsetting=Train-from-scratch2026.04 | 54.81 | |
| CodePrefBackbone=Apertus-8B-SFT, Evaluation Protocol=5-shot2025.11 | 54.28 | |
| self-rewarding#Params=1.15B, Backbone=Llama-3.2-1B-Instruct2026.05 | 54.2 | |
| Twin-merging#Params=1.48B, Backbone=Llama-3.2-1B-Instruct2026.05 | 53.4 | |
| SFTBackbone=Apertus-8B-SFT, Evaluation Protocol=5-shot2025.11 | 53.07 | |
| Hybrid-SCAsetting=Train-from-scratch2026.04 | 52.39 | |
| Hybrid-GDNsetting=Train-from-scratch2026.04 | 50.95 | |
| UM-190kshots=5-shot2025.11 | 50.14 | |
| UM-187kshots=5-shot2025.11 | 49.45 | |
| ORPOshots=5-shot2025.11 | 48.98 | |
| TuluDPOshots=5-shot2025.11 | 48.84 | |
| Fullsetting=Train-from-scratch2026.04 | 48.29 | |
| HelpSteershots=5-shot2025.11 | 48.07 | |
| UltraFBshots=5-shot2025.11 | 47.99 | |
| UM-170kshots=5-shot2025.11 | 47.96 | |
| Hybrid-SWAsetting=Train-from-scratch2026.04 | 47.46 | |
| SFTshots=5-shot2025.11 | 46.84 | |
| PEFT#Params=1.42B, Backbone=Llama-3.2-1B-Instruct2026.05 | 46.1 | |
| Llama3.2-1B-Instruct#Params=1.15B, Backbone=Llama-3.2-1B-Instruct2026.05 | 45.8 | |
| CodePrefshots=5-shot2025.11 | 44.98 | |
| Qwen2-0.5Bnumber of parameters=0.5B2024.07 | 40.1 | |
| Qwen1.5-1.8B# Non-Emb Params=1.2B2024.07 | 38.4 | |
| Qwen2-0.5B# Non-Emb Params=0.3B2024.07 | 36.5 | |
| Qwen1.5-1.8Bnumber of parameters=1.8B2024.07 | 35.3 | |
| Gemma-2B# Non-Emb Params=2.0B2024.07 | 17.7 | |
| Qwen1.5-0.5Bnumber of parameters=0.5B2024.07 | 11.3 |