Reasoning on GSM8K, Math, AIME, HumanEval, and LiveCodeBench
85.12GSM8K AccuracyReasoning (DeepSeek-R1-Distill-Llama-8B)
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Reasoning (DeepSeek-R1-Distill-Llama-8B)Model Family=Llama3.1-8B2026.01 | 85.12 | 64.4 | 33.33 | 76.63 | 29.91 | |
| ReasonAnyMerging=ReasonAny, Model Family=Llama3.1-8B2026.01 | 83.3 | 65.8 | 10 | 42.31 | 10.38 | |
| TIESMerging=TIES, Model Family=Llama3.1-8B2026.01 | 79.15 | 63.6 | 3.33 | 21.32 | 7.62 | |
| LinearMerging=Linear, Model Family=Llama3.1-8B2026.01 | 75.51 | 48.6 | 3.33 | 32.27 | 2.97 | |
| DAREMerging=DARE, Model Family=Llama3.1-8B2026.01 | 73.39 | 40.2 | 3.33 | 19.63 | 3.55 | |
| Task ArithmeticMerging=Task Arithmetic, Model Family=Llama3.1-8B2026.01 | 70.66 | 31.2 | 10 | 23.12 | 0.58 | |
| FuseLLMMerging=FuseLLM, Model Family=Llama3.1-8B2026.01 | 58.45 | 20.6 | 0 | 19.12 | 0 | |
| LEDMerging=LED, Model Family=Llama3.1-8B2026.01 | 56.33 | 46.6 | 6.67 | 49.42 | 6.38 | |
| Finance (WiroAI-Finance-Llama-8B)Model Family=Llama3.1-8B2026.01 | 54.36 | 13.6 | 0 | 0 | 0.38 |