Reasoning on GSM8K, Math, AIME, HumanEval, and LiveCodeBench (test)
87.23GSM8K AccuracyReasoning
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| ReasoningModel Type=Domain Expert (DeepSeek-R1-Distill-Qwen-7B)2026.01 | 87.23 | 86.2 | 60 | 89.61 | 30.37 | |
| ReasonAnyMethod Class=Model Merging2026.01 | 73.77 | 69.4 | 36.67 | 70.42 | 11.8 | |
| LEDMethod Class=Model Merging2026.01 | 72.48 | 60.6 | 30 | 65.23 | 17.6 | |
| BiomedicineModel Type=Domain Expert (Meditron3-Qwen2.5-7B)2026.01 | 69.4 | 74 | 6.67 | 37.95 | 3.2 | |
| Task ArithmeticMethod Class=Model Merging2026.01 | 62.17 | 42.8 | 26.67 | 48.25 | 3.8 | |
| LinearMethod Class=Model Merging2026.01 | 50.42 | 43.8 | 16.67 | 37.23 | 3.8 | |
| FuseLLMMethod Class=Model Merging2026.01 | 1.8 | 0.2 | 16.67 | 55.58 | 5 | |
| TIESMethod Class=Model Merging2026.01 | 0.83 | 2.6 | 0 | 40.27 | 5.3 | |
| DAREMethod Class=Model Merging2026.01 | 0.53 | 1 | 0 | 40.16 | 12.5 |