Math Problem Solving on AMC23
70.6AccuracyTaH+
Evaluation Results
| Method | Links | |
|---|---|---|
| TaH+Param.=4B2025.11 | 70.6 | |
| TaHParam.=4B2025.11 | 70.3 | |
| TaHModel Size=4B*2025.11 | 69.7 | |
| TaH+Model Size=4B*2025.11 | 68.1 | |
| OuroModel Size=4B*2025.11 | 64.4 | |
| StandardModel Size=4B*2025.11 | 64.2 | |
| SoftThinkParam.=4B2025.11 | 64.1 | |
| SoftThinkModel Size=4B*2025.11 | 63.1 | |
| StandardParam.=4B2025.11 | 62.8 | |
| RoutingParam.=4B2025.11 | 60.9 | |
| Fine-tunedBase LLM=Qwen-2.5-7B-Instruct, Parameters=x4, Evaluation Protocol=0 shot, CoT2026.05 | 57.5 | |
| Zero-shotBase LLM=Qwen-2.5-7B-Instruct, Parameters=x1, Evaluation Protocol=0 shot, CoT2026.05 | 52.5 | |
| DiDi-Merg.-LBase LLM=Qwen-2.5-7B-Instruct, Parameters=x2.0, Evaluation Protocol=0 shot, CoT2026.05 | 52.5 | |
| TaH+Param.=1.7B2025.11 | 51.2 | |
| Twin-MergingBase LLM=Qwen-2.5-7B-Instruct, Parameters=x2.25, Evaluation Protocol=0 shot, CoT2026.05 | 50 | |
| TaHParam.=1.7B2025.11 | 48.4 | |
| TaH+Model Size=1.7B2025.11 | 48.4 | |
| FREE-MergingBase LLM=Qwen-2.5-7B-Instruct, Parameters=x2.08, Evaluation Protocol=0 shot, CoT2026.05 | 47.5 | |
| SoftThinkParam.=1.7B2025.11 | 43.1 | |
| AlwaysThinkParam.=1.7B2025.11 | 42.5 | |
| StandardParam.=1.7B2025.11 | 42.2 | |
| RoutingParam.=1.7B2025.11 | 42.2 | |
| AlwaysThinkModel Size=1.7B2025.11 | 40.9 | |
| TaHModel Size=1.7B2025.11 | 40.9 | |
| OuroModel Size=1.7B2025.11 | 40.6 | |
| SoftThinkModel Size=1.7B2025.11 | 40.3 | |
| StandardModel Size=1.7B2025.11 | 39.7 | |
| Fine-tunedBase LLM=Llama-3.1-8B-Instruct, Parameters=x4, Evaluation Protocol=0 shot, CoT2026.05 | 37.5 | |
| DiDi-Merg.-LBase LLM=Llama-3.1-8B-Instruct, Parameters=x2.0, Evaluation Protocol=0 shot, CoT2026.05 | 35 | |
| TaHParam.=0.6B2025.11 | 32.5 | |
| TaH+Param.=0.6B2025.11 | 30.6 | |
| Twin-MergingBase LLM=Llama-3.1-8B-Instruct, Parameters=x2.25, Evaluation Protocol=0 shot, CoT2026.05 | 27.5 | |
| FREE-MergingBase LLM=Llama-3.1-8B-Instruct, Parameters=x2.08, Evaluation Protocol=0 shot, CoT2026.05 | 27.5 | |
| Zero-shotBase LLM=Llama-3.1-8B-Instruct, Parameters=x1, Evaluation Protocol=0 shot, CoT2026.05 | 25 | |
| TaH+Model Size=0.6B2025.11 | 24.7 | |
| SoftThinkParam.=0.6B2025.11 | 24.1 | |
| TaHModel Size=0.6B2025.11 | 24.1 | |
| StandardParam.=0.6B2025.11 | 23.4 | |
| StandardModel Size=0.6B2025.11 | 22.7 | |
| SoftThinkModel Size=0.6B2025.11 | 22.2 | |
| AlwaysThinkModel Size=0.6B2025.11 | 21.9 | |
| OuroModel Size=0.6B2025.11 | 19.7 | |
| AlwaysThinkParam.=0.6B2025.11 | 15.6 | |
| RoutingParam.=0.6B2025.11 | 10.9 |