Mathematical Problem Solving on MATH (Accuracy)
95.86AccuracyINT4
Evaluation Results
| Method | Links | |
|---|---|---|
| INT4Model=GLM-4.7 (358B)2026.04 | 95.86 | |
| BF16Model=GLM-4.7 (358B)2026.04 | 95.66 | |
| BDR-64Model=GLM-4.7 (358B)2026.04 | 95.59 | |
| BDR-16Model=GLM-4.7 (358B)2026.04 | 95.46 | |
| BDR-128Model=GLM-4.7 (358B)2026.04 | 95.32 | |
| Llama-4-MaverickCategory=Candidate Models2026.03 | 95.2 | |
| Llama-3.3-70BCategory=Candidate Models2026.03 | 94 | |
| FineRouterCategory=Routers2026.03 | 93.8 | |
| IPRCategory=Routers2026.03 | 93.6 | |
| Qwen3-235B-A22BCategory=Candidate Models2026.03 | 93.3 | |
| BDR-64 (K)Model=Qwen3-8B2026.04 | 92.71 | |
| BDR-128Model=Qwen3-8B2026.04 | 92.63 | |
| BF16Model=Qwen3-8B2026.04 | 92.59 | |
| kNNCategory=Routers2026.03 | 92.2 | |
| BDR-64Model=Qwen3-8B2026.04 | 92.18 | |
| BDR-16Model=Qwen3-8B2026.04 | 92.06 | |
| QuratingBase Model=OpenR1-Distill-7B2026.06 | 91.88 | |
| LESSBase Model=OpenR1-Distill-7B2026.06 | 91.76 | |
| DeepSeek-R1Category=Candidate Models2026.03 | 91.6 | |
| IF (Impl. w. GraSS)Base Model=OpenR1-Distill-7B2026.06 | 91.52 | |
| BM25Base Model=OpenR1-Distill-7B2026.06 | 91.2 | |
| Claude-Sonnet-4.5Category=Candidate Models2026.03 | 91.1 | |
| Random BaselineBase Model=OpenR1-Distill-7B2026.06 | 91.05 | |
| RDSBase Model=OpenR1-Distill-7B2026.06 | 91 | |
| OpenR1-Distill-7BBase Model=OpenR1-Distill-7B2026.06 | 90.98 | |
| DSIRBase Model=OpenR1-Distill-7B2026.06 | 90.85 | |
| DRIFTBase Model=OpenR1-Distill-7B2026.06 | 90.82 | |
| GraphRouterCategory=Routers2026.03 | 89.6 | |
| DeepSeek-v3Category=Candidate Models2026.03 | 89 | |
| GPT-OSS-120BCategory=Candidate Models2026.03 | 88.4 | |
| OriginalVocabulary size=32,0002025.08 | 88.4 | |
| VocabTailorVocabulary size=5,135 + [14], % Original Vocabulary=(16.09%)2025.08 | 88.4 | |
| Claude-Haiku-4.5Category=Candidate Models2026.03 | 88.1 | |
| RouterDCCategory=Routers2026.03 | 88 | |
| VPVocabulary size=10,300, % Original Vocabulary=(32.19%)2025.08 | 87.8 | |
| RouteLLMCategory=Routers2026.03 | 87.5 | |
| Self-DistillationBase Model=OpenR1-Distill-7B2026.06 | 83.78 | |
| ToolTreeBackbone Model=GPT-4o2026.03 | 78.19 | |
| Qwen3-32BCategory=Candidate Models2026.03 | 75.5 | |
| TokenBuncherModel=Qwen2.5-7B-Instruct2025.08 | 72.8 | |
| Mistral-LargeCategory=Candidate Models2026.03 | 71 | |
| BaseModel=Qwen2.5-7B-Instruct2025.08 | 70.4 | |
| Qwen-2.5-7B-InstructDecoding Method=Medusa-style tree-based greedy decoding, Decoding Head=multi-token head2025.11 | 69.9 | |
| ToolTreeBackbone Model=GPT-4o-mini2026.03 | 69.42 | |
| OctoToolsBackbone Model=GPT-4o2026.03 | 68.57 | |
| BM25Base Model=Olmo3-7B-Instruct-SFT2026.06 | 67.86 | |
| DRIFTBase Model=Olmo3-7B-Instruct-SFT2026.06 | 67.52 | |
| Self-DistillationBase Model=Olmo3-7B-Instruct-SFT2026.06 | 67.48 | |
| QuratingBase Model=Olmo3-7B-Instruct-SFT2026.06 | 67.29 | |
| RepNoiseModel=Qwen2.5-7B-Instruct2025.08 | 66.8 | |
| Olmo3-7B-Instruct-SFTBase Model=Olmo3-7B-Instruct-SFT2026.06 | 66.16 | |
| RDSBase Model=Olmo3-7B-Instruct-SFT2026.06 | 66.16 | |
| Random BaselineBase Model=Olmo3-7B-Instruct-SFT2026.06 | 66.09 | |
| IF (Impl. w. GraSS)Base Model=Olmo3-7B-Instruct-SFT2026.06 | 65.42 | |
| BaseModel=Qwen2.5-3B-Instruct2025.08 | 65.4 | |
| LESSBase Model=Olmo3-7B-Instruct-SFT2026.06 | 64.22 | |
| TokenBuncherModel=Qwen2.5-3B-Instruct2025.08 | 62.8 | |
| Few-ShotBackbone Model=GPT-4o2026.03 | 61.45 | |
| MLPCategory=Routers2026.03 | 59.5 | |
| OctoToolsBackbone Model=GPT-4o-mini2026.03 | 58.43 | |
| DSIRBase Model=Olmo3-7B-Instruct-SFT2026.06 | 54.63 | |
| HuggingGPTBackbone Model=GPT-4o2026.03 | 53.51 | |
| Few-ShotBackbone Model=GPT-4o-mini2026.03 | 53.26 | |
| TokenBuncherModel=Ministral-8B-Instruct2025.08 | 53 | |
| BaseModel=Ministral-8B-Instruct2025.08 | 52 | |
| RepNoiseModel=Qwen2.5-3B-Instruct2025.08 | 51.8 | |
| BaselineN=16, setting=English occurrence as language confusion, aggregation=Averaged across four models and four target languages2026.04 | 49.59 | |
| OriginalModel Family=LLaMA2026.05 | 48.34 | |
| TLPON=16, setting=English occurrence as language confusion, aggregation=Averaged across four models and four target languages2026.04 | 47.73 | |
| HuggingGPTBackbone Model=GPT-4o-mini2026.03 | 45.14 | |
| CTRAPModel=Qwen2.5-7B-Instruct2025.08 | 45 | |
| CTRAPModel=Qwen2.5-3B-Instruct2025.08 | 43.8 | |
| ORPON=16, setting=English occurrence as language confusion, aggregation=Averaged across four models and four target languages2026.04 | 43.78 | |
| Tulu 3 8BDecoding Method=Medusa-style tree-based greedy decoding, Decoding Head=multi-token head2025.11 | 43.7 | |
| Llama-3.1-8B-InstructDecoding Method=Medusa-style tree-based greedy decoding, Decoding Head=multi-token head2025.11 | 42.5 | |
| DPON=16, setting=English occurrence as language confusion, aggregation=Averaged across four models and four target languages2026.04 | 42 | |
| SOMAModel Family=LLaMA2026.05 | 41.62 | |
| SFTN=16, setting=English occurrence as language confusion, aggregation=Averaged across four models and four target languages2026.04 | 41.35 | |
| Claude 3.5 Haiku (target)Model=Claude 3.5 Haiku2026.04 | 41.3 | |
| Ministral-8B-InstructDecoding Method=Medusa-style tree-based greedy decoding, Decoding Head=multi-token head2025.11 | 40 | |
| RepNoiseModel=Ministral-8B-Instruct2025.08 | 39.8 | |
| OriginalModel Family=Qwen2026.05 | 36.48 | |
| CTRAPModel=Ministral-8B-Instruct2025.08 | 34.4 | |
| RouteLLMModel Family=LLaMA2026.05 | 33.88 | |
| MAP-Neo-7B-InstructDecoding Method=Medusa-style tree-based greedy decoding, Decoding Head=multi-token head2025.11 | 31.5 | |
| Llama 3.1 8B (baseline)Model=Llama 3.1 8B2026.04 | 31.46 | |
| History-FTModel Family=LLaMA2026.05 | 31.46 | |
| OLMo-2-7B-1124-InstructDecoding Method=Medusa-style tree-based greedy decoding, Decoding Head=multi-token head2025.11 | 31.3 | |
| SOMAModel Family=Qwen2026.05 | 31.14 | |
| Gemma-2-9B-itDecoding Method=Medusa-style tree-based greedy decoding, Decoding Head=multi-token head2025.11 | 29.8 | |
| EvaByte-SFTDecoding Method=Medusa-style tree-based greedy decoding, Decoding Head=multi-token head2025.11 | 29.8 | |
| GARFA (Exp A)Optimization Strategy=L*+L*2026.04 | 28.95 | |
| Exp C: Lr+L*Optimization Strategy=Lr+L*2026.04 | 28.12 | |
| Mistral-SmallCategory=Candidate Models2026.03 | 25.1 | |
| OLMo-2-7B-SFTDecoding Method=Medusa-style tree-based greedy decoding, Decoding Head=multi-token head2025.11 | 25.1 | |
| RouteLLMModel Family=Qwen2026.05 | 25.08 | |
| Exp B: L*+LrOptimization Strategy=L*+Lr2026.04 | 25.04 | |
| History-PrefixModel Family=LLaMA2026.05 | 25.03 | |
| Exp D: Lr+Lr (control)Optimization Strategy=Lr+Lr2026.04 | 22.87 | |
| History-FTModel Family=Qwen2026.05 | 22.57 |