Multi-task Language Understanding on MMLU (Accuracy)
94.7AccuracyClaude-Sonnet-4.5
Evaluation Results
| Method | Links | |
|---|---|---|
| Claude-Sonnet-4.5Category=Candidate Models2026.03 | 94.7 | |
| Qwen3-235B-A22BCategory=Candidate Models2026.03 | 93.9 | |
| FineRouterCategory=Routers2026.03 | 93.4 | |
| Llama-4-MaverickCategory=Candidate Models2026.03 | 93 | |
| IPRCategory=Routers2026.03 | 92.2 | |
| DeepSeek-R1Category=Candidate Models2026.03 | 92 | |
| kNNCategory=Routers2026.03 | 91.9 | |
| DS-V3.1-Terminus (no_think)Model=DS-V3.1-Terminus (no_think), Number of Parameters=671B, Privacy Protocol=Plaintext2026.03 | 91.72 | |
| AloePriModel=DS-V3.1-Terminus (no_think), Number of Parameters=671B, Privacy Protocol=AloePri2026.03 | 90.67 | |
| Self-Consistency2026.04 | 90.63 | |
| Weighted MV2026.04 | 90.47 | |
| Single Agent2026.04 | 90.11 | |
| GraphRouterCategory=Routers2026.03 | 90.1 | |
| Simple MV2026.04 | 90.09 | |
| EMS-SimSorting Strategy=Semantic Similarity2026.04 | 90.09 | |
| EMS-RelSorting Strategy=Historical Reliability2026.04 | 90.09 | |
| Conf-MV2026.04 | 89.86 | |
| Qwen3Model=Qwen3, Number of Parameters=32B, Privacy Protocol=Plaintext2026.03 | 89.61 | |
| GPT-OSS-120BCategory=Candidate Models2026.03 | 89.6 | |
| DeepSeek-v3Category=Candidate Models2026.03 | 89.2 | |
| RouteLLMCategory=Routers2026.03 | 89.2 | |
| RouterDCCategory=Routers2026.03 | 89 | |
| Qwen3Model=Qwen3, Number of Parameters=14B, Privacy Protocol=Plaintext2026.03 | 87.64 | |
| Llama-3.3-70BCategory=Candidate Models2026.03 | 87.4 | |
| AloePriModel=Qwen3, Number of Parameters=32B, Privacy Protocol=AloePri2026.03 | 86.24 | |
| OriginalModel=GPT-oss, Scope=–2026.04 | 85.6 | |
| Qwen3-32BCategory=Candidate Models2026.03 | 85.3 | |
| GRAPEModel=GPT-oss, Scope=Global, Number of experts pruned=2e2026.04 | 85.3 | |
| Claude-Haiku-4.5Category=Candidate Models2026.03 | 85.2 | |
| Count-guidedModel=GPT-oss, Scope=Local, Number of experts pruned=2e2026.04 | 84.8 | |
| Qwen3-MoE-InstructModel=Qwen3-MoE-Instruct, Number of Parameters=30B-A3B, Privacy Protocol=Plaintext2026.03 | 84.75 | |
| AloePriModel=Qwen3, Number of Parameters=14B, Privacy Protocol=AloePri2026.03 | 84.57 | |
| AloePriModel=Qwen3-MoE-Instruct, Number of Parameters=30B-A3B, Privacy Protocol=AloePri2026.03 | 84.49 | |
| EnumerateModel=GPT-oss, Scope=Local, Number of experts pruned=2e2026.04 | 83.7 | |
| AloePriModel=R1-Distill, Number of Parameters=32B, Privacy Protocol=AloePri2026.03 | 83.61 | |
| DEKModel=GPT-oss, Scope=Local, Number of experts pruned=2e2026.04 | 83.6 | |
| R1-DistillModel=R1-Distill, Number of Parameters=32B, Privacy Protocol=Plaintext2026.03 | 83.5 | |
| Linear-AcTBackbone=Qwen2.5-32B, Shots=5-shot2026.04 | 83.42 | |
| Llama 3.1 InstructModel Scale=70B2025.04 | 83.4 | |
| Router-guidedModel=GPT-oss, Scope=Local, Number of experts pruned=2e2026.04 | 83.4 | |
| GRAPEModel=GPT-oss, Scope=Global, Number of experts pruned=4e2026.04 | 83.4 | |
| FP16Model=Qwen2.5-32B, Avg. Bits=16.00, Evaluation Protocol=5-shot2025.08 | 83.32 | |
| ITIBackbone=Qwen2.5-32B, Shots=5-shot2026.04 | 83.02 | |
| OriginalBackbone=Qwen2.5-32B, Shots=5-shot2026.04 | 82.83 | |
| Count-guidedModel=GPT-oss, Scope=Local, Number of experts pruned=4e2026.04 | 82.5 | |
| EnumerateModel=GPT-oss, Scope=Local, Number of experts pruned=4e2026.04 | 82.5 | |
| Llama 3 InstructModel Scale=70B2025.04 | 82 | |
| MicroMixModel=Qwen2.5-32B, Avg. Bits=5.22, Evaluation Protocol=5-shot2025.08 | 81.79 | |
| ParamΔModel Scale=70B2025.04 | 81.7 | |
| FlatQuantModel=Qwen2.5-32B, Avg. Bits=4.71, Evaluation Protocol=5-shot2025.08 | 81.52 | |
| Router-guidedModel=GPT-oss, Scope=Local, Number of experts pruned=4e2026.04 | 81.3 | |
| BaseZero-shot=true2026.03 | 81 | |
| DEKModel=GPT-oss, Scope=Local, Number of experts pruned=4e2026.04 | 80.8 | |
| Mean-AcTBackbone=Qwen-2.5-14B, Shots=5-shot2026.04 | 80.28 | |
| PID-AcTBackbone=Qwen-2.5-14B, Shots=5-shot2026.04 | 80.16 | |
| AMXFP4Model=Qwen2.5-32B, Avg. Bits=5.00, Evaluation Protocol=5-shot2025.08 | 79.96 | |
| AtomModel=Qwen2.5-32B, Avg. Bits=4.21, Evaluation Protocol=5-shot2025.08 | 79.54 | |
| QuaRotModel=Qwen2.5-32B, Avg. Bits=4.12, Evaluation Protocol=5-shot2025.08 | 79.39 | |
| INT6Model=Qwen2.5-32B, Avg. Bits=6.00, Evaluation Protocol=5-shot2025.08 | 79.33 | |
| S-PIDBackbone=Qwen2.5-32B, Shots=5-shot2026.04 | 79.24 | |
| AloePriModel=R1-Distill, Number of Parameters=14B, Privacy Protocol=AloePri2026.03 | 79.2 | |
| ITIBackbone=Qwen-2.5-14B, Shots=5-shot2026.04 | 79.04 | |
| R1-DistillModel=R1-Distill, Number of Parameters=14B, Privacy Protocol=Plaintext2026.03 | 78.91 | |
| QUIKModel=Qwen2.5-32B, Avg. Bits=6.21, Evaluation Protocol=5-shot2025.08 | 78.89 | |
| Llama 3 BaseModel Scale=70B2025.04 | 78.8 | |
| OriginalBackbone=Qwen-2.5-14B, Shots=5-shot2026.04 | 78.8 | |
| Linear-AcTBackbone=Qwen-2.5-14B, Shots=5-shot2026.04 | 78.64 | |
| Llama 3.1 BaseModel Scale=70B2025.04 | 78.5 | |
| ODESteerBackbone=Qwen-2.5-14B, Shots=5-shot2026.04 | 78.08 | |
| SFTZero-shot=true2026.03 | 76 | |
| S-PIDBackbone=Qwen-2.5-14B, Shots=5-shot2026.04 | 75.9 | |
| A-LQRBackbone=Qwen2.5-32B, Shots=5-shot2026.04 | 75.4 | |
| GRPOZero-shot=true2026.03 | 75 | |
| MLPCategory=Routers2026.03 | 74.7 | |
| SA-SFTBackbone=Qwen2.5-7B-I, Tuning Strategy=full, Prompting Setting=5-shot2026.01 | 74.4 | |
| HelixTemperature=1.02026.02 | 74.3 | |
| Mistral-SmallCategory=Candidate Models2026.03 | 74.1 | |
| HelixTemperature=0.52026.02 | 73.73 | |
| BF16Rate=16.000, Evaluation Protocol=Zero-shot, Backbone Model=Qwen3-8B2026.03 | 73.02 | |
| HelixTemperature=1.52026.02 | 72.92 | |
| HelixTemperature=2.02026.02 | 72.82 | |
| WaterSICRate=4.125, Evaluation Protocol=Zero-shot, Backbone Model=Qwen3-8B2026.03 | 72.77 | |
| HelixTemperature=2.52026.02 | 72.52 | |
| HelixTemperature=3.02026.02 | 72.49 | |
| Huffman-GPTQRate=4.125, Evaluation Protocol=Zero-shot, Backbone Model=Qwen3-8B2026.03 | 72.28 | |
| HRTNRate=4.125, Evaluation Protocol=Zero-shot, Backbone Model=Qwen3-8B2026.03 | 72.15 | |
| LCO-KLDBackbone=Qwen-3-4B2026.03 | 72.11 | |
| GPTQRate=4.125, Evaluation Protocol=Zero-shot, Backbone Model=Qwen3-8B2026.03 | 71.76 | |
| A-LQRBackbone=Qwen-2.5-14B, Shots=5-shot2026.04 | 71.66 | |
| Huffman-GPTQRate=3.125, Evaluation Protocol=Zero-shot, Backbone Model=Qwen3-8B2026.03 | 70.96 | |
| WaterSICRate=3.125, Evaluation Protocol=Zero-shot, Backbone Model=Qwen3-8B2026.03 | 70.53 | |
| Mistral-LargeCategory=Candidate Models2026.03 | 70.5 | |
| IPROXTarget Model=Qwen2-7B, Sparsity Level (ρ)=0.7, #Params=3.3B2026.02 | 70.41 | |
| IPROXTarget Model=Qwen2-7B, Sparsity Level (ρ)=0.3, #Params=5.8B2026.02 | 70.36 | |
| Qwen2-7BTarget Model=Qwen2-7B, Proxy Model=Qwen2-7B, #Params=7B2026.02 | 70.35 | |
| IPROXTarget Model=Qwen2-7B, Sparsity Level (ρ)=0.5, #Params=4.4B2026.02 | 70.27 | |
| Qwen2-1.5BTarget Model=Qwen2-7B, Proxy Model=Qwen2-1.5B, #Params=1.5B2026.02 | 70.18 | |
| IPROXTarget Model=Qwen3-4B, Sparsity Level (ρ)=0.3, #Params=3.1B2026.02 | 70.15 | |
| IPROXTarget Model=Qwen3-4B, Sparsity Level (ρ)=0.5, #Params=2.2B2026.02 | 70.08 | |
| IPROXTarget Model=Qwen3-4B, Sparsity Level (ρ)=0.7, #Params=1.5B2026.02 | 69.94 |