Multiple Choice Question Answering on ARC Easy
99.7AccuracyBERT-Judge
Evaluation Results
| Method | Links | |
|---|---|---|
| BERT-Judgeclean_name=BERT-as-a-Judge2026.04 | 99.7 | |
| Regex2026.04 | 88.2 | |
| DPOBackbone=Qwen3-8B2026.03 | 85.5 | |
| RepEBackbone=Qwen3-8B2026.03 | 84.4 | |
| Prompt EngBackbone=Qwen3-8B2026.03 | 84.4 | |
| Base ModelBackbone=Qwen3-8B2026.03 | 84.4 | |
| PAROQType=linear, Backbone=Qwen3-14B2025.11 | 84.3 | |
| DSPABackbone=Qwen3-8B2026.03 | 84.3 | |
| FP16Type=-, Backbone=Qwen3-14B2025.11 | 84.2 | |
| Static-SAEBackbone=Qwen3-8B2026.03 | 84.1 | |
| E-QATType=linear, Backbone=Qwen3-14B2025.11 | 84 | |
| DSPABackbone=Gemma-2-9B2026.03 | 83.8 | |
| RepEBackbone=Gemma-2-9B2026.03 | 83.8 | |
| Prompt EngBackbone=Gemma-2-9B2026.03 | 83.7 | |
| Base ModelBackbone=Gemma-2-9B2026.03 | 83.7 | |
| DPOBackbone=Gemma-2-9B2026.03 | 83.6 | |
| Static-SAEBackbone=Gemma-2-9B2026.03 | 83.6 | |
| FP16Type=-, Backbone=Qwen3-8B2025.11 | 83.5 | |
| QTIPType=vector, Backbone=Qwen3-14B2025.11 | 83.5 | |
| PAROQType=linear, Backbone=Qwen3-8B2025.11 | 83.3 | |
| AWQType=linear, Backbone=Qwen3-14B2025.11 | 83.2 | |
| QTIPType=vector, Backbone=Qwen3-8B2025.11 | 82.8 | |
| BF16Model Size=8B2026.03 | 82.71 | |
| PAROQType=linear, Backbone=LLaMA-3.1-8B-Instruct2025.11 | 82.2 | |
| AWQType=linear, Backbone=Qwen3-8B2025.11 | 82.2 | |
| FP16Type=-, Backbone=LLaMA-3.1-8B-Instruct2025.11 | 81.8 | |
| E-QATType=linear, Backbone=Qwen3-8B2025.11 | 81.7 | |
| QTIPType=vector, Backbone=LLaMA-3.1-8B-Instruct2025.11 | 81.6 | |
| DPOBackbone=Gemma-2-2B2026.03 | 81 | |
| BF16Model Size=8B, Quantization Method=None (BF16), Adaptive Block Scaling (4/6)=false2025.12 | 80.9 | |
| E-QATType=linear, Backbone=LLaMA-3.1-8B-Instruct2025.11 | 80.9 | |
| GPTQ + 4/6Model Size=8B, Quantization Method=GPTQ, Adaptive Block Scaling (4/6)=true2025.12 | 80.8 | |
| Prompt EngBackbone=Gemma-2-2B2026.03 | 80.8 | |
| PAROQType=linear, Backbone=Qwen3-4B2025.11 | 80.7 | |
| AWQType=linear, Backbone=LLaMA-3.1-8B-Instruct2025.11 | 80.6 | |
| Static-SAEBackbone=Gemma-2-2B2026.03 | 80.6 | |
| SpQR†Bits=4.12 (W3)2026.07 | 80.6 | |
| FP16Type=-, Backbone=Qwen3-4B2025.11 | 80.5 | |
| DSPABackbone=Gemma-2-2B2026.03 | 80.4 | |
| Base ModelBackbone=Gemma-2-2B2026.03 | 80.4 | |
| FP16Bits=16.02026.07 | 80.4 | |
| RepEBackbone=Gemma-2-2B2026.03 | 80.3 | |
| HQQBits=4.02026.07 | 80.2 | |
| GPTQModel Size=8B, Quantization Method=GPTQ, Adaptive Block Scaling (4/6)=false2025.12 | 79.8 | |
| QTIPType=vector, Backbone=Qwen3-4B2025.11 | 79.8 | |
| OWQBits=3.02026.07 | 79.8 | |
| E-QATType=linear, Backbone=Qwen3-4B2025.11 | 79.7 | |
| OriginalBackbone=Mistral-7B, Pruning Ratio=0%2026.05 | 79.55 | |
| REAPModel=Qwen3-30B-A3B, Reduction ratio=25%, Compression type=Pruning, Evaluation protocol=One-shot2026.05 | 79.5 | |
| AWQBits=4.02026.07 | 79.5 | |
| RTNBits=4.02026.07 | 79.3 | |
| TASA b4.0Bits=4.02026.07 | 79.2 | |
| OriginalModel=Qwen3-30B-A3B, Reduction ratio=0%, Compression type=None, Evaluation protocol=One-shot2026.05 | 79.1 | |
| FrequencyModel=Qwen3-30B-A3B, Reduction ratio=25%, Compression type=Pruning, Evaluation protocol=One-shot2026.05 | 79.1 | |
| FAAR+2FAModel Size=8B2026.03 | 79.02 | |
| M-SMoEModel=Qwen3-30B-A3B, Reduction ratio=25%, Compression type=Merging, Evaluation protocol=One-shot2026.05 | 78.8 | |
| GPTQ+4/6Model Size=8B2026.03 | 78.72 | |
| RTNModel Size=8B, Quantization Method=RTN, Adaptive Block Scaling (4/6)=false2025.12 | 78.6 | |
| ConMoEModel=Qwen3-30B-A3B, Reduction ratio=25%, Compression type=None, Evaluation protocol=One-shot2026.05 | 78.6 | |
| MR-GPTQModel Size=8B2026.03 | 78.55 | |
| RTN + 4/6Model Size=8B, Quantization Method=RTN, Adaptive Block Scaling (4/6)=true2025.12 | 78.5 | |
| AWQModel Size=8B, Quantization Method=AWQ, Adaptive Block Scaling (4/6)=false2025.12 | 78.4 | |
| SmoothQuant + 4/6Model Size=8B, Quantization Method=SmoothQuant, Adaptive Block Scaling (4/6)=true2025.12 | 78.4 | |
| GPTQModel Size=8B2026.03 | 78.3 | |
| AWQ + 4/6Model Size=8B, Quantization Method=AWQ, Adaptive Block Scaling (4/6)=true2025.12 | 78.1 | |
| SmoothQuantModel Size=8B, Quantization Method=SmoothQuant, Adaptive Block Scaling (4/6)=false2025.12 | 78.1 | |
| GPTQBits=4.02026.07 | 78.1 | |
| AWQType=linear, Backbone=Qwen3-4B2025.11 | 77.9 | |
| M-SMoEModel=Qwen3-30B-A3B, Reduction ratio=50%, Compression type=Merging, Evaluation protocol=One-shot2026.05 | 77.7 | |
| MoEITSBackbone=DeepSeek-V2 Lite, Evaluation setting=15-shot, Average Accuracy (Av)=67.99, Parameter Reduction (P)=55.42%2026.04 | 77.61 | |
| RTNModel Size=8B2026.03 | 77.51 | |
| FrequencyModel=Qwen3-30B-A3B, Reduction ratio=50%, Compression type=Pruning, Evaluation protocol=One-shot2026.05 | 77.3 | |
| OriginalModel=OLMoE-1B-7B-0125, Reduction ratio=0%, Compression type=None, Evaluation protocol=One-shot2026.05 | 77 | |
| BaselineBackbone=Qwen1.5-MoE, Weight Bitwidth=162026.04 | 76.77 | |
| UnifiedQA-3bparameters=3B, Protocol=Transfer-learning, External Knowledge=Yes2023.07 | 76.5 | |
| REAPModel=Qwen3-30B-A3B, Reduction ratio=50%, Compression type=Pruning, Evaluation protocol=One-shot2026.05 | 76.5 | |
| TASA b3.0Bits=3.02026.07 | 76.1 | |
| ConMoEModel=OLMoE-1B-7B-0125, Reduction ratio=25%, Compression type=None, Evaluation protocol=One-shot2026.05 | 75.6 | |
| REAPModel=OLMoE-1B-7B-0125, Reduction ratio=25%, Compression type=Pruning, Evaluation protocol=One-shot2026.05 | 75.3 | |
| HQQBits=3.02026.07 | 75 | |
| AccBackbone=Mistral-7B, Pruning Ratio=25%, Relevance Estimation Source=Task-specific2026.05 | 74.83 | |
| FKL*Vocab size=64k, Number of shots=4-shot, Combined with SFT=true2025.12 | 74 | |
| UnifiedQA-3bparameters=3B, Protocol=Transfer-learning, External Knowledge=No2023.07 | 73.7 | |
| DSKD*Vocab size=64k, Number of shots=4-shot, Combined with SFT=true2025.12 | 73.6 | |
| Memory Grafting# Shots=0-shot, # Trainable Params=2.8B, # Activated (w/o token embed)=0.55B, # Trained Tokens=100B, # Experts (shared + routed, top-k)=1 + 47 (top-4)2026.05 | 73.4 | |
| FKLVocab size=64k, Number of shots=4-shot2025.12 | 73.2 | |
| DSKDVocab size=64k, Number of shots=4-shot2025.12 | 73.1 | |
| ConMoEModel=Qwen3-30B-A3B, Reduction ratio=50%, Compression type=None, Evaluation protocol=One-shot2026.05 | 73.1 | |
| SFTVocab size=64k, Number of shots=4-shot2025.12 | 72.9 | |
| HC-SMoEModel=OLMoE-1B-7B-0125, Reduction ratio=25%, Compression type=Merging, Evaluation protocol=One-shot2026.05 | 72.9 | |
| HC-SMoEModel=Qwen3-30B-A3B, Reduction ratio=25%, Compression type=Merging, Evaluation protocol=One-shot2026.05 | 72.8 | |
| MxMoEBackbone=Mixtral, Weight Bitwidth=2.252026.04 | 72.77 | |
| NoWagBackbone=Mixtral, Weight Bitwidth=2.082026.04 | 72.73 | |
| ALM*Vocab size=64k, Number of shots=4-shot, Combined with SFT=true2025.12 | 72.5 | |
| ALMVocab size=64k, Number of shots=4-shot2025.12 | 72.4 | |
| ULD*Vocab size=64k, Number of shots=4-shot, Combined with SFT=true2025.12 | 72.4 | |
| MoE Baseline# Shots=0-shot, # Trainable Params=2.8B, # Activated (w/o token embed)=0.55B, # Trained Tokens=100B, # Experts (shared + routed, top-k)=1 + 64 (top-4)2026.05 | 72.39 | |
| ULDVocab size=64k, Number of shots=4-shot2025.12 | 72.2 | |
| MoE-PrunerBackbone=DeepSeek-V2 Lite, Evaluation setting=15-shot, Average Accuracy (Av)=61.422026.04 | 71.89 | |
| MoE-I^2Backbone=DeepSeek-V2 Lite, Evaluation setting=15-shot, Average Accuracy (Av)=62.79, Parameter Reduction (P)=53.98%2026.04 | 71.8 |