Common-sense Reasoning on COPA
99.2AccuracyAUSteer
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| AUSteerModel=Llama-3.3-70B-Instruct2026.02 | 99.2 | — | |
| VanillaModel=Llama-3.3-70B-Instruct2026.02 | 98.6 | — | |
| AUSteerModel=Qwen3-30B-A3B2026.02 | 97.8 | — | |
| PaLM 2-Lprompting=1-shot2023.05 | 96 | — | |
| Traditional MTL2024.05 | 94.1 | — | |
| MoNEPruning ratio=25%, Backbone=Qwen2-57B-A14B, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 94 | — | |
| VanillaModel=Qwen3-30B-A3B2026.02 | 93.4 | — | |
| Qwen2-57B-A14BPruning ratio=0%, Backbone=Qwen2-57B-A14B, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 93 | — | |
| Qwen2-57B-A14BPruning ratio=0%, Base Model=Qwen2-57B-A14B, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 93 | — | |
| HIZOOModel=LLaMA-3.1-8B2025.10 | 93 | 1.5 | |
| MoonlightPruning ratio=0%, Backbone=Moonlight, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 92 | — | |
| MoonlightPruning ratio=0%, Base Model=Moonlight, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 92 | — | |
| MeZOModel=LLaMA-3.1-8B2025.10 | 92 | 1.54 | |
| ZO Fine-tunerModel=Qwen2.5-14B2025.10 | 92 | 1.34 | |
| T0Model Size=11B, Evaluation Protocol=Zero-shot2022.10 | 91.5 | — | |
| GPT-3Model Size=175B, Evaluation Protocol=Zero-shot2022.10 | 91 | — | |
| PaLMprompting=1-shot2023.05 | 91 | — | |
| FLAPPruning ratio=25%, Backbone=Qwen2-57B-A14B, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 91 | — | |
| ZO Fine-tunerModel=LLaMA-3.1-8B2025.10 | 91 | 1.35 | |
| LOZOModel=Qwen2.5-14B2025.10 | 91 | 1.4 | |
| FLIPPEDModel Size=11B, Evaluation Protocol=Zero-shot2022.10 | 90.75 | — | |
| PaLM 2-Mprompting=1-shot2023.05 | 90 | — | |
| FLAPPruning ratio=25%, Backbone=Moonlight, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 90 | — | |
| RSPruning ratio=25%, Backbone=Moonlight, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 90 | — | |
| MoNEPruning ratio=25%, Backbone=Moonlight, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 90 | — | |
| RSPruning ratio=25%, Backbone=Deepseek-V2-Lite, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 90 | — | |
| MoNEPruning ratio=25%, Backbone=Deepseek-V2-Lite, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 90 | — | |
| MoNEPruning ratio=50%, Base Model=Moonlight, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 90 | — | |
| FLIPPEDModel Size=3B, Evaluation Protocol=Zero-shot2022.10 | 89.88 | — | |
| DirectModel Size=3B, Evaluation Protocol=Zero-shot2022.10 | 89.63 | — | |
| Always Tell Me The OddsBackbone=Qwen2.5-14B-Instruct, Data Augmentation=+Syn, Training Strategy=+R2025.05 | 89.3 | — | |
| PaLM 2-Sprompting=1-shot2023.05 | 89 | — | |
| MC-SMoEPruning ratio=25%, Backbone=Moonlight, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 89 | — | |
| RSPruning ratio=50%, Base Model=Moonlight, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 89 | — | |
| RSPruning ratio=50%, Base Model=Deepseek-V2-Lite, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 89 | — | |
| MeZO-AdamUModel=LLaMA-3.1-8B2025.10 | 89 | 1.67 | |
| LOZOModel=LLaMA-3.1-8B2025.10 | 89 | 1.46 | |
| Deepseek-V2-LitePruning ratio=0%, Backbone=Deepseek-V2-Lite, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 88 | — | |
| MC-SMoEPruning ratio=25%, Backbone=Qwen2-57B-A14B, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 88 | — | |
| Deepseek-V2-LitePruning ratio=0%, Base Model=Deepseek-V2-Lite, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 88 | — | |
| FO-PromptModel=Llama-2-7b2026.04 | 88 | — | |
| Hybrid-PromptModel=Llama-2-7b2026.04 | 88 | — | |
| Hybrid-LoRAModel=Llama-2-7b2026.04 | 88 | — | |
| Always Tell Me The OddsBackbone=Qwen2.5-14B-Instruct, Data Augmentation=+Syn2025.05 | 87.9 | — | |
| DeepSeek-R1-Distill-Qwen-32BEvaluation Protocol=0-Shot2025.05 | 87.7 | — | |
| Task ArithmeticValidation=true2024.05 | 87.5 | — | |
| Always Tell Me The OddsBackbone=Qwen2.5-7B-Instruct2025.05 | 87.2 | — | |
| RSPruning ratio=25%, Backbone=Qwen2-57B-A14B, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 87 | — | |
| MC-SMoEPruning ratio=50%, Base Model=Qwen2-57B-A14B, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 87 | — | |
| HIZOOModel=Qwen2.5-14B2025.10 | 87 | 1.34 | |
| Llama-3-InstructEvaluation Protocol=Probe, Size=14B2025.05 | 86.8 | — | |
| Always Tell Me The OddsBackbone=Qwen2.5-14B-Instruct2025.05 | 86.5 | — | |
| GENICLBackbone=LLaMA-7B, Evaluation Setting=Few-shot (ICL)2025.05 | 86 | — | |
| MeZOModel=Qwen2.5-14B2025.10 | 86 | 1.28 | |
| MNPPSupervision Type=Unsupervised, Zero-shot Transfer=True2021.05 | 85.5 | — | |
| Individual2024.05 | 85.3 | — | |
| Always Tell Me The OddsBackbone=Qwen2.5-8B-Instruct2025.05 | 85.1 | — | |
| FullBackbone=OPT-30B2023.06 | 85 | — | |
| H2ONumber of shots=5-shot, KV cache budget=20%, Sparsification Method=Heavy Hitter Oracle2023.06 | 85 | — | |
| LLM-RBackbone=LLaMA-7B, Evaluation Setting=Few-shot (ICL)2025.05 | 85 | — | |
| OLMoEPruning ratio=0%, Backbone=OLMoE, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 85 | — | |
| MoNEPruning ratio=25%, Backbone=OLMoE, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 85 | — | |
| FLAPPruning ratio=25%, Backbone=Qwen3-30B-A3B, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 85 | — | |
| RSPruning ratio=25%, Backbone=Qwen3-30B-A3B, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 85 | — | |
| MoNEPruning ratio=25%, Backbone=Qwen3-30B-A3B, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 85 | — | |
| OLMoEPruning ratio=0%, Base Model=OLMoE, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 85 | — | |
| MoNEPruning ratio=50%, Base Model=Qwen2-57B-A14B, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 85 | — | |
| MoNEPruning ratio=50%, Base Model=Qwen3-30B-A3B, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 85 | — | |
| Hybrid-PrefixModel=Llama-2-7b2026.04 | 85 | — | |
| FO-LoRAModel=Vicuna-7b-v1.52026.04 | 85 | — | |
| Gated KalmaNetevaluation_mode=zero-shot, parameters=2.8B, context_length=< 2K tokens2025.11 | 85 | — | |
| MeZO-AdamUModel=Qwen2.5-14B2025.10 | 85 | 1.43 | |
| ZO Fine-tunerModel=OPT-30B2025.10 | 85 | 1.81 | |
| RoBERTa-WGSupervision Type=Supervised, Finetuning Dataset=WinoGrande, Zero-shot Transfer=True2021.05 | 84.4 | — | |
| H2OBackbone=OPT-30B2023.06 | 84 | — | |
| E5baseBackbone=LLaMA-7B, Evaluation Setting=Few-shot (ICL)2025.05 | 84 | — | |
| FLAPPruning ratio=25%, Backbone=Deepseek-V2-Lite, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 84 | — | |
| MC-SMoEPruning ratio=25%, Backbone=Deepseek-V2-Lite, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 84 | — | |
| Qwen3-30B-A3BPruning ratio=0%, Backbone=Qwen3-30B-A3B, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 84 | — | |
| MoNEPruning ratio=50%, Base Model=Deepseek-V2-Lite, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 84 | — | |
| Qwen3-30B-A3BPruning ratio=0%, Base Model=Qwen3-30B-A3B, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 84 | — | |
| FO-LoRAModel=Llama-2-7b2026.04 | 84 | — | |
| FO-PromptModel=Vicuna-7b-v1.52026.04 | 84 | — | |
| Hybrid-PromptModel=Vicuna-7b-v1.52026.04 | 84 | — | |
| Hybrid-LoRAModel=Vicuna-7b-v1.52026.04 | 84 | — | |
| Fisher MergingValidation=true2024.05 | 83.1 | — | |
| CBDSBackbone=LLaMA-7B, Evaluation Setting=Few-shot (ICL)2025.05 | 83 | — | |
| MC-SMoEPruning ratio=25%, Backbone=Qwen3-30B-A3B, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 83 | — | |
| MC-SMoEPruning ratio=50%, Base Model=Qwen3-30B-A3B, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 83 | — | |
| RSPruning ratio=50%, Base Model=Qwen3-30B-A3B, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 83 | — | |
| FO-PrefixModel=Llama-2-7b2026.04 | 83 | — | |
| Hybrid-PrefixModel=Vicuna-7b-v1.52026.04 | 83 | — | |
| KoCoModel Scale=1.6B Parameters, Training Tokens=8B Tokens, Pre-training Corpus=DCLM, Training Paradigm=Continual Pre-training2026.04 | 83 | — | |
| RoBERTa-LType=Encoder2025.05 | 83 | — | |
| MeZOModel=OPT-30B2025.10 | 83 | 1.93 | |
| EMR-MERGINGValidation=false2024.05 | 82.4 | — | |
| SpAttenBackbone=OPT-30B2023.06 | 82 | — | |
| EPRBackbone=LLaMA-7B, Evaluation Setting=Few-shot (ICL)2025.05 | 82 | — | |
| URL Prefix (MeCo)Model Scale=1.6B Parameters, Training Tokens=8B Tokens, Pre-training Corpus=DCLM2026.04 | 82 | — | |
| Standard CPTModel Scale=1.6B Parameters, Training Tokens=8B Tokens, Pre-training Corpus=DCLM, Training Paradigm=Continual Pre-training2026.04 | 82 | — |