Multiple Choice Question Answering on ARC Challenge (test)
84.1AccuracyPPT (Distribution)
Evaluation Results
| Method | Links | |
|---|---|---|
| PPT (Distribution)Backbone=QWEN2-7B2026.05 | 84.1 | |
| PPT (Mean)Backbone=QWEN2-7B2026.05 | 83.7 | |
| BaseBackbone=QWEN2-7B2026.05 | 83.6 | |
| TaTTrain Dataset=ARC-C2026.03 | 82.17 | |
| PPT (Distribution)Backbone=LLAMA3-8B2026.05 | 79.6 | |
| BaseBackbone=LLAMA3-8B2026.05 | 78.6 | |
| TaTTrain Dataset=OpenQA2026.03 | 78.41 | |
| PPT (Mean)Backbone=LLAMA3-8B2026.05 | 77.4 | |
| Linear ProbeTrain Dataset=ARC-E2026.03 | 75.55 | |
| Linear ProbeTrain Dataset=ARC-C2026.03 | 75.32 | |
| TaTTrain Dataset=ARC-E2026.03 | 73.81 | |
| Linear ProbeTrain Dataset=ComQA2026.03 | 71.72 | |
| Linear ProbeTrain Dataset=CosQA2026.03 | 70.09 | |
| TaTTrain Dataset=CosQA2026.03 | 69.71 | |
| Linear ProbeTrain Dataset=OpenQA2026.03 | 66.42 | |
| TaTTrain Dataset=Hellaswag2026.03 | 65.96 | |
| Few-shot AccuracyMode=Few-shot2026.03 | 65.3 | |
| TaTTrain Dataset=ComQA2026.03 | 60.92 | |
| TaTTrain Dataset=SiQA2026.03 | 59.9 | |
| TWLABackbone=Qwen3-32B, Weight bits=1.58, Activation bits=16, Setting=Zero-shot2026.06 | 55.33 | |
| VerbalizedBackbone=LLAMA3-8B2026.05 | 54.8 | |
| UnifiedQA_T5-FTModel Size=LARGE, Protocol=Fine-Tuning2022.04 | 54.42 | |
| UnifiedQA_T5*Model Size=LARGE, Reference=Khashabi et al. (2020)2022.04 | 54.4 | |
| UnifiedQA_T5Model Size=LARGE2022.04 | 54.33 | |
| TWLABackbone=LLaMA2-70B, Weight bits=1.58, Activation bits=16, Setting=Zero-shot2026.06 | 53.75 | |
| TaTTrain Dataset=BoolQ2026.03 | 53.5 | |
| Linear ProbeTrain Dataset=Hellaswag2026.03 | 52.92 | |
| TWLABackbone=Qwen3-14B, Weight bits=1.58, Activation bits=16, Setting=Zero-shot2026.06 | 52.3 | |
| Zero-shot AccuracyMode=Zero-shot2026.03 | 50.1 | |
| Linear ProbeTrain Dataset=SiQA2026.03 | 49.82 | |
| UNPRUNEDACT.=2.8B, MEM.=100%, Speedup=1.0x, SFT Protocol=w/o SFT2024.11 | 48.3 | |
| GenMC_T5Model Size=LARGE2022.04 | 47.41 | |
| TWLABackbone=Qwen3-8B, Weight bits=1.58, Activation bits=16, Setting=Zero-shot2026.06 | 45.99 | |
| Linear ProbeTrain Dataset=BoolQ2026.03 | 44.51 | |
| UnifiedQA_T5Model Size=BASE2022.04 | 42.58 | |
| UnifiedQA_T5-FTModel Size=BASE, Protocol=Fine-Tuning2022.04 | 42.43 | |
| TWLABackbone=LLaMA2-13B, Weight bits=1.58, Activation bits=16, Setting=Zero-shot2026.06 | 42.41 | |
| CD-MoE-SR#L=9, ACT.=2.8B, MEM.=72.5%, Speedup=1.26x, SFT Protocol=w/ lightweight SFT2024.11 | 42.4 | |
| CD-MoE-S#L=8, ACT.=2.4B, MEM.=73.1%, Speedup=1.34x, SFT Protocol=w/ lightweight SFT2024.11 | 41.5 | |
| TWLABackbone=LLaMA3-8B, Weight bits=1.58, Activation bits=16, Setting=Zero-shot2026.06 | 39.42 | |
| GenMC_T5Model Size=BASE2022.04 | 39 | |
| CD-MoE-SR#L=9, ACT.=2.8B, MEM.=72.5%, Speedup=1.26x, SFT Protocol=w/o SFT2024.11 | 37.9 | |
| CD-MoE-S#L=8, ACT.=2.4B, MEM.=73.1%, Speedup=1.34x, SFT Protocol=w/o SFT2024.11 | 37.8 | |
| LayerDrop#L=8, ACT.=2.4B, MEM.=72.2%, Speedup=1.34x, SFT Protocol=w/ lightweight SFT2024.11 | 37.8 | |
| TWLABackbone=LLaMA2-7B, Weight bits=1.58, Activation bits=16, Setting=Zero-shot2026.06 | 37.63 | |
| M-SMoE#L=9, ACT.=2.8B, MEM.=70.0%, Speedup=1.0x, SFT Protocol=w/o SFT2024.11 | 37.1 | |
| BlockDrop#L=8, ACT.=2.3B, MEM.=71.3%, Speedup=1.42x, SFT Protocol=w/o SFT2024.11 | 36.2 | |
| OPENLLAMA-3BACT.=3B, SFT Protocol=w/o SFT2024.11 | 36.1 | |
| RoBERTaModel Size=LARGE2022.04 | 35.97 | |
| LayerDrop#L=8, ACT.=2.4B, MEM.=72.2%, Speedup=1.34x, SFT Protocol=w/o SFT2024.11 | 35.3 | |
| AskLLM-OBackbone=OpenLLaMA 3B, Checkpoint Selection Metric=next token prediction accuracy2024.11 | 35.15 | |
| RoBERTaModel Size=BASE2022.04 | 34.85 | |
| BADSBackbone=OpenLLaMA 3B, Checkpoint Selection Metric=next token prediction accuracy2024.11 | 34.39 | |
| MixingBackbone=OpenLLaMA 3B, Checkpoint Selection Metric=next token prediction accuracy2024.11 | 33.79 | |
| Duplicate_MetaBackbone=OpenLLaMA 3B, Checkpoint Selection Metric=next token prediction accuracy2024.11 | 33.28 | |
| Meta_OnlyBackbone=OpenLLaMA 3B, Checkpoint Selection Metric=next token prediction accuracy2024.11 | 32.08 | |
| ClassActBackbone=OpenLLaMA 3B, Checkpoint Selection Metric=next token prediction accuracy2024.11 | 31.91 | |
| OPT-2.7BACT.=2.7B, SFT Protocol=w/o SFT2024.11 | 31.2 | |
| ALBERTModel Size=LARGE2022.04 | 31.19 | |
| BLOOM-3BACT.=3B, SFT Protocol=w/o SFT2024.11 | 30.4 | |
| ALBERTModel Size=BASE2022.04 | 30.21 | |
| GPT-NEO-2.7BACT.=2.7B, SFT Protocol=w/o SFT2024.11 | 29.8 | |
| Random_SelectBackbone=OpenLLaMA 3B, Checkpoint Selection Metric=next token prediction accuracy2024.11 | 28.92 | |
| Duo++k=22026.02 | 26.11 | |
| UnifiedQA_T5*Model Size=BASE, Reference=Khashabi et al. (2020)2022.04 | 25.8 | |
| Duo++k=52026.02 | 25.77 | |
| Duo2026.02 | 25.43 | |
| Duo++k=32026.02 | 25 | |
| MDLM2026.02 | 24.66 | |
| AR Transformer2026.02 | 23.04 | |
| CDSBackbone=OpenLLaMA 3B, Checkpoint Selection Metric=next token prediction accuracy2024.11 | 21.16 |