Multiple Choice Question Answering on PIQA
80.5AccuracyOriginal
Evaluation Results
| Method | Links | |
|---|---|---|
| OriginalModel=Qwen3-30B-A3B, Reduction ratio=0%2026.05 | 80.5 | |
| M-SMoEModel=Qwen3-30B-A3B, Reduction ratio=25%2026.05 | 80.4 | |
| FrequencyModel=Qwen3-30B-A3B, Reduction ratio=25%2026.05 | 80.4 | |
| ConMoEModel=Qwen3-30B-A3B, Reduction ratio=25%2026.05 | 80.4 | |
| REAPModel=deepseek-moe-16b-base, Reduction ratio=25%2026.05 | 79.8 | |
| REAPModel=Qwen3-30B-A3B, Reduction ratio=25%2026.05 | 79.7 | |
| OriginalModel=deepseek-moe-16b-base, Reduction ratio=0%2026.05 | 79.7 | |
| ConMoEModel=deepseek-moe-16b-base, Reduction ratio=25%2026.05 | 79.6 | |
| OriginalModel=OLMoE-1B-7B-0125, Reduction ratio=0%2026.05 | 79.6 | |
| FrequencyModel=Qwen3-30B-A3B, Reduction ratio=50%2026.05 | 79.4 | |
| FrequencyModel=deepseek-moe-16b-base, Reduction ratio=25%2026.05 | 79.1 | |
| REAPModel=Qwen3-30B-A3B, Reduction ratio=50%2026.05 | 79 | |
| HC-SMoEModel=deepseek-moe-16b-base, Reduction ratio=25%2026.05 | 78.5 | |
| M-SMoEModel=OLMoE-1B-7B-0125, Reduction ratio=25%2026.05 | 78.2 | |
| AR (autoregressive)Zero-shot=true, Likelihood-based evaluation=true, Number of Parameters=1.7B2026.05 | 78.1 | |
| ConMoEModel=deepseek-moe-16b-base, Reduction ratio=50%2026.05 | 78.1 | |
| ConMoEModel=Qwen3-30B-A3B, Reduction ratio=50%2026.05 | 78 | |
| REAPModel=deepseek-moe-16b-base, Reduction ratio=50%2026.05 | 78 | |
| REAPModel=OLMoE-1B-7B-0125, Reduction ratio=25%2026.05 | 77.9 | |
| ConMoEModel=OLMoE-1B-7B-0125, Reduction ratio=25%2026.05 | 77.8 | |
| FrequencyModel=OLMoE-1B-7B-0125, Reduction ratio=25%2026.05 | 77.5 | |
| FrequencyModel=deepseek-moe-16b-base, Reduction ratio=50%2026.05 | 77.3 | |
| M-SMoEModel=deepseek-moe-16b-base, Reduction ratio=25%2026.05 | 77.2 | |
| AR (autoregressive, ours)Zero-shot=true, Likelihood-based evaluation=true, Number of Parameters=1.7B2026.05 | 76.9 | |
| HC-SMoEModel=OLMoE-1B-7B-0125, Reduction ratio=25%2026.05 | 76.2 | |
| HC-SMoEModel=Qwen3-30B-A3B, Reduction ratio=25%2026.05 | 75.8 | |
| M-SMoEModel=Qwen3-30B-A3B, Reduction ratio=50%2026.05 | 75.1 | |
| JTP (d=1)Horizon (d)=1, Parameters=1.3B, Training tokens=100B2025.11 | 74.92 | |
| HC-SMoEModel=deepseek-moe-16b-base, Reduction ratio=50%2026.05 | 74.6 | |
| LLaDa-8B-BaseZero-shot=true, Likelihood-based evaluation=true, Number of Parameters=8B2026.05 | 74.4 | |
| MTP (d=1)Horizon (d)=1, Parameters=1.3B, Training tokens=100B2025.11 | 74.32 | |
| NextLat (d=2)Horizon (d)=2, Parameters=1.3B, Training tokens=100B2025.11 | 73.61 | |
| GPTParameters=1.3B, Training tokens=100B2025.11 | 73.45 | |
| JTP (d=2)Horizon (d)=2, Parameters=1.3B, Training tokens=100B2025.11 | 73.34 | |
| NoPEModel Architecture=GLA, Parameters=1.3B, Training Tokens=26B2025.11 | 73.3 | |
| Selective RoPEModel Architecture=GLA, Parameters=1.3B, Training Tokens=26B2025.11 | 73.1 | |
| NextLat (d=1)Horizon (d)=1, Parameters=1.3B, Training tokens=100B2025.11 | 73.07 | |
| MTP (d=2)Horizon (d)=2, Parameters=1.3B, Training tokens=100B2025.11 | 72.8 | |
| RoPEModel Architecture=GLA, Parameters=1.3B, Training Tokens=26B2025.11 | 72.4 | |
| REAPModel=OLMoE-1B-7B-0125, Reduction ratio=50%2026.05 | 71 | |
| M-SMoEModel=deepseek-moe-16b-base, Reduction ratio=50%2026.05 | 70.4 | |
| FrequencyModel=OLMoE-1B-7B-0125, Reduction ratio=50%2026.05 | 70.4 | |
| RoPEModel Architecture=GLA, Parameters=370M, Training Tokens=10B2025.11 | 69.6 | |
| RoPEModel Architecture=Gated DeltaNet, Parameters=~400M, Training Tokens=10B2025.11 | 69.5 | |
| Selective RoPEModel Architecture=GLA, Parameters=370M, Training Tokens=10B2025.11 | 69.1 | |
| Selective RoPEModel Architecture=Gated DeltaNet, Parameters=~400M, Training Tokens=10B, Variant=stable_selective_rope2025.11 | 69.1 | |
| Selective RoPEModel Architecture=FoX, Parameters=370M, Training Tokens=10B, Output-norm setting=OFF2025.11 | 68.7 | |
| NoPEModel Architecture=Gated DeltaNet, Parameters=~400M, Training Tokens=10B2025.11 | 68.6 | |
| RoPEModel Architecture=FoX, Parameters=370M, Training Tokens=10B, Output-norm setting=OFF2025.11 | 68.4 | |
| ConMoEModel=OLMoE-1B-7B-0125, Reduction ratio=50%2026.05 | 68.3 | |
| NoPEModel Architecture=GLA, Parameters=370M, Training Tokens=10B2025.11 | 67.8 | |
| R2LMShots=0-sh, Candidate Type=long-target2026.06 | 67.2 | |
| HC-SMoEModel=OLMoE-1B-7B-0125, Reduction ratio=50%2026.05 | 67.1 | |
| NoPEModel Architecture=FoX, Parameters=370M, Training Tokens=10B, Output-norm setting=OFF2025.11 | 67 | |
| R2LM-PI (plug-in)Shots=0-sh, Candidate Type=long-target2026.06 | 64.6 | |
| Causal dLLMShots=0-sh, Candidate Type=long-target2026.06 | 63.8 | |
| HC-SMoEModel=Qwen3-30B-A3B, Reduction ratio=50%2026.05 | 63 | |
| Duo (uniform diffusion)Zero-shot=true, Likelihood-based evaluation=true, Number of Parameters=1.7B2026.05 | 62.7 | |
| MDLM (masked diffusion)Zero-shot=true, Likelihood-based evaluation=true, Number of Parameters=1.7B2026.05 | 62.2 | |
| M-SMoEModel=OLMoE-1B-7B-0125, Reduction ratio=50%2026.05 | 62 | |
| Bidirectional dLLMShots=0-sh, Candidate Type=long-target2026.06 | 61.5 | |
| SMDM-1BZero-shot=true, Likelihood-based evaluation=true, Number of Parameters=1B2026.05 | 60.3 | |
| MDLMevaluation protocol=likelihood-based answer scoring2026.05 | 58.27 | |
| CFM (ours)Zero-shot=true, Likelihood-based evaluation=true, Number of Parameters=1.7B2026.05 | 57.2 | |
| Eso-LM (interpolating)Zero-shot=true, Likelihood-based evaluation=true, Number of Parameters=1.7B2026.05 | 55.6 | |
| CoFReevaluation protocol=likelihood-based answer scoring2026.05 | 55.33 | |
| ChanceZero-shot=true, Likelihood-based evaluation=true2026.05 | 51.6 |