Science Question Answering on ARC Challenge
96AccuracyQwen-3-30B-A3B
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Qwen-3-30B-A3BOpenness=Open-weights, Regional Origin=Non-European, Post-training=Instruction-tuned2026.02 | 96 | — | — | |
| Qwen-3-32BOpenness=Open-weights, Regional Origin=Non-European, Post-training=Instruction-tuned2026.02 | 95.2 | — | — | |
| BaseZero-shot=true2026.03 | 95 | — | — | |
| Llama 3.1 InstructModel Scale=70B2025.04 | 94.7 | — | — | |
| Llama-3.3-70BOpenness=Open-weights, Regional Origin=Non-European, Post-training=Instruction-tuned2026.02 | 94.5 | — | — | |
| Qwen-3-14BOpenness=Open-weights, Regional Origin=Non-European, Post-training=Instruction-tuned2026.02 | 94.3 | — | — | |
| Llama 3 InstructModel Scale=70B2025.04 | 94.3 | — | — | |
| ParamΔModel Scale=70B2025.04 | 94.3 | — | — | |
| SFTZero-shot=true2026.03 | 94 | — | — | |
| OLMo-3.1-32BOpenness=Fully-open, Regional Origin=Non-European, Post-training=Instruction-tuned2026.02 | 93.6 | — | — | |
| Gemma-3-27BOpenness=Open-weights, Regional Origin=Non-European, Post-training=Instruction-tuned2026.02 | 93.5 | — | — | |
| Mistral-3.2-24BOpenness=Open-weights, Regional Origin=European, Post-training=Instruction-tuned2026.02 | 93.4 | — | — | |
| GRPOZero-shot=true2026.03 | 93 | — | — | |
| Gemma-3-12BOpenness=Open-weights, Regional Origin=Non-European, Post-training=Instruction-tuned2026.02 | 92.3 | — | — | |
| ConfidenceModel=Qwen-2.5-7B-Instruct2026.01 | 91.1 | — | — | |
| EuroLLM-22BOpenness=Fully-open, Regional Origin=European, Post-training=Instruction-tuned2026.02 | 89.8 | — | — | |
| HD + ConfidenceModel=Qwen-2.5-7B-Instruct2026.01 | 89 | — | — | |
| HD + Learn2AggModel=Qwen-2.5-7B-Instruct2026.01 | 88.9 | — | — | |
| Llama 3.1 BaseModel Scale=70B2025.04 | 88.9 | — | — | |
| High DiversityModel=Qwen-2.5-7B-Instruct2026.01 | 88.4 | — | — | |
| Debate 5 × 5Model=Qwen-2.5-7B-Instruct2026.01 | 88.2 | — | — | |
| ConfidenceModel=Llama-3.1-8B-Instruct2026.01 | 88.2 | — | — | |
| Majority VoteModel=Qwen-2.5-7B-Instruct2026.01 | 88 | — | — | |
| Ministral-3-8BModel Category=Text-Experts, Model Scale=8B2026.03 | 88 | — | — | |
| High DiversityModel=Llama-3.1-8B-Instruct2026.01 | 87.7 | — | — | |
| Llama 3 BaseModel Scale=70B2025.04 | 87.7 | — | — | |
| OLMo-3-7BOpenness=Fully-open, Regional Origin=Non-European, Post-training=Instruction-tuned2026.02 | 86.1 | — | — | |
| EuroLLM-9BOpenness=Fully-open, Regional Origin=European, Post-training=Instruction-tuned2026.02 | 85.9 | — | — | |
| HyperCLOVAX-8B-OmniModel Category=Unified, Model Scale=8B, Native speech I/O=true, Result Reproducibility=reproduced2026.03 | 85.8 | — | — | |
| HD + ConfidenceModel=Llama-3.1-8B-Instruct2026.01 | 85.4 | — | — | |
| single-threshold support basisModel=LLaDA-8B-Instruct, Degree of Chebyshev polynomial=62025.10 | 84.98 | — | — | |
| Apertus-70BOpenness=Fully-open, Regional Origin=European, Post-training=Instruction-tuned2026.02 | 84.7 | — | — | |
| Exact computationModel=LLaDA-8B-Instruct2025.10 | 84.64 | — | — | |
| Llama-3.1-8BOpenness=Open-weights, Regional Origin=Non-European, Post-training=Instruction-tuned2026.02 | 84.3 | — | — | |
| Llama 3.1 InstructModel Scale=8B2025.04 | 83.7 | — | — | |
| Majority VoteModel=Llama-3.1-8B-Instruct2026.01 | 83.2 | — | — | |
| HD + Learn2AggModel=Llama-3.1-8B-Instruct2026.01 | 83.2 | — | — | |
| ParamΔModel Scale=8B2025.04 | 82.8 | — | — | |
| single-threshold support basisModel=LLaDA-8B-Instruct, Degree of Chebyshev polynomial=42025.10 | 82.42 | — | — | |
| Qwen-14Brole=Teacher model, shot=5-shot2024.07 | 82.25 | — | — | |
| LLaDABudget=1/1, Evaluation Protocol=fully generative2026.05 | 82.2 | — | — | |
| Llama 3 InstructModel Scale=8B2025.04 | 82.1 | — | — | |
| ME-DLM Stage 3Budget=1/1, Evaluation Protocol=fully generative2026.05 | 81.1 | — | — | |
| ME-DLM Stage 3Budget=1/2, Evaluation Protocol=fully generative2026.05 | 80.7 | — | — | |
| Qwen-1.5 14BRole=Teacher2024.07 | 80.59 | — | — | |
| Qwen-1.5 14BModel Type=Teacher2024.07 | 80.59 | — | — | |
| Single ModelModel=Qwen-2.5-7B-Instruct2026.01 | 80.5 | — | — | |
| Debate 5 × 5Model=Llama-3.1-8B-Instruct2026.01 | 80.2 | — | — | |
| ME-DLM Stage 2Budget=1/1, Evaluation Protocol=fully generative2026.05 | 79.7 | — | — | |
| Clean FT (No WM)Model=Llama-3-8B, Condition=Original (Pre-Attack)2026.03 | 79.6 | — | — | |
| Weight Quant.Model=Mistral-7B, Condition=Original (Pre-Attack)2026.03 | 79.6 | — | — | |
| Weight Quant.Model=Qwen2.5-7B, Condition=Original (Pre-Attack)2026.03 | 78.8 | — | — | |
| Weight Quant.Model=Qwen2.5-7B, Condition=Attacked (Post-Attack)2026.03 | 78.8 | — | — | |
| EmMarkModel=Qwen2.5-7B, Condition=Attacked (Post-Attack)2026.03 | 78.6 | — | — | |
| EmMarkModel=Mistral-7B, Condition=Original (Pre-Attack)2026.03 | 78.4 | — | — | |
| Hard-routing MoEBackbone=Qwen2.5-7B-Instruct, # Params (%)=14.2G (100%), # Trans. (%)=14.2G (100%), Fine-tuning setting=DFL, Rank (rl)=322026.02 | 78.21 | — | — | |
| GPT3.5Source Task=RACE2024.05 | 78.2 | — | — | |
| Single ModelModel=Llama-3.1-8B-Instruct2026.01 | 78 | — | — | |
| GPT3.5Source Task=BoolQ2024.05 | 77.8 | — | — | |
| EmMarkModel=Llama-3-8B, Condition=Attacked (Post-Attack)2026.03 | 77.8 | — | — | |
| EmMarkModel=Llama-3-8B, Condition=Original (Pre-Attack)2026.03 | 77.4 | — | — | |
| Weight Quant.Model=Llama-3-8B, Condition=Attacked (Post-Attack)2026.03 | 77.4 | — | — | |
| GPT3.5Source Task=ARC-Easy2024.05 | 77.2 | — | — | |
| Naive Top-kModel=Qwen2.5-7B, Condition=Original (Pre-Attack)2026.03 | 77.2 | — | — | |
| EmMarkModel=Mistral-7B, Condition=Attacked (Post-Attack)2026.03 | 77.2 | — | — | |
| LLaDABudget=1/2, Evaluation Protocol=fully generative2026.05 | 77.2 | — | — | |
| Naive Top-kModel=Qwen2.5-7B, Condition=Attacked (Post-Attack)2026.03 | 77 | — | — | |
| Clean FT (No WM)Model=Mistral-7B, Condition=Original (Pre-Attack)2026.03 | 77 | — | — | |
| Full DIFSWModel=Mistral-7B, Condition=Original (Pre-Attack)2026.03 | 77 | — | — | |
| Weight Quant.Model=Mistral-7B, Condition=Attacked (Post-Attack)2026.03 | 76.6 | — | — | |
| Naive Top-kModel=Llama-3-8B, Condition=Attacked (Post-Attack)2026.03 | 76.4 | — | — | |
| Naive Top-kModel=Mistral-7B, Condition=Original (Pre-Attack)2026.03 | 76.4 | — | — | |
| Full DIFSWModel=Mistral-7B, Condition=Attacked (Post-Attack)2026.03 | 76.4 | — | — | |
| Sparse-and-Orthogonal LoRABackbone=Qwen2.5-7B-Instruct, # Params (%)=87M (0.60%), # Trans. (%)=42M (0.30%), Fine-tuning setting=DFL, Rank (rl)=322026.02 | 76.23 | — | — | |
| GPT3.5Source Task=Commonsense-QA2024.05 | 76.2 | — | — | |
| Weight Quant.Model=Llama-3-8B, Condition=Original (Pre-Attack)2026.03 | 76 | — | — | |
| Full DIFSWModel=Llama-3-8B, Condition=Attacked (Post-Attack)2026.03 | 76 | — | — | |
| Full DIFSWModel=Llama-3-8B, Condition=Original (Pre-Attack)2026.03 | 75.8 | — | — | |
| Clean FT (No WM)Model=Qwen2.5-7B, Condition=Original (Pre-Attack)2026.03 | 75.8 | — | — | |
| Apertus-8BOpenness=Fully-open, Regional Origin=European, Post-training=Instruction-tuned2026.02 | 75.5 | — | — | |
| GPT3.5Source Task=MNLI2024.05 | 75.2 | — | — | |
| Sparse-and-Orthogonal LoRA (Single)Backbone=Qwen2.5-7B-Instruct, # Params (%)=87M (0.60%), # Trans. (%)=42M (0.30%), Fine-tuning setting=DFL, Rank (rl)=322026.02 | 75.17 | — | — | |
| GPT3.5Source Task=Zero-shot2024.05 | 74.6 | — | — | |
| GPT3.5Source Task=AG-news2024.05 | 74.4 | — | — | |
| ME-DLM Stage 2Budget=1/2, Evaluation Protocol=fully generative2026.05 | 74.4 | — | — | |
| Naive Top-kModel=Llama-3-8B, Condition=Original (Pre-Attack)2026.03 | 74.2 | — | — | |
| GPT3.5Source Task=QQP2024.05 | 74 | — | — | |
| ME-DLM Stage 3Budget=1/4, Evaluation Protocol=fully generative2026.05 | 73.6 | — | — | |
| Full DIFSWModel=Qwen2.5-7B, Condition=Original (Pre-Attack)2026.03 | 73.4 | — | — | |
| FPFTBackbone=Qwen2.5-7B-Instruct, # Params (%)=14.2G (100%), # Trans. (%)=14.2G (100%), Fine-tuning setting=DFL, Rank (rl)=322026.02 | 73.29 | — | — | |
| Full DIFSWModel=Qwen2.5-7B, Condition=Attacked (Post-Attack)2026.03 | 73.2 | — | — | |
| Naive Top-kModel=Mistral-7B, Condition=Attacked (Post-Attack)2026.03 | 73.2 | — | — | |
| GPT3.5Source Task=SST22024.05 | 72.6 | — | — | |
| EmMarkModel=Llama-2-7B, Condition=Attacked (Post-Attack)2026.03 | 72.6 | — | — | |
| Weight Quant.Model=Llama-2-7B, Condition=Attacked (Post-Attack)2026.03 | 72.4 | — | — | |
| EmMarkModel=DeepSeek-7B, Condition=Attacked (Post-Attack)2026.03 | 72.4 | — | — | |
| Qwen-1.5 7BRole=Teacher2024.07 | 72.3 | — | — | |
| Naive Top-kModel=Llama-2-7B, Condition=Original (Pre-Attack)2026.03 | 72.2 | — | — | |
| Clean FT (No WM)Model=DeepSeek-7B, Condition=Original (Pre-Attack)2026.03 | 72.2 | — | — | |
| EmMarkModel=Qwen2.5-7B, Condition=Original (Pre-Attack)2026.03 | 72 | — | — |