Science Question Answering on ARC-C
96.3AccuracyGPT-4
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| GPT-42024.07 | 96.3 | — | — | — | |
| Llama 3 405B2024.07 | 96.1 | — | — | — | |
| Nemotron 4 340B2024.07 | 94.3 | — | — | — | |
| Llama 3 70B2024.07 | 92.9 | — | — | — | |
| Mixtral 8x22B2024.07 | 91.9 | — | — | — | |
| GCPOBackbone=Qwen3-4B2026.05 | 91.4 | — | — | — | |
| Qwen3-4BParams=4B2025.12 | 89.83 | — | — | — | |
| LMSIPrompting Method=Self-Consistency2022.10 | 89.8 | — | — | — | |
| DAPOBackbone=Qwen3-4B2026.05 | 89.6 | — | — | — | |
| Wang et al. (2022b)description=Previous SOTA2022.10 | 88.7 | — | — | — | |
| Self-ConsistencyLMSI=false2022.10 | 88.7 | — | — | — | |
| TodyCommAttack Rate=< 50%2026.02 | 88.52 | 350 | — | — | |
| MEDALBackbone=Dream2025.12 | 88.5 | — | — | — | |
| DIVERBackbone=Qwen3-4B2026.05 | 88.4 | — | — | — | |
| LMSIPrompting Method=CoT-Prompting2022.10 | 88.3 | — | — | — | |
| Qwen 3-8BModel Category=Open weight models, Instruction-tuned=true2026.03 | 87.4 | — | — | — | |
| LMSIPrompting Method=Standard-Prompting2022.10 | 87.2 | — | — | — | |
| Standard-PromptingLMSI=false2022.10 | 87.1 | — | — | — | |
| Gemma 3-12BModel Category=Open weight models, Instruction-tuned=true2026.03 | 86.9 | — | — | — | |
| Div-R1Backbone=Qwen3-4B2026.05 | 86.9 | — | — | — | |
| AgentPruneAttack Rate=< 50%2026.02 | 86.73 | 301 | — | — | |
| DQOBackbone=Qwen3-4B2026.05 | 86.5 | — | — | — | |
| Qwen2.5-3BParams=3B2025.12 | 86.44 | — | — | — | |
| TodyCommAttack Rate== 50%2026.02 | 85.95 | 361 | — | — | |
| LLaDA + Geometric AlignmentModel Type=dLLM, Gen Length=2562026.04 | 85.9 | — | — | — | |
| Gemma 2-9BModel Category=Open weight models, Instruction-tuned=true2026.03 | 85.6 | — | — | — | |
| LLaDA + Geometric AlignmentModel Type=dLLM, Gen Length=1282026.04 | 85.5 | — | — | — | |
| CoT-PromptingLMSI=false2022.10 | 85.2 | — | — | — | |
| LLaDA + Geometric AlignmentModel Type=dLLM, Gen Length=5122026.04 | 85.2 | — | — | — | |
| GRPOBackbone=Qwen3-4B2026.05 | 84.6 | — | — | — | |
| Qwen 2.5-7BModel Category=Open weight models, Instruction-tuned=true2026.03 | 83.7 | — | — | — | |
| LLaDA + SFTModel Type=dLLM, Gen Length=1282026.04 | 83.7 | — | — | — | |
| LLaDA + Outcome RLModel Type=dLLM, Gen Length=2562026.04 | 83.5 | — | — | — | |
| LLaDA + SFTModel Type=dLLM, Gen Length=2562026.04 | 83.2 | — | — | — | |
| LLaDA + Outcome RLModel Type=dLLM, Gen Length=1282026.04 | 83.1 | — | — | — | |
| G-DesignerAttack Rate=< 50%2026.02 | 83.05 | 585 | — | — | |
| MEDALBackbone=LLaDA1.52025.12 | 82.8 | — | — | — | |
| LLaMA3-8BArchitecture=AR Models2026.04 | 82.4 | — | — | — | |
| LLaDA + Outcome RLModel Type=dLLM, Gen Length=5122026.04 | 82.2 | — | — | — | |
| MEDALBackbone=LLaDA2025.12 | 82.1 | — | — | — | |
| Complete GraphAttack Rate=< 50%2026.02 | 81.17 | 793 | — | — | |
| TodyCommAttack Rate=> 50%2026.02 | 81.05 | 366 | — | — | |
| Qwen3-1.7BParams=1.7B2025.12 | 80.34 | — | — | — | |
| Bst5Backbone=Dream2025.12 | 80 | — | — | — | |
| Llama 3 8B2024.07 | 79.7 | — | — | — | |
| DAPOBackbone=Qwen3-1.7B2026.05 | 79.4 | — | — | — | |
| Qwen2.5-1.5BParams=1.5B2025.12 | 79.32 | — | — | — | |
| Random GraphAttack Rate=< 50%2026.02 | 79.15 | 642 | — | — | |
| GCPOBackbone=Qwen3-1.7B2026.05 | 79.1 | — | — | — | |
| Gervasio-8BModel Category=Open weight models, Instruction-tuned=true2026.03 | 79 | — | — | — | |
| AMALIA-9B-DPOModel Category=Fully open models, Training=DPO, Instruction-tuned=true2026.03 | 78.9 | — | — | — | |
| Gemma 7B2024.07 | 78.6 | — | — | — | |
| DIVERBackbone=Qwen3-1.7B2026.05 | 78.3 | — | — | — | |
| Mistral 7B2024.07 | 78.2 | — | — | — | |
| LLaDA + SFTModel Type=dLLM, Gen Length=5122026.04 | 78.1 | — | — | — | |
| DreamBackbone=Dream2025.12 | 78 | — | — | — | |
| AMALIA-9B-SFTModel Category=Fully open models, Training=SFT, Instruction-tuned=true2026.03 | 77.9 | — | — | — | |
| Div-R1Backbone=Qwen3-1.7B2026.05 | 77.8 | — | — | — | |
| Bst5Backbone=LLaDA1.52025.12 | 77.6 | — | — | — | |
| SmolLM3-3BParams=3B2025.12 | 77.29 | — | — | — | |
| Bst5Backbone=LLaDA2025.12 | 77.2 | — | — | — | |
| DQOBackbone=Qwen3-1.7B2026.05 | 76.4 | — | — | — | |
| Llama 3.1-8BModel Category=Open weight models, Instruction-tuned=true2026.03 | 75.8 | — | — | — | |
| GRPOBackbone=Qwen3-1.7B2026.05 | 75.6 | — | — | — | |
| Direct Fine-tuningrank=1, training_steps=+200 steps2026.02 | 75.48 | — | — | — | |
| Direct Fine-tuningrank=1, training_steps=+700 steps2026.02 | 75.13 | — | — | — | |
| LLaDA1.5Backbone=LLaDA1.52025.12 | 74.9 | — | — | — | |
| In-Squeezereduction=128 ... -> 1, strategy=Min steps2026.02 | 74.78 | — | — | — | |
| Direct Fine-tuningrank=1, training_steps=+0 steps2026.02 | 74.7 | — | — | — | |
| Cont-Squeezereduction=128 -> 1, training_steps=+700 steps2026.02 | 74.18 | — | — | — | |
| Cont-Squeezereduction=128 -> 1, training_steps=+200 steps2026.02 | 74.13 | — | — | — | |
| AgentPruneAttack Rate== 50%2026.02 | 74.02 | 353 | — | — | |
| In-Squeezereduction=128 ... -> 1, strategy=Standard2026.02 | 73.61 | — | — | — | |
| Ministral-8BModel Category=Open weight models, Instruction-tuned=true2026.03 | 73.6 | — | — | — | |
| G-DesignerAttack Rate== 50%2026.02 | 72.91 | 617 | — | — | |
| LLaDABackbone=LLaDA2025.12 | 72.2 | — | — | — | |
| llama-3.2-3BParams=3B2025.12 | 72.2 | — | — | — | |
| EuroLLM-9BModel Category=Fully open models, Instruction-tuned=true2026.03 | 71.2 | — | — | — | |
| LlamaBackbone=Llama2025.12 | 70.5 | — | — | — | |
| Qwen2-1.5BParams=1.5B2025.12 | 70.17 | — | — | — | |
| AgentPruneAttack Rate=> 50%2026.02 | 69.79 | 369 | — | — | |
| Gemma2-9BArchitecture=AR Models2026.04 | 68.4 | — | — | — | |
| Apertus-8BModel Category=Fully open models, Instruction-tuned=true2026.03 | 68.3 | — | — | — | |
| Qwen3-0.6BParams=0.6B2025.12 | 68.14 | — | — | — | |
| Complete GraphAttack Rate== 50%2026.02 | 67.67 | 785 | — | — | |
| G-DesignerAttack Rate=> 50%2026.02 | 67.22 | 610 | — | — | |
| Qwen3-4B(Base)Backbone=Qwen3-4B2026.05 | 66.9 | — | — | — | |
| gemma2-2BParams=2B2025.12 | 66.44 | — | — | — | |
| PCMind-2.1-Kaiyuan-2BParams=2B2025.12 | 66.1 | — | — | — | |
| Random GraphAttack Rate=> 50%2026.02 | 64.88 | 628 | — | — | |
| YuLan-Mini-2.4BParams=2.4B2025.12 | 64.75 | — | — | — | |
| Complete GraphAttack Rate=> 50%2026.02 | 64.33 | 769 | — | — | |
| Random GraphAttack Rate== 50%2026.02 | 64.1 | 614 | — | — | |
| Qwen2.5 7BArchitecture=Autoregressive, Parameters=7B, Number of shots=02026.01 | 63.7 | — | — | — | |
| weight decayBackbone=Qwen-0.6B2025.12 | 62.4 | — | — | — | |
| Mistral-7BModel Category=Open weight models, Instruction-tuned=true2026.03 | 61.8 | — | — | — | |
| Attn-QATModel=Llama 3.1-70B, Precision=Attn-QAT2026.02 | 61.53 | — | — | — | |
| Llama 3.1-70BPrecision=BF162026.02 | 61.35 | — | — | — | |
| Attn-QATModel=Qwen3-14B, Precision=Attn-QAT2026.02 | 60.84 | — | — | — | |
| Qwen2 7BArchitecture=Autoregressive, Parameters=7B, Number of shots=02026.01 | 60.6 | — | — | — |