Accuracy Evaluation on ARC Challenge (Reasoning)
97.2AccuracyQwen3
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3Params=32B2026.04 | 97.2 | |
| Llama 3 405BEvaluation Protocol=0-shot2026.03 | 96.9 | |
| I-DLMParams=32B, Decoding=ISD (N=4, sampling)2026.04 | 96.8 | |
| GPT-4oEvaluation Protocol=0-shot2026.03 | 96.7 | |
| Claude 3.5 SonnetEvaluation Protocol=0-shot2026.03 | 96.7 | |
| I-DLMParams=8B, Decoding=ISD (N=4, sampling)2026.04 | 95.8 | |
| Qwen3Params=8B2026.04 | 95.8 | |
| Llama 3 70BEvaluation Protocol=0-shot2026.03 | 94.8 | |
| SDARParams=30B2026.04 | 93.2 | |
| CharacterFlywheel V7Evaluation Protocol=0-shot2026.03 | 93.1 | |
| SDARParams=8B2026.04 | 91.9 | |
| KALEBackbone=Qwen3-32B-thinking2026.01 | 91.09 | |
| LLaDA-2.1-miniParams=16B2026.04 | 90.2 | |
| CoT2-MetaStrategy=Ours (CoT2-Meta), Inference Budget=C=162026.03 | 86.3 | |
| MFSBackbone=Qwen2.5-3B-Instruct2026.01 | 85.05 | |
| SFTBackbone=Qwen3-32B-thinking2026.01 | 84.22 | |
| ReST-MCTS*Strategy=ReST-MCTS*, Inference Budget=C=162026.03 | 84.2 | |
| Llama 3 8BEvaluation Protocol=0-shot2026.03 | 83.4 | |
| Vanilla ToTStrategy=Vanilla ToT, Inference Budget=C=162026.03 | 82.4 | |
| VanillaBackbone=Qwen3-32B-thinking2026.01 | 81.06 | |
| Best-of-16Strategy=Best-of-16, Inference Budget=C=162026.03 | 80.1 | |
| phi-DecodingBackbone=Qwen2.5-3B-Instruct2026.01 | 79.69 | |
| Predictive DecodingBackbone=Qwen2.5-3B-Instruct2026.01 | 78.69 | |
| Guided DecodingBackbone=Qwen2.5-3B-Instruct2026.01 | 78.34 | |
| Tree-of-ThoughtsBackbone=Qwen2.5-3B-Instruct2026.01 | 78.23 | |
| AR (CoT)Backbone=Qwen2.5-3B-Instruct2026.01 | 77.47 | |
| Engram-40BShots=25-shot2026.01 | 76.4 | |
| Greedy CoTStrategy=Greedy CoT, Inference Budget=C=162026.03 | 76.2 | |
| ASCTotal AI agents=30, Cooperative agents=10, GPU=NVIDIA Tesla P1002026.02 | 75 | |
| Engram-27BShots=25-shot2026.01 | 73.8 | |
| MoE-27BShots=25-shot2026.01 | 70.1 | |
| LAWTotal AI agents=30, Cooperative agents=10, GPU=NVIDIA Tesla P1002026.02 | 65 | |
| DefaultModel=Qwen2.5-14B2026.02 | 61.9 | |
| MixAT-GCGModel=Qwen2.5-14B2026.02 | 61.9 | |
| B3D-RWKV-7.2BModel Type=Strictly Causal Diffusion LM, Few-shot examples=02026.05 | 61.6 | |
| CATModel=Qwen2.5-14B2026.02 | 60.7 | |
| Dream-7BModel Type=Diffusion LM, Few-shot examples=02026.05 | 59.8 | |
| Dense-4BShots=25-shot2026.01 | 59.3 | |
| DATModel=Qwen2.5-14B2026.02 | 58.4 | |
| Qwen3-8BModel Type=Causal LM, Few-shot examples=02026.05 | 56.6 | |
| GATotal AI agents=30, Cooperative agents=10, GPU=NVIDIA Tesla P1002026.02 | 56 | |
| RWKV-7-7.2BModel Type=Causal LM, Few-shot examples=02026.05 | 55.5 | |
| Teacher (Qwen3-0.6B)Role=Teacher, Model=Qwen3-0.6B, Evaluation Protocol=generation-based evaluation2026.03 | 55 | |
| Diffusion-onlyModel=Llama3-8B2026.02 | 53.6 | |
| LLaMA3-8BModel Type=Causal LM, Few-shot examples=02026.05 | 53.6 | |
| CBModel=Llama3-8B2026.02 | 53.3 | |
| DefaultModel=Llama3-8B2026.02 | 53 | |
| DATModel=Llama3-8B2026.02 | 52.1 | |
| CATModel=Llama3-8B2026.02 | 52 | |
| MixAT-GCGModel=Llama3-8B2026.02 | 51.3 | |
| FP16Model=Hymba-Instruct-1.5B, Compression Ratio=1x2026.01 | 48.89 | |
| QMCModel=Hymba-Instruct-1.5B, Memory=2bits-MLC, Compression Ratio=4.44x2026.01 | 48.63 | |
| LATModel=Llama3-8B2026.02 | 48.5 | |
| Hybrid KDASequence Mixer=KDA, Attention Layers=7, Evaluation Protocol=generation-based evaluation2026.03 | 48.4 | |
| FP16Model=Phi-1.5B, Compression Ratio=1x2026.01 | 48.12 | |
| LLaDA-8BModel Type=Diffusion LM, Few-shot examples=02026.05 | 47.5 | |
| QMCModel=Hymba-Instruct-1.5B, Memory=3bits-MLC, Compression Ratio=4.44x2026.01 | 47.35 | |
| FP16Model=Qwen2.5-1.5B-Instruct, Compression Ratio=1x2026.01 | 46.93 | |
| RTNModel=Phi-1.5B, Format=INT4, Compression Ratio=4x2026.01 | 46.25 | |
| Hybrid MambaSequence Mixer=Mamba2, Attention Layers=7, Evaluation Protocol=generation-based evaluation2026.03 | 46.2 | |
| QMCModel=Phi-1.5B, Memory=2bits-MLC, Compression Ratio=4.44x2026.01 | 45.99 | |
| FP16Model=LLaMA-3.2-3B, Compression Ratio=1x2026.01 | 45.9 | |
| QMCModel=Qwen2.5-1.5B-Instruct, Memory=3bits-MLC, Compression Ratio=4.44x2026.01 | 45.82 | |
| QMCModel=Phi-1.5B, Memory=3bits-MLC, Compression Ratio=4.44x2026.01 | 45.56 | |
| QMCModel=Qwen2.5-1.5B-Instruct, Memory=2bits-MLC, Compression Ratio=4.44x2026.01 | 45.56 | |
| QMCModel=LLaMA-3.2-3B, Memory=2bits-MLC, Compression Ratio=4.44x2026.01 | 45.05 | |
| MXINT4Model=Phi-1.5B, Compression Ratio=4x2026.01 | 44.8 | |
| MXINT4Model=Qwen2.5-1.5B-Instruct, Compression Ratio=4x2026.01 | 43.26 | |
| MXINT4Model=Hymba-Instruct-1.5B, Compression Ratio=4x2026.01 | 41.13 | |
| RandASTotal AI agents=30, Cooperative agents=10, GPU=NVIDIA Tesla P1002026.02 | 41 | |
| QMCModel=LLaMA-3.2-3B, Memory=3bits-MLC, Compression Ratio=4.44x2026.01 | 40.36 | |
| RTNModel=Qwen2.5-1.5B-Instruct, Format=INT4, Compression Ratio=4x2026.01 | 40.1 | |
| RTNModel=Hymba-Instruct-1.5B, Format=INT4, Compression Ratio=4x2026.01 | 36.86 | |
| MXINT4Model=LLaMA-3.2-3B, Compression Ratio=4x2026.01 | 36.01 | |
| RTNModel=LLaMA-3.2-3B, Format=INT4, Compression Ratio=4x2026.01 | 35.84 | |
| SmoothQuantBit=W8A8, Format=HiF2026.02 | 34 | |
| Pure KDASequence Mixer=KDA, Attention Layers=0, Evaluation Protocol=generation-based evaluation2026.03 | 33.9 | |
| SmoothQuantBit=W8A8, Format=MXFP2026.02 | 33.2 | |
| RTNBit=W8A8, Format=INT2026.02 | 33.1 | |
| SVDQuantBit=W8A8, Format=INT2026.02 | 32.8 | |
| RTNBit=W8A8, Format=HiF2026.02 | 32.7 | |
| SVDQuantBit=W8A8, Format=HiF2026.02 | 32.4 | |
| BF16Format=BF162026.02 | 32.3 | |
| SVDQuantBit=W8A8, Format=MXFP2026.02 | 32.2 | |
| RTNBit=W4A4, Format=NVFP2026.02 | 32.2 | |
| RTNBit=W8A8, Format=MXFP2026.02 | 32.1 | |
| SmoothQuantBit=W8A8, Format=INT2026.02 | 32.1 | |
| SVDQuantBit=W4A4, Format=NVFP2026.02 | 32.1 | |
| SmoothQuantBit=W4A4, Format=MXFP2026.02 | 31.5 | |
| SVDQuantBit=W4A4, Format=HiF2026.02 | 31.5 | |
| SmoothQuantBit=W4A4, Format=NVFP2026.02 | 31.3 | |
| SVDQuantBit=W4A4, Format=MXFP2026.02 | 31.3 | |
| RTNBit=W4A4, Format=HiF2026.02 | 31 | |
| RTNBit=W4A4, Format=MXFP2026.02 | 30.8 | |
| SmoothQuantBit=W4A4, Format=HiF2026.02 | 30.6 | |
| DiffuMambaModel Type=Strictly Causal Diffusion LM, Few-shot examples=02026.05 | 28.3 | |
| SVDQuantBit=W4A4, Format=INT2026.02 | 24.9 | |
| RTNBit=W4A4, Format=INT2026.02 | 22 | |
| SmoothQuantBit=W4A4, Format=INT2026.02 | 22 | |
| Pure MambaSequence Mixer=Mamba2, Attention Layers=0, Evaluation Protocol=generation-based evaluation2026.03 | 19.9 |