Natural Language Understanding on ARC Challenge
95.3AccuracyLLaMA-3.1-405B Base
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| LLaMA-3.1-405B Base#Shots=25-shot, Architecture=Dense, # activated params=405B, # total params=405B2026.01 | 95.3 | — | |
| DeepSeek-V3-Base#Shots=25-shot, Architecture=MoE, # activated params=37B, # total params=671B2026.01 | 95.3 | — | |
| Yuan3.0-1T Base#Shots=25-shot, Architecture=MoE, # activated params=68.5B, # total params=1010B2026.01 | 94.3 | — | |
| Full-Attn# Shots=25-shot, Model Architecture=80B MoE (Hybrid 1:11), Attention Variant=Full-Attn2026.02 | 78.4 | — | |
| HySparse# Shots=25-shot, Model Architecture=80B MoE (Hybrid 1:11), Attention Variant=HySparse2026.02 | 77.6 | — | |
| HySparse# Shots=25-shot, Model Architecture=7B Dense (Hybrid 1:3), Attention Variant=HySparse2026.02 | 75 | — | |
| Hybrid SWA# Shots=25-shot, Model Architecture=7B Dense (Hybrid 1:3), Attention Variant=Hybrid SWA2026.02 | 74.9 | — | |
| Full-Attn# Shots=25-shot, Model Architecture=7B Dense (Hybrid 1:3), Attention Variant=Full-Attn2026.02 | 70.2 | — | |
| Hybrid SWA# Shots=25-shot, Model Architecture=80B MoE (Hybrid 1:11), Attention Variant=Hybrid SWA2026.02 | 63.9 | — | |
| NLSModel=Llama 3 70B, Precision=W8A8, Time Steps (T)=12026.05 | 62.54 | 78.64 | |
| Arcanazero-shot=true2024.10 | 61.4 | — | |
| Vicuna-v1.5zero-shot=true2024.10 | 56.6 | — | |
| LLaMA-2-Chatzero-shot=true2024.10 | 54.9 | — | |
| WizardLMzero-shot=true2024.10 | 47.5 | — | |
| NLSModel=Llama 2 7B, Precision=W6A6, Time Steps (T)=42026.05 | 43.34 | 66.89 | |
| LLaMA-2zero-shot=true2024.10 | 40.3 | — |