Question Answering on OpenBookQA (Normalized Accuracy)
55.6Normalized AccuracyLlama-Instruct
Evaluation Results
| Method | Links | |
|---|---|---|
| Llama-InstructShots=10-shot2026.03 | 55.6 | |
| HATified-SFTShots=10-shot2026.03 | 52.6 | |
| BaseModel=Llama-70B2026.04 | 49 | |
| BaseModel=Qwen2.5-7B2026.04 | 47.4 | |
| BaseModel=Mistral-7B2026.04 | 46.2 | |
| BaseModel=Llama-13B2026.04 | 45.6 | |
| Llama 3-8BShots=0-shot2024.12 | 45 | |
| Llama 3-8B E8T2Shots=0-shot2024.12 | 44.8 | |
| GlobalGCBackbone=Llama-2 7B, Zero-Shot=true2025.02 | 33.6 | |
| ClipByValueBackbone=Llama-2 7B, Zero-Shot=true2025.02 | 33.2 | |
| GlobalGCBackbone=Mixtral 8x1B, Zero-Shot=true2025.02 | 33 | |
| AGCBackbone=Llama-2 7B, Zero-Shot=true2025.02 | 32.8 | |
| AdaGCBackbone=Llama-2 7B, Zero-Shot=true2025.02 | 32.8 | |
| Mamba-2Model Scale=440M, Evaluation Protocol=Zero-shot2026.04 | 32.8 | |
| Mamba-2-1.5BTraining Tokens=100B, Context Length=2K, Scale=1.5B2026.03 | 32.6 | |
| Mamba-3-MIMO-1.5BTraining Tokens=100B, Context Length=2K, Scale=1.5B, MIMO Rank=42026.03 | 32.6 | |
| Mamba-2 + PoSTModel Scale=440M, Evaluation Protocol=Zero-shot2026.04 | 32.6 | |
| ClippyBackbone=Llama-2 7B, Zero-Shot=true2025.02 | 32.4 | |
| AdaGCBackbone=Mixtral 8x1B, Zero-Shot=true2025.02 | 32.2 | |
| Mamba-3-SISO-1.5BTraining Tokens=100B, Context Length=2K, Scale=1.5B2026.03 | 32 | |
| RWKV-7 + PoSTModel Scale=180M, Evaluation Protocol=Zero-shot2026.04 | 32 | |
| Gated DeltaNet + PoSTModel Scale=180M, Evaluation Protocol=Zero-shot2026.04 | 31.8 | |
| EFLAParameters=1.3B, Zero-shot=true2025.12 | 31.6 | |
| GDN-1.5BTraining Tokens=100B, Context Length=2K, Scale=1.5B2026.03 | 31.6 | |
| BaseModel=GPT-2 (774M)2026.04 | 31.2 | |
| EFLA + Loose βParameters=340M, Zero-shot=true2025.12 | 30.8 | |
| GatedFWA-NSAClass=Attention, Implementation=Triton, Linear=true2025.12 | 30.8 | |
| Mamba-2Model Scale=180M, Evaluation Protocol=Zero-shot2026.04 | 30.6 | |
| Gated DeltaNetModel Scale=180M, Evaluation Protocol=Zero-shot2026.04 | 30.6 | |
| ClipByValueBackbone=Llama-2 1.3B, Zero-Shot=true2025.02 | 30.4 | |
| AdaGCBackbone=Llama-2 1.3B, Zero-Shot=true2025.02 | 30.4 | |
| GlobalGCBackbone=Llama-2 1.3B, Zero-Shot=true2025.02 | 30.2 | |
| Mamba-2-880MTraining Tokens=100B, Context Length=2K, Scale=880M2026.03 | 30.2 | |
| Transformer (LLaMA) + SWA + NSAClass=Attention, Implementation=Triton, Linear=true2025.12 | 30.18 | |
| HGRN2Class=RNN-Like, Implementation=Triton, Linear=true2025.12 | 30 | |
| ClippyBackbone=Llama-2 1.3B, Zero-Shot=true2025.02 | 30 | |
| Mamba-3-SISO-880MTraining Tokens=100B, Context Length=2K, Scale=880M2026.03 | 30 | |
| Mamba-2 + PoSTModel Scale=180M, Evaluation Protocol=Zero-shot2026.04 | 30 | |
| GatedFWAClass=Attention, Implementation=Triton, Linear=true2025.12 | 29.93 | |
| DeltaNetParameters=1.3B, Zero-shot=true2025.12 | 29.8 | |
| PLDRv51-SOC-110M-5Zero-shot=true2026.03 | 29.8 | |
| Transformer-1.5BTraining Tokens=100B, Context Length=2K, Scale=1.5B2026.03 | 29.6 | |
| Transformer (LLaMA)Class=Attention, Implementation=CUDA, Linear=false2025.12 | 29.4 | |
| Transformer (LLaMA) + SWAClass=Attention, Implementation=CUDA, Linear=true2025.12 | 29.22 | |
| GLAClass=RNN-Like, Implementation=Triton, Linear=true2025.12 | 29.2 | |
| PLDRv51-SOC-110M-4Zero-shot=true2026.03 | 29.2 | |
| DeltaNetClass=RNN-Like, Implementation=Triton, Linear=true2025.12 | 29 | |
| PolychromaticLMstage=Base2026.03 | 29 | |
| RWKV-7Model Scale=180M, Evaluation Protocol=Zero-shot2026.04 | 29 | |
| GDN-880MTraining Tokens=100B, Context Length=2K, Scale=880M2026.03 | 28.6 | |
| Mamba-3-MIMO-880MTraining Tokens=100B, Context Length=2K, Scale=880M, MIMO Rank=42026.03 | 28.6 | |
| RetNetClass=RNN-Like, Implementation=CUDA, Linear=true2025.12 | 28.4 | |
| Mamba-3-MIMO-440MTraining Tokens=100B, Context Length=2K, Scale=440M, MIMO Rank=42026.03 | 28.4 | |
| ABL-SOC-110M-3Zero-shot=true2026.03 | 28.2 | |
| MambaClass=RNN-Like, Implementation=CUDA, Linear=true2025.12 | 28 | |
| GDN-440MTraining Tokens=100B, Context Length=2K, Scale=440M2026.03 | 27.6 | |
| ABL-SOC-110M-1Zero-shot=true2026.03 | 27.6 | |
| PLDRv51-SOC-110M-3Zero-shot=true2026.03 | 27.2 | |
| DeltaNetParameters=340M, Zero-shot=true2025.12 | 27 | |
| EFLA + Adaptive DecayParameters=340M, Zero-shot=true2025.12 | 27 | |
| PLDRv51-SOC-110M-1Zero-shot=true2026.03 | 27 | |
| PolychromaticLMstage=SFT2026.03 | 26.8 | |
| Transformer-880MTraining Tokens=100B, Context Length=2K, Scale=880M2026.03 | 26.8 | |
| EFLAParameters=340M, Zero-shot=true2025.12 | 26.6 | |
| ABL-SOC-110M-2Zero-shot=true2026.03 | 26.6 | |
| RFMoEScale=L, Size=870.6M, FLOPs=613.2M, Zero-shot=true2026.04 | 26.6 | |
| RFMoEScale=M, Size=307.3M, FLOPs=249.2M, Zero-shot=true2026.04 | 26.4 | |
| PLDRv51-SOC-110M-2Zero-shot=true2026.03 | 26.2 | |
| GPT-Neo-125MZero-shot=true2026.03 | 26.2 | |
| Transformer-440MTraining Tokens=100B, Context Length=2K, Scale=440M2026.03 | 26 | |
| Mamba-2-440MTraining Tokens=100B, Context Length=2K, Scale=440M2026.03 | 26 | |
| Mamba-3-SISO-440MTraining Tokens=100B, Context Length=2K, Scale=440M2026.03 | 26 | |
| SUB-SOC-110M-1Zero-shot=true2026.03 | 26 | |
| ReMoEScale=S, Size=92.44M, FLOPs=90.93M, Zero-shot=true2026.04 | 26 | |
| SUB-SOC-110M-2Zero-shot=true2026.03 | 25.2 | |
| Random2026.03 | 25 | |
| AoEScale=S, Size=93.85M, FLOPs=88.57M, Zero-shot=true2026.04 | 25 | |
| MoEScale=L, Size=808.4M, FLOPs=608.4M, Zero-shot=true2026.04 | 25 | |
| MoEScale=S, Size=92.44M, FLOPs=90.93M, Zero-shot=true2026.04 | 24.6 | |
| MoEScale=M, Size=289.9M, FLOPs=248.0M, Zero-shot=true2026.04 | 24.6 | |
| RFMoEScale=S, Size=95.32M, FLOPs=91.08M, Zero-shot=true2026.04 | 24.4 | |
| Mamba-2-180MTraining Tokens=100B, Context Length=2K, Scale=180M2026.03 | 23.2 | |
| Mamba-3-SISO-180MTraining Tokens=100B, Context Length=2K, Scale=180M2026.03 | 22.8 | |
| GDN-180MTraining Tokens=100B, Context Length=2K, Scale=180M2026.03 | 22 | |
| Mamba-3-MIMO-180MTraining Tokens=100B, Context Length=2K, Scale=180M, MIMO Rank=42026.03 | 22 | |
| Transformer-180MTraining Tokens=100B, Context Length=2K, Scale=180M2026.03 | 21.8 | |
| DualAlpha (α)=63/64, Data repetitions=12025.12 | 17.6 | |
| DualAlpha (α)=3/4, Data repetitions=322025.12 | 14.4 | |
| AutoregressiveAlpha (α)=1, Data repetitions=12025.12 | 13.6 | |
| AutoregressiveAlpha (α)=1, Data repetitions=322025.12 | 9.9 | |
| DualAlpha (α)=1/8, Data repetitions=1282025.12 | 8.5 | |
| GAINModel=Llama-13B2026.04 | 0.4 | |
| GAINModel=Mistral-7B2026.04 | 0 | |
| GAINModel=Llama-70B2026.04 | -0.2 | |
| AutoregressiveAlpha (α)=1, Data repetitions=1282025.12 | -0.5 | |
| GAINModel=Qwen2.5-7B2026.04 | -1 | |
| GAINModel=GPT-2 (774M)2026.04 | -1.8 | |
| LoRAModel=GPT-2 (774M)2026.04 | -2.2 | |
| LoRAModel=Qwen2.5-7B2026.04 | -2.8 | |
| LoRAModel=Mistral-7B2026.04 | -3 |