Multiple-choice Question Answering on ARC Challenge
74.7AccLlama3.1-8B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Llama3.1-8BBase Model=Llama3.1-8B, Evaluation Protocol=zero-shot2026.02 | 74.7 | — | |
| Llama3.1-8B + ODESTEERBase Model=Llama3.1-8B, Steering Method=ODESTEER, Evaluation Protocol=zero-shot2026.02 | 74.5 | — | |
| Llama-3.1-8B-InstructSetting=Teacher2026.02 | 71.67 | — | |
| FLAN-T5 xlargeSetting=Teacher2026.02 | 68.24 | — | |
| Llama-2-13b-chat (OTTER)Backbone=Llama-2-13b-chat, Evaluation protocol=Zero-shot, Variant=OTTER2024.04 | 66.2 | 2.5 | |
| UnifiedQA-3bparameters=3B, Protocol=Transfer-learning, External Knowledge=No2023.07 | 64.5 | — | |
| Llama-2-13b-chat (Naive)Backbone=Llama-2-13b-chat, Evaluation protocol=Zero-shot, Variant=Naive2024.04 | 64.4 | 13.7 | |
| UnifiedQA-3bparameters=3B, Protocol=Transfer-learning, External Knowledge=Yes2023.07 | 64.2 | — | |
| Vicuna-13b (OTTER)Backbone=Vicuna-13b, Evaluation protocol=Zero-shot, Variant=OTTER2024.04 | 63.3 | 2.4 | |
| Llama-2-13b (Naive)Backbone=Llama-2-13b, Evaluation protocol=Zero-shot, Variant=Naive2024.04 | 62.9 | 6 | |
| Vicuna-13b (Naive)Backbone=Vicuna-13b, Evaluation protocol=Zero-shot, Variant=Naive2024.04 | 62.9 | 8.3 | |
| Llama-2-13b (OTTER)Backbone=Llama-2-13b, Evaluation protocol=Zero-shot, Variant=OTTER2024.04 | 62.8 | 1.5 | |
| Teacher SelectionSetting=Student, Backbone=FLAN-T5 large (783M)2026.02 | 61.12 | — | |
| Similarity-based RouterSetting=Student, Backbone=FLAN-T5 large (783M)2026.02 | 60.6 | — | |
| Fine-tuningSetting=Student, Backbone=FLAN-T5 large (783M)2026.02 | 60.26 | — | |
| Distilling-Step-by-StepSetting=Student, Backbone=FLAN-T5 large (783M)2026.02 | 60 | — | |
| TinyLLMSetting=Student, Backbone=FLAN-T5 large (783M)2026.02 | 59.83 | — | |
| Knowledge AggregationSetting=Student, Backbone=FLAN-T5 large (783M)2026.02 | 59.74 | — | |
| PLM ClassifierSetting=Student, Backbone=FLAN-T5 large (783M)2026.02 | 59.66 | — | |
| Plackett-Luce RankingSetting=Student, Backbone=FLAN-T5 large (783M)2026.02 | 59.48 | — | |
| Llama-2-7b-chat (OTTER)Backbone=Llama-2-7b-chat, Evaluation protocol=Zero-shot, Variant=OTTER2024.04 | 57.4 | 1.3 | |
| Llama-2-7b-chat (Naive)Backbone=Llama-2-7b-chat, Evaluation protocol=Zero-shot, Variant=Naive2024.04 | 56.5 | 12.4 | |
| UnifiedQA-largeparameters=770M, Protocol=Transfer-learning, External Knowledge=No2023.07 | 55.2 | — | |
| BF16Model Size=8B2026.03 | 54.96 | — | |
| Vicuna-7b (OTTER)Backbone=Vicuna-7b, Evaluation protocol=Zero-shot, Variant=OTTER2024.04 | 54.1 | 2.3 | |
| FAAR+2FAModel Size=8B2026.03 | 53.95 | — | |
| GPTQModel Size=8B2026.03 | 53.81 | — | |
| MR-GPTQModel Size=8B2026.03 | 53.7 | — | |
| OriginalBackbone=Mistral-7B, Pruning Ratio=0%2026.05 | 53.67 | — | |
| Vicuna-7b (Naive)Backbone=Vicuna-7b, Evaluation protocol=Zero-shot, Variant=Naive2024.04 | 53.5 | 8.6 | |
| GPTQ+4/6Model Size=8B2026.03 | 53.16 | — | |
| RTNModel Size=8B2026.03 | 52.49 | — | |
| BioMistral-7BSetting=Teacher2026.02 | 51.59 | — | |
| InferenceSetting=Student, Backbone=FLAN-T5 large (783M)2026.02 | 51.07 | — | |
| UnifiedQA-largeparameters=770M, Protocol=Transfer-learning, External Knowledge=Yes2023.07 | 49.5 | — | |
| FKL*Vocab size=64k, Number of shots=4-shot, Combined with SFT=true2025.12 | 46.9 | — | |
| Teacher SelectionSetting=Student, Backbone=FLAN-T5 base (248M)2026.02 | 46.7 | — | |
| SFTVocab size=64k, Number of shots=4-shot2025.12 | 46.5 | — | |
| FKLVocab size=64k, Number of shots=4-shot2025.12 | 46.5 | — | |
| AccBackbone=Mistral-7B, Pruning Ratio=25%, Relevance Estimation Source=Task-specific2026.05 | 46.42 | — | |
| ALMVocab size=64k, Number of shots=4-shot2025.12 | 46.4 | — | |
| PLM ClassifierSetting=Student, Backbone=FLAN-T5 base (248M)2026.02 | 46.35 | — | |
| Similarity-based RouterSetting=Student, Backbone=FLAN-T5 base (248M)2026.02 | 46.35 | — | |
| ALM*Vocab size=64k, Number of shots=4-shot, Combined with SFT=true2025.12 | 46.3 | — | |
| DSKD*Vocab size=64k, Number of shots=4-shot, Combined with SFT=true2025.12 | 46.3 | — | |
| Plackett-Luce RankingSetting=Student, Backbone=FLAN-T5 base (248M)2026.02 | 46.27 | — | |
| ULD*Vocab size=64k, Number of shots=4-shot, Combined with SFT=true2025.12 | 46.2 | — | |
| TinyLLMSetting=Student, Backbone=FLAN-T5 base (248M)2026.02 | 46.18 | — | |
| Distilling-Step-by-StepSetting=Student, Backbone=FLAN-T5 base (248M)2026.02 | 46.01 | — | |
| DSKDVocab size=64k, Number of shots=4-shot2025.12 | 45.9 | — | |
| Knowledge AggregationSetting=Student, Backbone=FLAN-T5 base (248M)2026.02 | 45.75 | — | |
| Llama-2-7b (OTTER)Backbone=Llama-2-7b, Evaluation protocol=Zero-shot, Variant=OTTER2024.04 | 45.5 | 1.7 | |
| UnifiedQA-baseparameters=220M, Protocol=Transfer-learning, External Knowledge=Yes2023.07 | 45.2 | — | |
| ULDVocab size=64k, Number of shots=4-shot2025.12 | 45.1 | — | |
| UnifiedQA-baseparameters=220M, Protocol=Transfer-learning, External Knowledge=No2023.07 | 44.8 | — | |
| FKL*Vocab size=32k, Number of shots=4-shot, Combined with SFT=true2025.12 | 44.4 | — | |
| ULD*Vocab size=32k, Number of shots=4-shot, Combined with SFT=true2025.12 | 44.2 | — | |
| DSKD*Vocab size=32k, Number of shots=4-shot, Combined with SFT=true2025.12 | 44 | — | |
| ALM*Vocab size=32k, Number of shots=4-shot, Combined with SFT=true2025.12 | 43.9 | — | |
| SFTVocab size=32k, Number of shots=4-shot2025.12 | 43.7 | — | |
| Fine-tuningSetting=Student, Backbone=FLAN-T5 base (248M)2026.02 | 43.61 | — | |
| Llama 2-chatSetting=Teacher2026.02 | 43.35 | — | |
| ALMVocab size=32k, Number of shots=4-shot2025.12 | 43.3 | — | |
| DSKDVocab size=32k, Number of shots=4-shot2025.12 | 43.3 | — | |
| FKLVocab size=32k, Number of shots=4-shot2025.12 | 43.2 | — | |
| Cosine SimilarityBackbone=Mistral-7B, Pruning Ratio=25%, Relevance Estimation Source=Task-specific2026.05 | 42.41 | — | |
| ULDVocab size=32k, Number of shots=4-shot2025.12 | 42.4 | — | |
| FKLVocab size=16k, Number of shots=4-shot2025.12 | 42.4 | — | |
| FKL*Vocab size=16k, Number of shots=4-shot, Combined with SFT=true2025.12 | 42 | — | |
| ALM*Vocab size=16k, Number of shots=4-shot, Combined with SFT=true2025.12 | 41.8 | — | |
| ULD*Vocab size=16k, Number of shots=4-shot, Combined with SFT=true2025.12 | 41.8 | — | |
| ALMVocab size=16k, Number of shots=4-shot2025.12 | 41.4 | — | |
| DSKD*Vocab size=16k, Number of shots=4-shot, Combined with SFT=true2025.12 | 41.2 | — | |
| PerplexityBackbone=Mistral-7B, Pruning Ratio=25%, Relevance Estimation Source=Task-specific2026.05 | 40.96 | — | |
| DSKDVocab size=16k, Number of shots=4-shot2025.12 | 40.9 | — | |
| SFTVocab size=16k, Number of shots=4-shot2025.12 | 40.1 | — | |
| ULDVocab size=16k, Number of shots=4-shot2025.12 | 40.1 | — | |
| NextLat (d=2)Horizon (d)=2, Parameters=1.3B, Training tokens=100B2025.11 | 40.1 | — | |
| NextLat (d=1)Horizon (d)=1, Parameters=1.3B, Training tokens=100B2025.11 | 39.68 | — | |
| JTP (d=1)Horizon (d)=1, Parameters=1.3B, Training tokens=100B2025.11 | 39.25 | — | |
| GPTParameters=1.3B, Training tokens=100B2025.11 | 39.16 | — | |
| MTP (d=2)Horizon (d)=2, Parameters=1.3B, Training tokens=100B2025.11 | 39.08 | — | |
| MTP (d=1)Horizon (d)=1, Parameters=1.3B, Training tokens=100B2025.11 | 38.91 | — | |
| JTP (d=2)Horizon (d)=2, Parameters=1.3B, Training tokens=100B2025.11 | 38.57 | — | |
| Out. Cosine-SimBackbone=Mistral-7B, Pruning Ratio=25%, Relevance Estimation Source=Task-specific2026.05 | 38.41 | — | |
| Out. Norm-SimBackbone=Mistral-7B, Pruning Ratio=25%, Relevance Estimation Source=Task-specific2026.05 | 38.41 | — | |
| BF16Model Size=1B2026.03 | 36.86 | — | |
| FAAR+2FAModel Size=1B2026.03 | 36.09 | — | |
| Llama-2-7b (Naive)Backbone=Llama-2-7b, Evaluation protocol=Zero-shot, Variant=Naive2024.04 | 36 | 27.4 | |
| NuResFormerAnchor Source=Internal Anchor, Mixing Granularity=Headwise, Normalization Strategy=Only Q/K Norms, Mixing Path Components=All, Evaluation Protocol=Zero-shot2026.01 | 34.73 | — | |
| RTNModel Size=1B2026.03 | 34.28 | — | |
| GPTQ+4/6Model Size=1B2026.03 | 33.81 | — | |
| NuResFormerAnchor Source=Internal Anchor, Mixing Granularity=Headwise, Normalization Strategy=Full, Mixing Path Components=All, Evaluation Protocol=Zero-shot2026.01 | 33.79 | — | |
| NuResFormerAnchor Source=Internal Anchor, Mixing Granularity=Elementwise, Normalization Strategy=Full, Mixing Path Components=All, Evaluation Protocol=Zero-shot2026.01 | 33.7 | — | |
| ResFormer (Value Residual)Anchor Source=Baselines, Evaluation Protocol=Zero-shot2026.01 | 33.62 | — | |
| MR-GPTQModel Size=1B2026.03 | 33.47 | — | |
| Dynamic ExoFormerAnchor Source=External Anchor, Mixing Granularity=Elementwise, Normalization Strategy=Full, Evaluation Protocol=Zero-shot, Mixing Logic=Dynamic2026.01 | 33.36 | — | |
| GPTQModel Size=1B2026.03 | 33.07 | — | |
| Naïve CombinationAnchor Source=Baselines, Evaluation Protocol=Zero-shot2026.01 | 32.94 | — | |
| NuResFormerAnchor Source=Internal Anchor, Mixing Granularity=Scalar, Normalization Strategy=Only Q/K Norms, Mixing Path Components=All, Evaluation Protocol=Zero-shot2026.01 | 32.85 | — |