Question Answering on ARC-C
94.1AccuracyDRAG
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| DRAGBackbone LLM=Phi-3.5-mini-instruct2025.06 | 94.1 | — | |
| DRAGBackbone LLM=Llama-3.1-8B-Instruct2025.06 | 93.1 | — | |
| DRAGBackbone LLM=Qwen2.5-3B-Instruct2025.06 | 93 | — | |
| DRAGBackbone LLM=Llama-3.2-3B-Instruct2025.06 | 93 | — | |
| DRAGBackbone LLM=Gemma-2-2B-it2025.06 | 91.5 | — | |
| Qwen3-8B + SFTshot=0-shot, mask rate=0.1, base model=Qwen3-8B2026.05 | 91.3 | — | |
| Qwen3-8B + SFT + WeMask(TF)shot=0-shot, mask rate=0.1, base model=Qwen3-8B2026.05 | 91.3 | — | |
| Qwen3-8B + WeMask(SFT)shot=0-shot, mask rate=0.1, base model=Qwen3-8B2026.05 | 91.3 | — | |
| DRAGBackbone LLM=GLM-edge-1.5B-chat2025.06 | 90 | — | |
| GPT-3.5Size=>20B2024.03 | 88.81 | — | |
| FPModel=LLaDA-Instruct, Quantization=Full Precision2026.06 | 88.47 | — | |
| Qwen3-4B + SFT + WeMask(TF)shot=0-shot, mask rate=0.1, base model=Qwen3-4B2026.05 | 87.54 | — | |
| FAIR-CalibModel=LLaDA-Instruct, Quantization=W4A42026.06 | 87.46 | — | |
| FlatQuantModel=LLaDA-Instruct, Quantization=W4A42026.06 | 87.31 | — | |
| FPModel=LLaDA-1.5, Quantization=Full Precision2026.06 | 87.12 | — | |
| Qwen3-4B + WeMask(SFT)shot=0-shot, mask rate=0.1, base model=Qwen3-4B2026.05 | 87.03 | — | |
| Qwen3-8B + Gated Attentionshot=0-shot, mask rate=0.1, base model=Qwen3-8B2026.05 | 86.95 | — | |
| Qwen3-4B + SFTshot=0-shot, mask rate=0.1, base model=Qwen3-4B2026.05 | 86.69 | — | |
| C2DLM2025.11 | 86.26 | — | |
| DRAGBackbone LLM=LLaMA-2-7B2025.06 | 86.2 | — | |
| FAIR-CalibModel=LLaDA-1.5, Quantization=W4A42026.06 | 86.1 | — | |
| QuaRotModel=LLaDA-1.5, Quantization=W4A42026.06 | 85.82 | — | |
| LLaDA-8B-InstructSetting=Direct2025.11 | 85.75 | — | |
| LLaDA-8B-InstructSetting=SFT2025.11 | 85.67 | — | |
| QuaRotModel=LLaDA-Instruct, Quantization=W4A42026.06 | 85.42 | — | |
| FlatQuantModel=LLaDA-1.5, Quantization=W4A42026.06 | 85.12 | — | |
| Qwen3-4B + Gated Attentionshot=0-shot, mask rate=0.1, base model=Qwen3-4B2026.05 | 84.81 | — | |
| MiniRAGBackbone LLM=Phi-3.5-mini-instruct2025.06 | 82.7 | — | |
| PaLM-2Size=>20B2024.03 | 82.37 | — | |
| RTNModel=LLaDA-1.5, Quantization=W4A42026.06 | 82.37 | — | |
| MeanLearnSize=13B2024.03 | 82 | — | |
| SimRAGBackbone LLM=Llama-3.1-8B-Instruct2025.06 | 81.4 | — | |
| MixLoRABase Model=LLaMA-3 8B, Trainable Parameters=3.0%2024.04 | 79.9 | — | |
| LoRA-MixerBackbone=LLaMA3-8B, Adapter=Q/K/V/O, Routing Strategy=RSL2025.06 | 79.89 | — | |
| RTNModel=LLaDA-Instruct, Quantization=W4A42026.06 | 79.66 | — | |
| MixDoRABase Model=LLaMA-3 8B, Trainable Parameters=3.0%2024.04 | 78.9 | — | |
| BaseBackbone=LLaMA3-8B, Adapter=Q/K/V/O2025.06 | 78.65 | — | |
| MeanLearnSize=7B2024.03 | 78.06 | — | |
| MeanLearnSize=8B2024.03 | 77.97 | — | |
| DoRABase Model=LLaMA-3 8B, Trainable Parameters=2.6%2024.04 | 76.4 | — | |
| PhatgooseBackbone=LLaMA3-8B, Adapter=Q/K/V/O, Routing Strategy=sigmoid gating based on cosine similarity2025.06 | 75.94 | — | |
| LoRABase Model=LLaMA-3 8B, Trainable Parameters=2.6%2024.04 | 75.7 | — | |
| LLaMA-3Size=8B2024.03 | 75.25 | — | |
| Orca-2Size=7B2024.03 | 73.9 | — | |
| Orca-2Size=13B2024.03 | 70.85 | — | |
| MixLoRABase Model=LLaMA-2 13B, Trainable Parameters=2.5%2024.04 | 69.9 | — | |
| PaLM 2-Lprompting=1-shot2023.05 | 69.2 | — | |
| CRAGBackbone LLM=SelfRAG-LLaMA-2-7B2025.06 | 68.6 | — | |
| MiniRAGBackbone LLM=Gemma-2-2B-it2025.06 | 68.6 | — | |
| MixDoRABase Model=LLaMA-2 13B, Trainable Parameters=2.5%2024.04 | 68.4 | — | |
| DoRABase Model=LLaMA-2 13B, Trainable Parameters=2.4%2024.04 | 67.7 | — | |
| MiniRAGBackbone LLM=Qwen2.5-3B-Instruct2025.06 | 67.7 | — | |
| LoRABase Model=LLaMA-2 13B, Trainable Parameters=2.4%2024.04 | 67.6 | — | |
| Self-RAGBackbone LLM=SelfRAG-LLaMA-2-7B2025.06 | 67.3 | — | |
| MiniRAGBackbone LLM=Llama-3.2-3B-Instruct2025.06 | 65.3 | — | |
| PaLM 2-Mprompting=1-shot2023.05 | 64.9 | — | |
| MiniRAGBackbone LLM=GLM-edge-1.5B-chat2025.06 | 62.3 | — | |
| FPModel=Dream-Instruct, Quantization=Full Precision2026.06 | 61.43 | 71.01 | |
| Gemma-7BParameters=7B2024.07 | 61.1 | — | |
| Qwen2-7BParameters=7B2024.07 | 60.6 | — | |
| PaLMprompting=1-shot2023.05 | 60.1 | — | |
| Mistral-7BParameters=7B2024.07 | 60 | — | |
| DoRABase Model=LLaMA-2 7B, Trainable Parameters=2.9%2024.04 | 59.8 | — | |
| VicunaSize=13B2024.03 | 59.77 | — | |
| WizardLMSize=13B2024.03 | 59.77 | — | |
| PaLM 2-Sprompting=1-shot2023.05 | 59.6 | — | |
| Llama-3-8BParameters=8B2024.07 | 59.3 | — | |
| FPModel=Dream-Base, Quantization=Full Precision2026.06 | 59.13 | 70.01 | |
| MixDoRABase Model=LLaMA-2 7B, Trainable Parameters=2.9%2024.04 | 58.2 | — | |
| MixLoRABase Model=LLaMA-2 7B, Trainable Parameters=2.9%2024.04 | 58.1 | — | |
| FAIR-CalibModel=Dream-Instruct, Quantization=W4A42026.06 | 58.02 | 66.66 | |
| VicunaSize=7B2024.03 | 57.63 | — | |
| FlatQuantModel=Dream-Instruct, Quantization=W4A42026.06 | 55.29 | 63.98 | |
| LLaMA-2Size=13B2024.03 | 55.25 | — | |
| DenseBackbone=LLaMA-3.1-8B, Pruning Method=None, Evaluation Protocol=Zero-shot, FFN Sparsity=0%2026.01 | 54.95 | — | |
| MixDoRABase Model=Gemma 2B, Trainable Parameters=4.3%2024.04 | 54.3 | — | |
| Qwen1.5-7BParameters=7B2024.07 | 54.2 | — | |
| FAIR-CalibModel=Dream-Base, Quantization=W4A42026.06 | 53.92 | 64.64 | |
| LLaMA2-7B# tokens for training=2T, Shot number=25, Training data source=Different from RedPajama2023.10 | 53 | — | |
| ZEPHYR-7B-beta2024.05 | 52.5 | — | |
| QuaRotModel=Dream-Instruct, Quantization=W4A42026.06 | 52.3 | 61.14 | |
| FP16Model=LLaMA-30B, Precision=W4A16, Zero-shot=true2026.05 | 51.54 | — | |
| FP16Model=LLaMA-30B, Precision=W3A16, Zero-shot=true2026.05 | 51.54 | — | |
| CATBackbone=MISTRAL-7B2024.05 | 51.5 | — | |
| SARQC-GS(Saliency)Model=LLaMA-30B, Precision=W4A16, Zero-shot=true2026.05 | 51.02 | — | |
| AWQModel=LLaMA-30B, Precision=W4A16, Zero-shot=true2026.05 | 50.94 | — | |
| SARQC-GS(Identity)Model=LLaMA-30B, Precision=W4A16, Zero-shot=true2026.05 | 50.94 | — | |
| LoRABase Model=LLaMA-2 7B, Trainable Parameters=2.9%2024.04 | 50.9 | — | |
| MISTRAL-7B2024.05 | 50.8 | — | |
| FlatQuantModel=Dream-Base, Quantization=W4A42026.06 | 50.51 | 62.08 | |
| PHI-3-MINI2024.05 | 50.5 | — | |
| SARQC-GBS(Saliency)Model=LLaMA-30B, Precision=W4A16, Zero-shot=true2026.05 | 50.26 | — | |
| GPTAQModel=LLaMA-30B, Precision=W4A16, Zero-shot=true2026.05 | 50.17 | — | |
| GradOTSetting=Plug-in, Backbone=LLaMA-13B2025.07 | 50.1 | — | |
| GPTQModel=LLaMA-30B, Precision=W4A16, Zero-shot=true2026.05 | 50.09 | — | |
| SARQC-GBS(Identity)Model=LLaMA-30B, Precision=W4A16, Zero-shot=true2026.05 | 50.09 | — | |
| No pruning (Original)Backbone=Gemma-2 2B, Sparsity=0, Evaluation Protocol=Zero-shot2026.03 | 49.7 | — | |
| Llama-2-13B#Bit=16, Base Model=Llama-2-13B, Zero-shot=true2025.03 | 49.2 | — | |
| QuaRotModel=Dream-Base, Quantization=W4A42026.06 | 48.72 | 57.49 | |
| CAPOBackbone=ZEPHYR-7B-beta2024.05 | 48.5 | — |