Question Answering on ARC-C (Accuracy)
0.71AccuracyLogicGaze
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| LogicGazeBackbone=LLaMA3-8B, Retrieval Setting=Yes, Evaluation Protocol=RAG2026.01 | 0.71 | — | |
| IterDRAGBackbone=LLaMA3-8B, Retrieval Setting=Yes, Evaluation Protocol=RAG2026.01 | 0.704 | — | |
| RankRAGBackbone=LLaMA3-8B, Retrieval Setting=Yes, Evaluation Protocol=RAG2026.01 | 0.696 | — | |
| AutoRAGBackbone=LLaMA3-8B, Retrieval Setting=Yes, Evaluation Protocol=RAG2026.01 | 0.681 | — | |
| RQ-RAGBackbone=LLaMA3-8B, Retrieval Setting=Yes, Evaluation Protocol=RAG2026.01 | 0.68 | — | |
| Self-RAGBackbone=LLaMA3-8B, Retrieval Setting=Yes, Evaluation Protocol=RAG2026.01 | 0.668 | — | |
| SFTBackbone=LLaMA3-8B, Retrieval Setting=No, Evaluation Protocol=SFT2026.01 | 0.624 | — | |
| Zero-Shot-InstructBackbone=LLaMA3-8B, Retrieval Setting=No, Evaluation Protocol=Zero-Shot-Instruct2026.01 | 0.589 | — | |
| BaselineModel=Qwen3-14B, Zero-shot protocol=true2025.12 | 0.5887 | — | |
| Zero-Shot-Instruct+RetBackbone=LLaMA3-8B, Retrieval Setting=Yes, Evaluation Protocol=Zero-Shot-Instruct2026.01 | 0.587 | — | |
| BitDelta (scalar)Model=Qwen3-14B, Zero-shot protocol=true2025.12 | 0.587 | — | |
| Vector (row/col)Model=Qwen3-14B, Zero-shot protocol=true2025.12 | 0.587 | — | |
| SFT+RetBackbone=LLaMA3-8B, Retrieval Setting=Yes, Evaluation Protocol=SFT2026.01 | 0.586 | — | |
| BaselineModel=Phi-4-reasoning, Zero-shot protocol=true2025.12 | 0.5572 | — | |
| Vector (row/col)Model=Phi-4-reasoning, Zero-shot protocol=true2025.12 | 0.5563 | — | |
| FP16Model=LLaMA3-8B-Instruct, Precision=MXFP42026.04 | 0.556 | — | |
| BitDelta (scalar)Model=Phi-4-reasoning, Zero-shot protocol=true2025.12 | 0.5546 | — | |
| FP16Model=LLaMA3.1-8B-Instruct, Precision=MXFP42026.04 | 0.552 | — | |
| Vector (row/col)Model=Llama-3.1-8B-Instruct, Zero-shot protocol=true2025.12 | 0.5358 | — | |
| Llama3 8BCompression Ratio=Baseline2025.09 | 0.535 | — | |
| DuQuant++Model=LLaMA3-8B-Instruct, Precision=MXFP42026.04 | 0.528 | — | |
| MR-GPTQModel=LLaMA3.1-8B-Instruct, Precision=MXFP42026.04 | 0.526 | — | |
| BitDelta (scalar)Model=Llama-3.1-8B-Instruct, Zero-shot protocol=true2025.12 | 0.5255 | — | |
| FlatQuantModel=LLaMA3.1-8B-Instruct, Precision=MXFP42026.04 | 0.521 | — | |
| DuQuant++Model=LLaMA3.1-8B-Instruct, Precision=MXFP42026.04 | 0.521 | — | |
| QuaRot*Model=LLaMA3.1-8B-Instruct, Precision=MXFP42026.04 | 0.519 | — | |
| BaselineModel=Llama-3.1-8B-Instruct, Zero-shot protocol=true2025.12 | 0.517 | — | |
| DuQuant++*Model=LLaMA3.1-8B-Instruct, Precision=MXFP42026.04 | 0.517 | — | |
| MR-GPTQModel=LLaMA3-8B-Instruct, Precision=MXFP42026.04 | 0.513 | — | |
| DuQuant++*Model=LLaMA3-8B-Instruct, Precision=MXFP42026.04 | 0.507 | — | |
| FlatQuantModel=LLaMA3-8B-Instruct, Precision=MXFP42026.04 | 0.494 | — | |
| SAIL-7BBackbone=SAIL-7B, Retrieval Setting=Yes, Evaluation Protocol=RAG2026.01 | 0.484 | — | |
| QuaRot*Model=LLaMA3-8B-Instruct, Precision=MXFP42026.04 | 0.482 | — | |
| QuaRotModel=LLaMA3.1-8B-Instruct, Precision=MXFP42026.04 | 0.463 | — | |
| QuaRotModel=LLaMA3-8B-Instruct, Precision=MXFP42026.04 | 0.459 | — | |
| ReplaceMeCompression Ratio=0.222025.09 | 0.437 | — | |
| CoSpaDiCompression Ratio=0.22025.09 | 0.416 | — | |
| w/o trainingQuantization=w/o Quant, Model=LLaMA2-7B2025.02 | 0.407 | — | |
| ByteFlow NetTokenizer=Byte, Parameter Scale=1.3B, Training Tokens=500B, Evaluation Protocol=Zero-shot2026.03 | 0.4036 | — | |
| IFDQuantization=GPTQ 4 bits, Model=LLaMA2-7B2025.02 | 0.4021 | — | |
| PASERQuantization=GPTQ 4 bits, Model=LLaMA2-7B2025.02 | 0.4018 | — | |
| PASERQuantization=RTN 4 bits, Model=LLaMA2-7B2025.02 | 0.3981 | — | |
| NuggetsQuantization=GPTQ 4 bits, Model=LLaMA2-7B2025.02 | 0.3956 | — | |
| NuggetsQuantization=RTN 4 bits, Model=LLaMA2-7B2025.02 | 0.3921 | — | |
| Instruction MiningQuantization=GPTQ 4 bits, Model=LLaMA2-7B2025.02 | 0.3893 | — | |
| IFDQuantization=RTN 4 bits, Model=LLaMA2-7B2025.02 | 0.3883 | — | |
| Full DataQuantization=GPTQ 4 bits, Model=LLaMA2-7B2025.02 | 0.3878 | — | |
| w/o trainingQuantization=GPTQ 4 bits, Model=LLaMA2-7B2025.02 | 0.3869 | — | |
| Full DataQuantization=RTN 4 bits, Model=LLaMA2-7B2025.02 | 0.3848 | — | |
| RandomQuantization=GPTQ 4 bits, Model=LLaMA2-7B2025.02 | 0.3822 | — | |
| Instruction MiningQuantization=RTN 4 bits, Model=LLaMA2-7B2025.02 | 0.3818 | — | |
| w/o trainingQuantization=RTN 4 bits, Model=LLaMA2-7B2025.02 | 0.3807 | — | |
| ReplaceMeCompression Ratio=0.312025.09 | 0.379 | — | |
| AU-NetTokenizer=Byte, Parameter Scale=1.3B, Training Tokens=500B, Evaluation Protocol=Zero-shot2026.03 | 0.3743 | — | |
| RandomQuantization=RTN 4 bits, Model=LLaMA2-7B2025.02 | 0.3713 | — | |
| LLaMATokenizer=BPE, Parameter Scale=1.3B, Training Tokens=500B, Evaluation Protocol=Zero-shot2026.03 | 0.3695 | — | |
| one-MoAModel Scale=2B, Setting=Zero-shot2026.05 | 0.3686 | — | |
| LLM-PrunerCompression Ratio=0.22025.09 | 0.366 | — | |
| qd-MoAModel Scale=2B, Setting=Zero-shot2026.05 | 0.3652 | — | |
| LlamaMoEModel Scale=2B, Setting=Zero-shot2026.05 | 0.365 | — | |
| MambaByteTokenizer=Byte, Parameter Scale=1.3B, Training Tokens=500B, Evaluation Protocol=Zero-shot2026.03 | 0.3642 | — | |
| LlamaByteTokenizer=Byte, Parameter Scale=1.3B, Training Tokens=500B, Evaluation Protocol=Zero-shot2026.03 | 0.3618 | — | |
| SpaceByteTokenizer=Byte, Parameter Scale=1.3B, Training Tokens=500B, Evaluation Protocol=Zero-shot2026.03 | 0.3605 | — | |
| +49-class Div.Class Divisions=49-class, Training Tokens=150B, Model Size=15B-A1.5B2026.02 | 0.3525 | — | |
| +3-class Div.Class Divisions=3-class, Training Tokens=150B, Model Size=15B-A1.5B2026.02 | 0.3514 | — | |
| Baseline MoEClass Divisions=None, Training Tokens=150B, Model Size=15B-A1.5B2026.02 | 0.3435 | — | |
| Qwen3-0.6BCompression Ratio=None2026.02 | 0.343 | — | |
| Bloom-7b1Ratio=0%, Evaluation Protocol=Zero-shot2024.03 | 0.3353 | — | |
| CoSpaDiCompression Ratio=0.32025.09 | 0.335 | — | |
| AIRAhybrid-DShot=0-shot, Scale=1B, Token Budget=37.5B tokens, Architecture Variant=Stretched2026.05 | 0.32 | 0.329 | |
| HyWIARatio=25%, Evaluation Protocol=Zero-shot2024.03 | 0.3114 | — | |
| LLM-Pruner Vector⋆Ratio=25%, Evaluation Protocol=Zero-shot2024.03 | 0.3089 | — | |
| Nemotron-2 Approx.Shot=0-shot, Scale=1B, Token Budget=37.5B tokens2026.05 | 0.307 | 0.332 | |
| AIRAhybrid-AShot=0-shot, Scale=1B, Token Budget=37.5B tokens, Architecture Variant=Stretched2026.05 | 0.305 | 0.335 | |
| AIRAhybrid-CShot=0-shot, Scale=1B, Token Budget=37.5B tokens, Architecture Variant=Stretched2026.05 | 0.302 | 0.322 | |
| LLM-Pruner Element²⋆Ratio=25%, Evaluation Protocol=Zero-shot2024.03 | 0.3012 | — | |
| Mamba (Mb + M)Shot=0-shot, Scale=1B, Token Budget=37.5B tokens2026.05 | 0.301 | 0.325 | |
| Composer (2Mb-M-3A)Shot=0-shot, Scale=1B, Token Budget=37.5B tokens2026.05 | 0.3 | 0.331 | |
| AIRAhybrid-BShot=0-shot, Scale=1B, Token Budget=37.5B tokens, Architecture Variant=Stretched2026.05 | 0.3 | 0.311 | |
| AIRAhybrid-BShot=0-shot, Scale=1B, Token Budget=37.5B tokens, Architecture Variant=Stacked2026.05 | 0.299 | 0.334 | |
| AIRAhybrid-EShot=0-shot, Scale=1B, Token Budget=37.5B tokens, Architecture Variant=Stretched2026.05 | 0.298 | 0.325 | |
| COMPOT†Compression Ratio=0.22026.02 | 0.296 | — | |
| AIRAhybrid-DShot=0-shot, Scale=1B, Token Budget=37.5B tokens, Architecture Variant=Stacked2026.05 | 0.295 | 0.309 | |
| Nemotron-H Approx.Shot=0-shot, Scale=1B, Token Budget=37.5B tokens2026.05 | 0.291 | 0.313 | |
| Zero-ShotBackbone=LLaMA3-8B, Retrieval Setting=No, Evaluation Protocol=Zero-Shot2026.01 | 0.289 | — | |
| LLM-PrunerCompression Ratio=0.32025.09 | 0.288 | — | |
| COMPOTCompression Ratio=0.22026.02 | 0.288 | — | |
| ByteFlow NetTokenizer=Byte, Parameter Scale=600M, Training Tokens=50B, Evaluation Protocol=Zero-shot2026.03 | 0.2836 | — | |
| Zero-Shot+RetBackbone=LLaMA3-8B, Retrieval Setting=Yes, Evaluation Protocol=Zero-Shot2026.01 | 0.28 | — | |
| AIRAhybrid-EShot=0-shot, Scale=1B, Token Budget=37.5B tokens, Architecture Variant=Stacked2026.05 | 0.277 | 0.303 | |
| ReplaceMeCompression Ratio=0.412025.09 | 0.275 | — | |
| AU-NetTokenizer=Byte, Parameter Scale=600M, Training Tokens=50B, Evaluation Protocol=Zero-shot2026.03 | 0.2743 | — | |
| CoSpaDiCompression Ratio=0.22026.02 | 0.271 | — | |
| COMPOT†Compression Ratio=0.32026.02 | 0.271 | — | |
| COMPOTCompression Ratio=0.32026.02 | 0.271 | — | |
| CoSpaDiCompression Ratio=0.42025.09 | 0.266 | — | |
| LLaMATokenizer=BPE, Parameter Scale=600M, Training Tokens=50B, Evaluation Protocol=Zero-shot2026.03 | 0.2595 | — | |
| LLM-PrunerCompression Ratio=0.42025.09 | 0.258 | — | |
| MambaByteTokenizer=Byte, Parameter Scale=600M, Training Tokens=50B, Evaluation Protocol=Zero-shot2026.03 | 0.2542 | — | |
| LlamaByteTokenizer=Byte, Parameter Scale=600M, Training Tokens=50B, Evaluation Protocol=Zero-shot2026.03 | 0.2518 | — |