Question Answering on ARC-E (0-shot)
87.88AccuracyMoving Average Regularized Looping
Evaluation Results
| Method | Links | |
|---|---|---|
| Moving Average Regularized LoopingModel=Gemma, Size=9B, Shots=0-shot2026.02 | 87.88 | |
| BaselineModel=Gemma, Size=9B, Shots=0-shot2026.02 | 87.84 | |
| Uniform Regularized LoopingModel=Gemma, Size=9B, Shots=0-shot2026.02 | 87.84 | |
| Auto-Align Regularized LoopingModel=Gemma, Size=9B, Shots=0-shot2026.02 | 87.84 | |
| Noise AblationModel=Gemma, Size=9B, Shots=0-shot2026.02 | 87.84 | |
| FP16Backbone=LLaMA-3.1-8B, Quantization=None, Protocol=0-shot2026.02 | 82.4 | |
| TRLModel Backbone=Qwen2.5-7B, Evaluation Protocol=0-shot2025.09 | 81.57 | |
| OPPOModel Backbone=Qwen2.5-7B, Evaluation Protocol=0-shot2025.09 | 81.36 | |
| BaselineModel=Gemma, Size=2B, Shots=0-shot2026.02 | 80.3 | |
| Auto-Align Regularized LoopingModel=Gemma, Size=2B, Shots=0-shot2026.02 | 80.22 | |
| Moving Average Regularized LoopingModel=Gemma, Size=2B, Shots=0-shot2026.02 | 80.18 | |
| Uniform Regularized LoopingModel=Gemma, Size=2B, Shots=0-shot2026.02 | 80.01 | |
| Noise AblationModel=Gemma, Size=2B, Shots=0-shot2026.02 | 79.84 | |
| FP16q (mean bit precision)=16.000, tok/s (throughput)=103.8, Evaluation Protocol=0-shot2025.12 | 79.63 | |
| BaselineModel=Llama, Size=8B, Shots=0-shot2026.02 | 77.69 | |
| Auto-Align Regularized LoopingModel=Llama, Size=8B, Shots=0-shot2026.02 | 77.61 | |
| Noise AblationModel=Llama, Size=8B, Shots=0-shot2026.02 | 77.57 | |
| Moving Average Regularized LoopingModel=Llama, Size=8B, Shots=0-shot2026.02 | 77.4 | |
| Uniform Regularized LoopingModel=Llama, Size=8B, Shots=0-shot2026.02 | 77.31 | |
| AstroBackbone=LLaMA-3.1-8B, Quantization=W3A16g128, Protocol=0-shot2026.02 | 77.1 | |
| OPPOModel Backbone=Qwen2.5-3B, Evaluation Protocol=0-shot2025.09 | 75.08 | |
| AQLM-1x16 + PV-Tuningq (mean bit precision)=2.213, tok/s (throughput)=49.0, PV-Tuning=true, Evaluation Protocol=0-shot2025.12 | 74.92 | |
| TRLModel Backbone=Qwen2.5-3B, Evaluation Protocol=0-shot2025.09 | 74.54 | |
| AQLM-1x16q (mean bit precision)=2.213, tok/s (throughput)=49.0, Evaluation Protocol=0-shot2025.12 | 73.99 | |
| CodeGEMM-m2v8g128 + PV-Tuningq (mean bit precision)=2.127, tok/s (throughput)=214.4, PV-Tuning=true, Evaluation Protocol=0-shot2025.12 | 73.91 | |
| MagRBackbone=LLaMA-3.1-8B, Quantization=W3A16g128, Protocol=0-shot2026.02 | 73.6 | |
| CodeGEMM-m1v4g128 + PV-Tuningq (mean bit precision)=2.126, tok/s (throughput)=228.3, PV-Tuning=true, Evaluation Protocol=0-shot2025.12 | 73.15 | |
| OmniQuantBackbone=LLaMA-3.1-8B, Quantization=W3A16g128, Protocol=0-shot2026.02 | 71.6 | |
| AQLM-2x8 + PV-Tuningq (mean bit precision)=2.005, tok/s (throughput)=124.5, PV-Tuning=true, Evaluation Protocol=0-shot2025.12 | 71.25 | |
| CodeGEMM-m1v4g128q (mean bit precision)=2.126, tok/s (throughput)=228.3, Evaluation Protocol=0-shot2025.12 | 63.97 | |
| CodeGEMM-m2v8g128q (mean bit precision)=2.127, tok/s (throughput)=214.4, Evaluation Protocol=0-shot2025.12 | 60.14 | |
| FP16Bits=–, Backbone=LLAMA-2-7B2026.04 | 53.58 | |
| FBI-LLMBits=W1A16, Backbone=LLAMA-2-7B2026.04 | 53 | |
| AQLM-2x8q (mean bit precision)=2.005, tok/s (throughput)=124.5, Evaluation Protocol=0-shot2025.12 | 46.25 | |
| LBLLMBits=W(1+1)A4, Backbone=LLAMA-2-7B2026.04 | 44.65 | |
| OnebitBits=W1A16, Backbone=LLAMA-2-7B2026.04 | 42.47 | |
| FlexRound-q2g128q (mean bit precision)=2.125, tok/s (throughput)=205.3, Evaluation Protocol=0-shot2025.12 | 24.75 |