Question Answering on ARC-Challenge 0-shot (test)
90.4AccuracyMIPO
Evaluation Results
| Method | Links | |
|---|---|---|
| MIPOBackbone=Qwen2.5-7B-Instruct, Training=MIPO, Few-shot=zero-shot2026.03 | 90.4 | |
| SFTBackbone=Qwen2.5-7B-Instruct, Training=SFT, Ground-truth usage=true, Few-shot=zero-shot2026.03 | 90 | |
| RLVRBackbone=Qwen2.5-7B-Instruct, Training=RLVR, Ground-truth usage=true, Few-shot=zero-shot2026.03 | 88.6 | |
| Qwen2.5-7B-InstructBackbone=Qwen2.5-7B-Instruct, Few-shot=zero-shot2026.03 | 88.2 | |
| SFTBackbone=Qwen2.5-7B-Instruct, Training=SFT, Few-shot=zero-shot2026.03 | 87.8 | |
| MIPOBackbone=Qwen2.5-3B-Instruct, Training=MIPO, Few-shot=zero-shot2026.03 | 80.13 | |
| SFTBackbone=Qwen2.5-3B-Instruct, Training=SFT, Ground-truth usage=true, Few-shot=zero-shot2026.03 | 79.6 | |
| Qwen2.5-3B-InstructBackbone=Qwen2.5-3B-Instruct, Few-shot=zero-shot2026.03 | 79.4 | |
| RLVRBackbone=Qwen2.5-3B-Instruct, Training=RLVR, Ground-truth usage=true, Few-shot=zero-shot2026.03 | 79.4 | |
| SFTBackbone=Qwen2.5-3B-Instruct, Training=SFT, Few-shot=zero-shot2026.03 | 78.67 | |
| MIPOBackbone=Llama-3.2-3B-Instruct, Training=MIPO, Few-shot=zero-shot2026.03 | 70.93 | |
| SFTBackbone=Llama-3.2-3B-Instruct, Training=SFT, Ground-truth usage=true, Few-shot=zero-shot2026.03 | 70.53 | |
| RLVRBackbone=Qwen2.5-1.5B-Instruct, Training=RLVR, Ground-truth usage=true, Few-shot=zero-shot2026.03 | 70.47 | |
| RLVRBackbone=Llama-3.2-3B-Instruct, Training=RLVR, Ground-truth usage=true, Few-shot=zero-shot2026.03 | 70.33 | |
| Llama-3.2-3B-InstructBackbone=Llama-3.2-3B-Instruct, Few-shot=zero-shot2026.03 | 68.6 | |
| MIPOBackbone=Qwen2.5-1.5B-Instruct, Training=MIPO, Few-shot=zero-shot2026.03 | 65.93 | |
| SFTBackbone=Llama-3.2-3B-Instruct, Training=SFT, Few-shot=zero-shot2026.03 | 63.67 | |
| Qwen2.5-1.5B-InstructBackbone=Qwen2.5-1.5B-Instruct, Few-shot=zero-shot2026.03 | 63.6 | |
| SFTBackbone=Qwen2.5-1.5B-Instruct, Training=SFT, Ground-truth usage=true, Few-shot=zero-shot2026.03 | 62.2 | |
| OPPOModel Backbone=Qwen2.5-7B, Evaluation Protocol=0-shot2025.09 | 55.8 | |
| TRLModel Backbone=Qwen2.5-7B, Evaluation Protocol=0-shot2025.09 | 55.55 | |
| FP16q (mean bit precision)=16.000, tok/s (throughput)=103.8, Evaluation Protocol=0-shot2025.12 | 55.03 | |
| SFTBackbone=Qwen2.5-1.5B-Instruct, Training=SFT, Few-shot=zero-shot2026.03 | 52.47 | |
| OPPOModel Backbone=Qwen2.5-3B, Evaluation Protocol=0-shot2025.09 | 49.57 | |
| AQLM-1x16 + PV-Tuningq (mean bit precision)=2.213, tok/s (throughput)=49.0, PV-Tuning=true, Evaluation Protocol=0-shot2025.12 | 48.89 | |
| TRLModel Backbone=Qwen2.5-3B, Evaluation Protocol=0-shot2025.09 | 48.89 | |
| CodeGEMM-m2v8g128 + PV-Tuningq (mean bit precision)=2.127, tok/s (throughput)=214.4, PV-Tuning=true, Evaluation Protocol=0-shot2025.12 | 47.7 | |
| CodeGEMM-m1v4g128 + PV-Tuningq (mean bit precision)=2.126, tok/s (throughput)=228.3, PV-Tuning=true, Evaluation Protocol=0-shot2025.12 | 46.33 | |
| AQLM-1x16q (mean bit precision)=2.213, tok/s (throughput)=49.0, Evaluation Protocol=0-shot2025.12 | 46.16 | |
| AQLM-2x8 + PV-Tuningq (mean bit precision)=2.005, tok/s (throughput)=124.5, PV-Tuning=true, Evaluation Protocol=0-shot2025.12 | 45.73 | |
| RLVRBackbone=Llama-3.2-1B-Instruct, Training=RLVR, Ground-truth usage=true, Few-shot=zero-shot2026.03 | 42.2 | |
| FP16Bits=–, Backbone=LLAMA-2-7B2026.04 | 40.53 | |
| MIPOBackbone=Llama-3.2-1B-Instruct, Training=MIPO, Few-shot=zero-shot2026.03 | 39.53 | |
| CodeGEMM-m1v4g128q (mean bit precision)=2.126, tok/s (throughput)=228.3, Evaluation Protocol=0-shot2025.12 | 38.48 | |
| CodeGEMM-m2v8g128q (mean bit precision)=2.127, tok/s (throughput)=214.4, Evaluation Protocol=0-shot2025.12 | 37.03 | |
| Llama-3.2-1B-InstructBackbone=Llama-3.2-1B-Instruct, Few-shot=zero-shot2026.03 | 33.2 | |
| LBLLMBits=W(1+1)A4, Backbone=LLAMA-2-7B2026.04 | 32.51 | |
| AQLM-2x8q (mean bit precision)=2.005, tok/s (throughput)=124.5, Evaluation Protocol=0-shot2025.12 | 30.89 | |
| OnebitBits=W1A16, Backbone=LLAMA-2-7B2026.04 | 30.2 | |
| FBI-LLMBits=W1A16, Backbone=LLAMA-2-7B2026.04 | 29.9 | |
| SFTBackbone=Llama-3.2-1B-Instruct, Training=SFT, Few-shot=zero-shot2026.03 | 28.53 | |
| SFTBackbone=Llama-3.2-1B-Instruct, Training=SFT, Ground-truth usage=true, Few-shot=zero-shot2026.03 | 27.13 | |
| FlexRound-q2g128q (mean bit precision)=2.125, tok/s (throughput)=205.3, Evaluation Protocol=0-shot2025.12 | 24.57 | |
| GELUArchitecture=GELU MLP, Evaluation mode=zero-shot2026.05 | 22.47 | |
| Yat (sb+ca)Parameterization=shared (b, ε), Scaling=constant per-channel, Evaluation mode=zero-shot2026.05 | 22.44 | |
| Yat (pn+ca)Parameterization=per-neuron, Scaling=constant per-channel, Evaluation mode=zero-shot2026.05 | 22.07 | |
| Yat (pn+α)Parameterization=per-neuron, Scaling=learnable per-channel, Evaluation mode=zero-shot2026.05 | 22.01 | |
| Yat (sb+α)Parameterization=shared (b, ε), Scaling=learnable per-channel, Evaluation mode=zero-shot2026.05 | 21.39 |