Question Answering on ARC Easy (Accuracy)
98.2AccuracyMistral Small 24B Inst 2501
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Mistral Small 24B Inst 2501Classifier=Self-labeled2026.01 | 98.2 | — | 0 | — | — | |
| Mistral Small 24B Inst 2501Classifier=Majority-labeled2026.01 | 98.2 | — | 0 | — | — | |
| Qwen3 14B BaseClassifier=Self-labeled2026.01 | 98 | — | — | — | — | |
| Qwen3 14B BaseClassifier=Majority-labeled2026.01 | 98 | — | — | — | — | |
| Mistral Small 24B Base 2501Classifier=Self-labeled2026.01 | 97.8 | — | 0 | — | — | |
| Mistral Small 24B Base 2501Classifier=Majority-labeled2026.01 | 97.8 | — | 0 | — | — | |
| Qwen3 14BClassifier=Self-labeled2026.01 | 97.5 | — | — | — | — | |
| Qwen3 14BClassifier=Majority-labeled2026.01 | 97.4 | — | — | — | — | |
| Qwen3 8B BaseClassifier=Self-labeled2026.01 | 97.3 | — | 0.1 | — | — | |
| Qwen3 8B BaseClassifier=Majority-labeled2026.01 | 97.3 | — | 0.1 | — | — | |
| Qwen3 8B BaseClassifier=Self-labeled2026.01 | 97.3 | — | — | — | — | |
| Qwen3 8B BaseClassifier=Majority-labeled2026.01 | 97.3 | — | — | — | — | |
| Qwen3 4B BaseClassifier=Self-labeled2026.01 | 96.4 | — | — | — | — | |
| Qwen3 4B BaseClassifier=Majority-labeled2026.01 | 96.4 | — | — | — | — | |
| Qwen3 8BClassifier=Self-labeled2026.01 | 96.2 | — | — | — | — | |
| Qwen3 8BClassifier=Majority-labeled2026.01 | 96.2 | — | — | — | — | |
| TASOModel=Qwen2.5 3B, #Param.=2.06 M, Evaluation Protocol=Zero-shot2025.09 | 94.84 | — | — | — | — | |
| DoRAModel=Qwen2.5 3B, #Param.=65.98 M, rank=32, Evaluation Protocol=Zero-shot2025.09 | 94.73 | — | — | — | — | |
| LoRAModel=Qwen2.5 3B, #Param.=59.87 M, rank=32, Evaluation Protocol=Zero-shot2025.09 | 94.63 | — | — | — | — | |
| AdaLoRAModel=Qwen2.5 3B, #Param.=61.32 M, rank=32, Evaluation Protocol=Zero-shot2025.09 | 94.63 | — | — | — | — | |
| LoRAModel=Qwen2.5 3B, #Param.=14.97 M, rank=8, Evaluation Protocol=Zero-shot2025.09 | 94.58 | — | — | — | — | |
| Qwen3 4BClassifier=Self-labeled2026.01 | 94.1 | — | — | — | — | |
| Qwen3 4BClassifier=Majority-labeled2026.01 | 94.1 | — | — | — | — | |
| VERAModel=Qwen2.5 3B, #Param.=1.42 M, rank=1024, Evaluation Protocol=Zero-shot2025.09 | 94.05 | — | — | — | — | |
| AdapterModel=Qwen2.5 3B, #Param.=67.10 M, Evaluation Protocol=Zero-shot2025.09 | 93.84 | — | — | — | — | |
| Fine-tuneModel=Qwen2.5 3B, #Param.=3151.91 M, Evaluation Protocol=Zero-shot2025.09 | 93.74 | — | — | — | — | |
| Llama 3.1 8B InstClassifier=Self-labeled2026.01 | 93.2 | — | 0 | — | — | |
| Llama 3.1 8B InstClassifier=Majority-labeled2026.01 | 93.2 | — | 0 | — | — | |
| IA3Model=Qwen2.5 3B, #Param.=1.35 M, Evaluation Protocol=Zero-shot2025.09 | 93.16 | — | — | — | — | |
| LoRA-XSModel=Qwen2.5 3B, #Param.=4.13 M, rank=128, Evaluation Protocol=Zero-shot2025.09 | 92.89 | — | — | — | — | |
| Llama 3.1 8BClassifier=Self-labeled2026.01 | 92.1 | — | 0.3 | — | — | |
| Llama 3.1 8BClassifier=Self-labeled2026.01 | 92.1 | — | 0.3 | — | — | |
| Llama 3.1 8BClassifier=Majority-labeled2026.01 | 92 | — | 0.1 | — | — | |
| Llama 3.1 8BClassifier=Majority-labeled2026.01 | 92 | — | 0.1 | — | — | |
| Qwen3 1.7B BaseClassifier=Self-labeled2026.01 | 91.9 | — | — | — | — | |
| Qwen3 1.7B BaseClassifier=Majority-labeled2026.01 | 91.8 | — | — | — | — | |
| Mistral Nemo Base 2407Classifier=Self-labeled2026.01 | 91 | — | -0.1 | — | — | |
| Mistral Nemo Base 2407Classifier=Majority-labeled2026.01 | 91 | — | -0.1 | — | — | |
| Mistral Nemo Base 2407Classifier=Self-labeled2026.01 | 91 | — | -0.1 | — | — | |
| Mistral Nemo Base 2407Classifier=Majority-labeled2026.01 | 91 | — | -0.1 | — | — | |
| W8A8-DirectModel=Qwen3-30B2026.04 | 90.61 | 0.04 | — | — | — | |
| W16A16-DirectModel=Qwen3-30B2026.04 | 90.57 | 0 | — | — | — | |
| MXFP4-DynamicModel=Qwen3-30B2026.04 | 90.57 | 0 | — | — | — | |
| Tensor RMS + CFormat=Tensor RMS + C, Bit width (b)=3.00, Quantization Aware Training (QAT)=true, Backbone=Llama 3.1 8B2025.05 | 90.2 | — | — | — | — | |
| MXFP4-W2FP8+DynModel=Qwen3-30B2026.04 | 90.19 | -0.38 | — | — | — | |
| BaselineFormat=Baseline, Bit width (b)=16.00, Quantization Aware Training (QAT)=true, Backbone=Llama 3.1 8B2025.05 | 90 | — | — | — | — | |
| BaselineBit width (b)=16.002025.05 | 90 | — | — | — | — | |
| MXFP4-OSCModel=Qwen3-30B2026.04 | 89.98 | -0.59 | — | — | — | |
| TS (C)Contextual=true2026.02 | 89.5 | 8.7 | — | — | — | |
| TS (C)Algorithm Family=Contextual Linear2026.02 | 89.5 | — | — | — | — | |
| MXFP4-W2FP8Model=Qwen3-30B2026.04 | 89.35 | -1.22 | — | — | — | |
| LinUCBAlgorithm Family=Contextual Linear2026.02 | 89.2 | — | — | — | — | |
| EXP3Algorithm Family=Non-Contextual2026.02 | 89 | — | — | — | — | |
| MXFP4-DirectModel=Qwen3-30B2026.04 | 88.85 | -1.73 | — | — | — | |
| LinUCB+KLAlgorithm Family=Contextual Linear2026.02 | 88.8 | — | — | — | — | |
| Mistral 7B v0.3Classifier=Self-labeled2026.01 | 88.6 | — | 0 | — | — | |
| Mistral 7B v0.3Classifier=Self-labeled2026.01 | 88.6 | — | 0 | — | — | |
| Mistral 7B Inst v0.3Classifier=Self-labeled2026.01 | 88.5 | — | 0.3 | — | — | |
| Mistral 7B v0.3Classifier=Majority-labeled2026.01 | 88.4 | — | -0.2 | — | — | |
| Mistral 7B v0.3Classifier=Majority-labeled2026.01 | 88.4 | — | -0.2 | — | — | |
| Llama 3.2 3B InstClassifier=Self-labeled2026.01 | 88.3 | — | 0.3 | — | — | |
| Llama 3.2 3B InstClassifier=Majority-labeled2026.01 | 88.3 | — | 0.2 | — | — | |
| Mistral 7B Inst v0.3Classifier=Majority-labeled2026.01 | 88.3 | — | 0.1 | — | — | |
| W16A16-DirectModel=Qwen3-8B2026.04 | 87.96 | 0 | — | — | — | |
| Qwen3 1.7BClassifier=Self-labeled2026.01 | 87.9 | — | — | — | — | |
| Mashup LearningBackbone=Mistral-7B-Instruct-v0.2, Setup=Lots-of-LoRAs2026.03 | 87.9 | — | — | — | — | |
| Qwen3 1.7BClassifier=Majority-labeled2026.01 | 87.8 | — | — | — | — | |
| W8A8-DirectModel=Qwen3-8B2026.04 | 87.75 | -0.21 | — | — | — | |
| TS (NC)Algorithm Family=Non-Contextual2026.02 | 87.7 | — | — | — | — | |
| MXFP4-W2FP8+DynModel=Qwen3-8B2026.04 | 86.95 | -1.01 | — | — | — | |
| MXFP4-OSCModel=Qwen3-8B2026.04 | 86.91 | -1.05 | — | — | — | |
| LinEXP3Algorithm Family=Contextual Linear2026.02 | 86.9 | — | — | — | — | |
| TASOModel=LLaMA3.2 3B, #Param.=1.67 M, Evaluation Protocol=Zero-shot2025.09 | 86.74 | — | — | — | — | |
| Mistral Nemo Inst 2407Classifier=Self-labeled2026.01 | 86.7 | — | 2.7 | — | — | |
| MXFP4-DynamicModel=Qwen3-8B2026.04 | 86.66 | -1.3 | — | — | — | |
| AdapterModel=LLaMA3.2 3B, #Param.=50.32 M, Evaluation Protocol=Zero-shot2025.09 | 86.53 | — | — | — | — | |
| LoRAModel=LLaMA3.2 3B, #Param.=48.63 M, rank=32, Evaluation Protocol=Zero-shot2025.09 | 86.53 | — | — | — | — | |
| Tensor RMS + SpFormat=Tensor RMS + Sp, Bit width (b)=3.05, Quantization Aware Training (QAT)=true, Backbone=Llama 3.1 8B2025.05 | 86.5 | — | — | — | — | |
| DoRAModel=LLaMA3.2 3B, #Param.=53.73 M, rank=32, Evaluation Protocol=Zero-shot2025.09 | 86.32 | — | — | — | — | |
| MXFP4-W2FP8Model=Qwen3-8B2026.04 | 86.11 | -1.85 | — | — | — | |
| e-FTRLAlgorithm Family=Non-Contextual2026.02 | 85.9 | — | — | — | — | |
| AdaLoRAModel=LLaMA3.2 3B, #Param.=49.51 M, rank=32, Evaluation Protocol=Zero-shot2025.09 | 85.58 | — | — | — | — | |
| LoRAModel=LLaMA3.2 3B, #Param.=12.16 M, rank=8, Evaluation Protocol=Zero-shot2025.09 | 85.21 | — | — | — | — | |
| Fine-tuneModel=LLaMA3.2 3B, #Param.=3266.58 M, Evaluation Protocol=Zero-shot2025.09 | 85.11 | — | — | — | — | |
| Mashup Learning at initBackbone=Mistral-7B-Instruct-v0.2, Setup=Lots-of-LoRAs2026.03 | 85 | — | — | — | — | |
| Llama 3.2 3BClassifier=Self-labeled2026.01 | 84.8 | — | 0.3 | — | — | |
| IA3Model=LLaMA3.2 3B, #Param.=0.91 M, Evaluation Protocol=Zero-shot2025.09 | 84.79 | — | — | — | — | |
| MXFP4-DirectModel=Qwen3-8B2026.04 | 84.51 | -3.45 | — | — | — | |
| Llama 3.2 3BClassifier=Majority-labeled2026.01 | 84.5 | — | 0 | — | — | |
| Tensor RMS + CBit width (b)=3.002025.05 | 84.4 | — | — | — | — | |
| VERAModel=LLaMA3.2 3B, #Param.=1.09 M, rank=1024, Evaluation Protocol=Zero-shot2025.09 | 84.27 | — | — | — | — | |
| Block AbsmaxFormat=Block Absmax, Bit width (b)=3.25, Quantization Aware Training (QAT)=true, Backbone=Llama 3.1 8B2025.05 | 84.2 | — | — | — | — | |
| LoRA-XSModel=LLaMA3.2 3B, #Param.=3.21 M, rank=128, Evaluation Protocol=Zero-shot2025.09 | 84.06 | — | — | — | — | |
| Mistral Nemo Inst 2407Classifier=Majority-labeled2026.01 | 84 | — | -0.1 | — | — | |
| From scratchBackbone=Mistral-7B-Instruct-v0.2, Setup=Lots-of-LoRAs2026.03 | 84 | — | — | — | — | |
| FP16Base Model=Qwen3-8B2026.01 | 83.59 | — | — | — | — | |
| Text-to-LoRA used as initBackbone=Mistral-7B-Instruct-v0.2, Setup=Lots-of-LoRAs2026.03 | 83.1 | — | — | — | — | |
| HiF4+HiGPTQModel=Qwen2.5-14B, A-W Quant Type=HiF4+HiGPTQ2026.02 | 83.04 | 3.7 | — | — | — | |
| DCRBase Model=Qwen2.5-7B2026.02 | 83 | — | — | — | — | |
| MoonlightPruning ratio=0%, Backbone=Moonlight, Evaluation protocol=Zero-shot, Calibration samples=100, Calibration dataset=Zyda22025.07 | 82.49 | — | — | — | — |