Question Answering on ARC Challenge (Accuracy)
87.3Accuracy (ARC)Frozen LLM graph
Evaluation Results
| Method | Links | |
|---|---|---|
| Frozen LLM graphNumber of Parameters=17.6M2026.04 | 87.3 | |
| TSDSIteration=2, Selection Ratio=0.5, Base Model=3b base model2026.04 | 83.45 | |
| EVOSELECTIteration=1, Selection Ratio=0.2, Base Model=3b base model2026.04 | 82.85 | |
| DiversityIteration=2, Selection Ratio=0.5, Base Model=3b base model2026.04 | 82.68 | |
| EVOSELECTIteration=2, Selection Ratio=0.2, Base Model=3b base model2026.04 | 82.42 | |
| RandomIteration=1, Selection Ratio=0.2, Base Model=3b base model2026.04 | 82.34 | |
| AttributionIteration=2, Selection Ratio=0.5, Base Model=3b base model2026.04 | 82.34 | |
| TSDSIteration=2, Selection Ratio=0.2, Base Model=3b base model2026.04 | 82.29 | |
| EVOSELECTIteration=1, Selection Ratio=0.5, Base Model=3b base model2026.04 | 82.08 | |
| DiversityIteration=1, Selection Ratio=0.5, Base Model=3b base model2026.04 | 82 | |
| AllIteration=2, Selection Ratio=1.0, Base Model=3b base model2026.04 | 82 | |
| DiversityIteration=2, Selection Ratio=0.2, Base Model=3b base model2026.04 | 81.91 | |
| EVOSELECTIteration=2, Selection Ratio=0.5, Base Model=3b base model2026.04 | 81.91 | |
| RandomIteration=2, Selection Ratio=0.5, Base Model=3b base model2026.04 | 81.83 | |
| AttributionIteration=2, Selection Ratio=0.2, Base Model=3b base model2026.04 | 81.83 | |
| DiversityIteration=1, Selection Ratio=0.2, Base Model=3b base model2026.04 | 81.74 | |
| Teacher-OPPOBackbone=Phi-4-mini, Training Framework=DeepScaleR2026.05 | 81.6 | |
| RandomIteration=2, Selection Ratio=0.2, Base Model=3b base model2026.04 | 81.57 | |
| BaseIteration=0, Selection Ratio=N/A, Base Model=3b base model2026.04 | 81.4 | |
| AttributionIteration=1, Selection Ratio=0.2, Base Model=3b base model2026.04 | 81.4 | |
| DAPO + OPPOBackbone=Phi-4-mini, Training Framework=DeepScaleR2026.05 | 81.4 | |
| Attr-DivIteration=2, Selection Ratio=0.5, Base Model=3b base model2026.04 | 81.31 | |
| Attr-DivIteration=1, Selection Ratio=0.5, Base Model=3b base model2026.04 | 81.23 | |
| TSDSIteration=1, Selection Ratio=0.5, Base Model=3b base model2026.04 | 81.23 | |
| Attr-DivIteration=2, Selection Ratio=0.2, Base Model=3b base model2026.04 | 81.06 | |
| Self-OPPOBackbone=Phi-4-mini, Training Framework=DeepScaleR2026.05 | 81 | |
| Attr-DivIteration=1, Selection Ratio=0.2, Base Model=3b base model2026.04 | 80.89 | |
| AttributionIteration=1, Selection Ratio=0.5, Base Model=3b base model2026.04 | 80.8 | |
| Dr.GRPO + OPPOBackbone=Phi-4-mini, Training Framework=DeepScaleR2026.05 | 80.8 | |
| GPT4oBackbone Architecture=GPT-4o2026.05 | 80.71 | |
| RandomIteration=1, Selection Ratio=0.5, Base Model=3b base model2026.04 | 80.63 | |
| TSDSIteration=1, Selection Ratio=0.2, Base Model=3b base model2026.04 | 80.63 | |
| AllIteration=1, Selection Ratio=1.0, Base Model=3b base model2026.04 | 80.55 | |
| IT-PrunModel=Llama3-8B, Ratio=31.20%, Zero-shot=true, Supervised Fine-Tuning (SFT)=true2026.05 | 79.86 | |
| SDPOBackbone=Phi-4-mini, Training Framework=DeepScaleR2026.05 | 79.8 | |
| DAPOBackbone=Phi-4-mini, Training Framework=DeepScaleR2026.05 | 79.4 | |
| Dr.GRPOBackbone=Phi-4-mini, Training Framework=DeepScaleR2026.05 | 78.8 | |
| GRPOBackbone=Phi-4-mini, Training Framework=DeepScaleR2026.05 | 78.6 | |
| Head on Qwen2.5-1.5BEvaluation Protocol=Parameter-matched learned head, Number of Parameters=22.8M2026.04 | 78.2 | |
| ShortGPTModel=Llama3-8B, Ratio=21.90%, Zero-shot=true, Supervised Fine-Tuning (SFT)=true2026.05 | 77.82 | |
| ShortGPTModel=Llama3-8B, Ratio=12.50%, Zero-shot=true, Supervised Fine-Tuning (SFT)=true2026.05 | 77.05 | |
| IT-PrunModel=Llama3-8B, Ratio=12.50%, Zero-shot=true, Supervised Fine-Tuning (SFT)=true2026.05 | 76.96 | |
| SPEARModel=Qwen2.5-1.5B2026.05 | 76.93 | |
| GRPOModel=Qwen2.5-1.5B2026.05 | 76.3 | |
| IT-PrunModel=Llama3-8B, Ratio=21.90%, Zero-shot=true, Supervised Fine-Tuning (SFT)=true2026.05 | 76.02 | |
| Qwen2.5-1.5B-InstructDecoding Strategy=Single-model greedy2026.04 | 75.9 | |
| ShortGPTModel=Llama3-8B, Ratio=31.20%, Zero-shot=true, Supervised Fine-Tuning (SFT)=true2026.05 | 75.34 | |
| Phi-3-mini-4K-InstructDecoding Strategy=Single-model greedy2026.04 | 75.3 | |
| RLTF-SDModel=Qwen2.5-1.5B2026.05 | 75.03 | |
| AMATA 8BModel Scale=8B, Backbone Architecture=AMATA2026.05 | 74.11 | |
| Mistral-7B-Instruct-v0.3Decoding Strategy=Single-model greedy2026.04 | 73.9 | |
| Feedback SFTModel=Qwen2.5-1.5B2026.05 | 73.89 | |
| GiGPO 8BModel Scale=8B, Backbone Architecture=GiGPO2026.05 | 73.84 | |
| MKAModel=Llama3-8B, Ratio=21.90%, Zero-shot=true, Supervised Fine-Tuning (SFT)=true2026.05 | 73.12 | |
| MKAModel=Llama3-8B, Ratio=12.50%, Zero-shot=true, Supervised Fine-Tuning (SFT)=true2026.05 | 73.04 | |
| SPA-RL 8BModel Scale=8B, Backbone Architecture=SPA-RL2026.05 | 72.98 | |
| SMART 8BModel Scale=8B, Backbone Architecture=SMART2026.05 | 72.81 | |
| Base modelBackbone=Phi-4-mini, Training Framework=DeepScaleR2026.05 | 72.8 | |
| SLEBModel=Llama3-8B, Ratio=12.50%, Zero-shot=true, Supervised Fine-Tuning (SFT)=true2026.05 | 72.35 | |
| MKAModel=Llama3-8B, Ratio=31.20%, Zero-shot=true, Supervised Fine-Tuning (SFT)=true2026.05 | 69.71 | |
| SPEARModel=Llama3.2-3B2026.05 | 69.17 | |
| SelfRag 8BModel Scale=8B, Backbone Architecture=Self-RAG2026.05 | 68.14 | |
| Feedback SFTModel=Llama3.2-3B2026.05 | 67.95 | |
| Gemma-2-2B-ITDecoding Strategy=Single-model greedy2026.04 | 67.1 | |
| GRPOModel=Llama3.2-3B2026.05 | 65.7 | |
| GPTQModel=Llama-3.1-405B-Instruct, Prec.=W4A16 g128, Evaluation Protocol=Zero-shot2026.05 | 65.1 | |
| OSAQ+GPTQModel=Llama-3.1-405B-Instruct, Prec.=W4A16 g128, Evaluation Protocol=Zero-shot2026.05 | 65 | |
| RADIT 8BModel Scale=8B, Backbone Architecture=RADIT2026.05 | 64.88 | |
| LeanQuantModel=Llama-3.1-405B-Instruct, Prec.=W4A16 g128, Evaluation Protocol=Zero-shot2026.05 | 64.8 | |
| Qwen2.5-DenseLLM=Qwen2.5, Model=Dense, Activated Parameters=1.5 B2026.05 | 64.42 | |
| TALEModel=LLaMA-2-13B, Sparsity=10%, Evaluation Protocol=0-shot2025.10 | 64.4 | |
| OSAQ+GPTQModel=Mistral-Large-123B-Instruct, Prec.=W4A16, Evaluation Protocol=Zero-shot2026.05 | 64.1 | |
| TALEModel=LLaMA-2-13B, Sparsity=25%, Evaluation Protocol=0-shot2025.10 | 64.1 | |
| LeanQuantModel=Mistral-Large-123B-Instruct, Prec.=W4A16, Evaluation Protocol=Zero-shot2026.05 | 64 | |
| GPTQModel=Mistral-Large-123B-Instruct, Prec.=W4A16, Evaluation Protocol=Zero-shot2026.05 | 64 | |
| RLTF-SDModel=Llama3.2-3B2026.05 | 63.28 | |
| Qwen2.5-Dense2MoELLM=Qwen2.5, Model=Ours, Activated Parameters=1.2 B2026.05 | 62.2 | |
| RTNModel=Qwen3.5-27B, Group Size (g)=162026.05 | 61.9 | |
| AWQModel=Qwen3.5-27B, Group Size (g)=162026.05 | 61.9 | |
| AAACModel=Qwen3.5-27B, Group Size (g)=162026.05 | 61.9 | |
| FP16Model=MIXTRAL-8X7B, Quantization=2-bit2026.05 | 61.9 | |
| FP16Model=MIXTRAL-8X7B, Quantization=3-bit2026.05 | 61.9 | |
| BF16Model=Qwen3.5-27B2026.05 | 61.7 | |
| TILEQvModel=MIXTRAL-8X7B, Quantization=3-bit, Extra Bits (Scales/Low-rank)=0.16/0.032026.05 | 61.4 | |
| IF4Model=Qwen3.5-27B, Group Size (g)=162026.05 | 61.3 | |
| Full PrecisionModel=Qwen-30B-A3B2026.05 | 61 | |
| Uniform SparsityBackbone=Mistral-7B2026.03 | 60.5 | |
| MALSBackbone=Mistral-7B2026.03 | 60.5 | |
| TILEQsModel=MIXTRAL-8X7B, Quantization=3-bit, Extra Bits (Scales/Low-rank)=0.16/0.032026.05 | 60.4 | |
| COVERCALModel=Mistral-7B, Quantization=GPTQ INT4, Calibration Samples=128, Seeds=32026.04 | 60.3 | |
| LOPROModel=MIXTRAL-8X7B, Quantization=3-bit, Extra Bits (Scales/Low-rank)=0.21/0.082026.05 | 60.3 | |
| RandomModel=Qwen32026.04 | 60.2 | |
| BaselineModel=Qwen32026.04 | 59.8 | |
| Moderate-GiModel=Qwen2.52026.04 | 59.6 | |
| KLModel=Qwen32026.04 | 59.5 | |
| O-LoRAModel=Qwen32026.04 | 59.5 | |
| Max-ActVarModel=Mistral-7B, Quantization=GPTQ INT4, Calibration Samples=128, Seeds=32026.04 | 59.4 | |
| EWCModel=Qwen32026.04 | 59.3 | |
| HCInferModel=Qwen-30B-A3B2026.05 | 59.25 | |
| Moderate-GiModel=Qwen32026.04 | 59.2 |