Question Answering on SQuAD (EM/F1)
89.8F1SMP-S
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SMP-SRemaining Weights=50%, Knowledge Distillation=true2022.10 | 89.8 | 82.8 | |
| SMP-LRemaining Weights=50%, Knowledge Distillation=true2022.10 | 89.4 | 82.2 | |
| Prompt tuningBackbone=Llama-3.2-3B-Base, Evaluation Protocol=Prompt tuning, Context Setting=w/ context2026.02 | 88.41 | 80.39 | |
| Standard MHAMemory (GB)=15.22025.12 | 88.4 | — | |
| CAPRemaining Weights=50%, Knowledge Distillation=true2022.10 | 88.2 | 80.9 | |
| MAHAMemory (GB)=6.72025.12 | 88.2 | — | |
| BERT baseRemaining Weights=100%, Knowledge Distillation=false2022.10 | 88.1 | 80.4 | |
| BigBirdMemory (GB)=10.32025.12 | 87.9 | — | |
| MovementRemaining Weights=50%, Knowledge Distillation=true2022.10 | 87.6 | 79.8 | |
| LongformerMemory (GB)=9.12025.12 | 87.6 | — | |
| SMP-SRemaining Weights=10%, Knowledge Distillation=true2022.10 | 87.2 | 79.3 | |
| SMP-LRemaining Weights=10%, Knowledge Distillation=true2022.10 | 86.9 | 78.9 | |
| PerformerMemory (GB)=8.52025.12 | 86.7 | — | |
| CAPRemaining Weights=10%, Knowledge Distillation=true2022.10 | 85.6 | 77.1 | |
| S-STEParams=774M, Pre-train=S-STE, Fine-tune=S-STE2024.09 | 85.5 | 75.5 | |
| ReformerMemory (GB)=7.82025.12 | 85.2 | — | |
| COIECDModel=FLAN-T5-3B2024.02 | 84.99 | 73.84 | |
| Soft-MovementRemaining Weights=10%, Knowledge Distillation=true2022.10 | 84.9 | 76.6 | |
| DenseParams=774M, Pre-train=Dense, Fine-tune=Dense2024.09 | 84.9 | 74.3 | |
| SMP-SRemaining Weights=10%, Knowledge Distillation=false2022.10 | 84.6 | 75.1 | |
| T-SR-STE+DFParams=774M, Pre-train=T-SR-STE+DF [20], Fine-tune=Dense2024.09 | 84.6 | 74.3 | |
| SMP-LRemaining Weights=10%, Knowledge Distillation=false2022.10 | 84.3 | 75 | |
| MovementRemaining Weights=10%, Knowledge Distillation=true2022.10 | 84.3 | 75.6 | |
| SMP-SRemaining Weights=3%, Knowledge Distillation=true2022.10 | 84.1 | 75 | |
| DenseParams=350M, Pre-train=Dense, Fine-tune=Dense2024.09 | 83.6 | 73.2 | |
| RegularModel=FLAN-T5-3B2024.02 | 83.53 | 71.2 | |
| SMP-LRemaining Weights=3%, Knowledge Distillation=true2022.10 | 83.4 | 74 | |
| SCModel=FLAN-T5-3B2024.02 | 83.28 | 70.9 | |
| CDModel=FLAN-T5-3B2024.02 | 83.1 | 71.25 | |
| CAPRemaining Weights=3%, Knowledge Distillation=true2022.10 | 83 | 73.8 | |
| DeBERTa-V3 base + PiFiModel Backbone=DeBERTa-V3-base, PiFi Integration=true, Source LLM=Llama-3.1-8B2025.06 | 82.83 | 69.87 | |
| S-STEParams=350M, Pre-train=S-STE, Fine-tune=S-STE2024.09 | 82.7 | 72.2 | |
| T-SR-STEParams=350M, Pre-train=T-SR-STE, Fine-tune=Dense2024.09 | 82.6 | 72.3 | |
| DeBERTa-V3 baseModel Backbone=DeBERTa-V3-base2025.06 | 82.49 | 69.4 | |
| T-SR-STE+DFParams=350M, Pre-train=T-SR-STE+DF [20], Fine-tune=Dense2024.09 | 82.4 | 71.9 | |
| SR-STEParams=350M, Pre-train=SR-STE, Fine-tune=Dense2024.09 | 82.4 | 72 | |
| Soft-MovementRemaining Weights=3%, Knowledge Distillation=true2022.10 | 82.3 | 72.7 | |
| L0-regularizationRemaining Weights=10%, Knowledge Distillation=true2022.10 | 81.9 | 72.4 | |
| CADModel=FLAN-T5-3B2024.02 | 81.88 | 68.62 | |
| MovementRemaining Weights=10%, Knowledge Distillation=false2022.10 | 81.7 | 71.9 | |
| DeBERTa base + PiFiModel Backbone=DeBERTa-base, PiFi Integration=true, Source LLM=Llama-3.1-8B2025.06 | 81.52 | 69.65 | |
| Soft-MovementRemaining Weights=10%, Knowledge Distillation=false2022.10 | 81.5 | 71.3 | |
| SMP-SRemaining Weights=3%, Knowledge Distillation=false2022.10 | 81.4 | 70.9 | |
| Prompt tuningBackbone=Llama-3.2-1B-Base, Evaluation Protocol=Prompt tuning, Context Setting=w/ context2026.02 | 81.09 | 71.89 | |
| SMP-LRemaining Weights=3%, Knowledge Distillation=false2022.10 | 81 | 70.7 | |
| DeBERTa baseModel Backbone=DeBERTa-base2025.06 | 80.94 | 67.87 | |
| RoBERTa base + PiFiModel Backbone=RoBERTa-base, PiFi Integration=true, Source LLM=Llama-3.1-8B2025.06 | 80.89 | 68.97 | |
| ELECTRA base + PiFiModel Backbone=ELECTRA-base, PiFi Integration=true, Source LLM=Llama-3.1-8B2025.06 | 80.45 | 67.99 | |
| ELSAalpha_hat=0.22026.01 | 80.44 | 88.04 | |
| MagnitudeRemaining Weights=10%, Knowledge Distillation=true2022.10 | 80.1 | 70.2 | |
| RoBERTa baseModel Backbone=RoBERTa-base2025.06 | 80.06 | 68.09 | |
| FedCAdaalpha_hat=0.22026.01 | 80.03 | 87.29 | |
| Soft-MovementRemaining Weights=3%, Knowledge Distillation=false2022.10 | 79.9 | 69.5 | |
| w/ ContextCompression Rate=1x2026.02 | 79.81 | 59.97 | |
| RoFedalpha_hat=0.22026.01 | 79.33 | 87.16 | |
| ELSAalpha_hat=0.12026.01 | 79.24 | 87.17 | |
| OURS (PIC)Compression Rate=4x, Parameter Size=0.5B2026.02 | 79.03 | 60.14 | |
| FedProxalpha_hat=0.22026.01 | 79.01 | 87.7 | |
| RaSAalpha_hat=0.22026.01 | 78.99 | 86.95 | |
| DenseParams=124M, Pre-train=Dense, Fine-tune=Dense2024.09 | 78.8 | 67.6 | |
| S-STEParams=124M, Pre-train=S-STE, Fine-tune=S-STE2024.09 | 78.8 | 68 | |
| FedCAdaalpha_hat=0.12026.01 | 78.52 | 86.45 | |
| T-SR-STE+DFParams=124M, Pre-train=T-SR-STE+DF [20], Fine-tune=Dense2024.09 | 78.5 | 67.5 | |
| RoFedalpha_hat=0.12026.01 | 78.5 | 86.55 | |
| RaSAalpha_hat=0.12026.01 | 78.38 | 86.79 | |
| FedProxalpha_hat=0.12026.01 | 78.33 | 87.31 | |
| BERT base + PiFiModel Backbone=BERT-base, PiFi Integration=true, Source LLM=Llama-3.1-8B2025.06 | 78.09 | 66.17 | |
| FedAMSalpha_hat=0.22026.01 | 78.02 | 87.05 | |
| MovementRemaining Weights=3%, Knowledge Distillation=true2022.10 | 78 | 67.5 | |
| PCC-LargeCompression Rate=4x, Parameter Size=8B2026.02 | 77.76 | 60.04 | |
| Llama 3.3 70BModel Size=70B2025.01 | 77.7 | 66.2 | |
| FedAMSalpha_hat=0.12026.01 | 77.66 | 86.21 | |
| SR-STEParams=124M, Pre-train=SR-STE, Fine-tune=Dense2024.09 | 77.5 | 66.2 | |
| FedAvgalpha_hat=0.22026.01 | 77.22 | 87.72 | |
| T-SR-STEParams=124M, Pre-train=T-SR-STE, Fine-tune=Dense2024.09 | 77.2 | 66.3 | |
| FedAvgalpha_hat=0.12026.01 | 76.92 | 87.06 | |
| FedAvg (Random)alpha_hat=0.22026.01 | 76.52 | 85.98 | |
| MovementRemaining Weights=3%, Knowledge Distillation=false2022.10 | 76.3 | 65.2 | |
| BERT baseModel Backbone=BERT-base2025.06 | 76.06 | 63.81 | |
| PCC-LiteCompression Rate=4x, Parameter Size=0.77B2026.02 | 75.83 | 57.44 | |
| ComprExITBackbone=Llama-3.2-3B-Base, Evaluation Protocol=Context Compression2026.02 | 75.68 | 59.26 | |
| FedAvg (Random)alpha_hat=0.12026.01 | 75.33 | 85.66 | |
| Llama 3.2 3BModel Size=3B2025.01 | 73.3 | 60.7 | |
| BeaconBackbone=Llama-3.2-3B-Base, Evaluation Protocol=Context Compression2026.02 | 72.65 | 59.81 | |
| COIECDModel=LLaMA2-13B2024.02 | 70.86 | 57.1 | |
| SeCoBackbone=LLaMA-3.2-1B-Instruct, Compression constraint=16x2026.05 | 70.71 | 51.06 | |
| CADModel=LLaMA2-13B2024.02 | 70.52 | 56.46 | |
| ATACompressor2026.03 | 70.52 | 52.1 | |
| SeCoBackbone=LLaMA-3.2-1B-Instruct, Compression constraint=32x2026.05 | 69.47 | 50.34 | |
| RegularModel=LLaMA2-13B2024.02 | 68.92 | 54.46 | |
| SCModel=LLaMA2-13B2024.02 | 68.85 | 54.55 | |
| ComprExITBackbone=Llama-3.2-1B-Base, Evaluation Protocol=Context Compression2026.02 | 68.08 | 51.38 | |
| CDModel=LLaMA2-13B2024.02 | 68.04 | 53.89 | |
| 500xBackbone=Llama-3.2-3B-Base, Evaluation Protocol=Context Compression2026.02 | 65.56 | 52.08 | |
| SeCoBackbone=Qwen3-4B-Instruct, Compression constraint=16x2026.05 | 65.56 | 44.98 | |
| ProbeRAGArchitecture=LLaMA-2-7B-Chat-HF2025.10 | 65.4 | 52.1 | |
| GeAR2025.01 | 64.5 | 60 | |
| SeCoBackbone=Qwen3-4B-Instruct, Compression constraint=32x2026.05 | 63.46 | 42.96 | |
| CANOEArchitecture=LLaMA-2-7B-Chat-HF2025.10 | 63.2 | 45.6 | |
| ICAEBackbone=Llama-3.2-3B-Base, Evaluation Protocol=Context Compression2026.02 | 62.53 | 48.51 |