Language Modeling on WikiText-2 (DKL and PPL)
3.52PPLBF16
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| BF16Algorithm=BF16, Bits=16, Backbone=Llama 3.1 70B Instruct2025.05 | 3.52 | 0 | |
| BF16Base Model=Llama 3.1 70B Instruct, Bits=162025.05 | 3.52 | 0 | |
| YAQA-BBase Model=Llama 3.1 70B Instruct, Bits=4, Quantizer=QTIP, Recovery Finetuning=true, Incoherence processing=true2025.05 | 3.67 | 0.029 | |
| YAQA-ABase Model=Llama 3.1 70B Instruct, Bits=4, Quantizer=QTIP, Recovery Finetuning=true, Incoherence processing=true2025.05 | 3.68 | 0.032 | |
| YAQA-BAlgorithm=YAQA-B, Bits=4, Backbone=Llama 3.1 70B Instruct2025.05 | 3.69 | 0.03 | |
| LDLQBase Model=Llama 3.1 70B Instruct, Bits=4, Quantizer=QTIP, Recovery Finetuning=true, Incoherence processing=true2025.05 | 3.71 | 0.036 | |
| GuidedQuantAlgorithm=GuidedQuant, Bits=4, Backbone=Llama 3.1 70B Instruct2025.05 | 3.72 | 0.035 | |
| YAQA-AAlgorithm=YAQA-A, Bits=4, Backbone=Llama 3.1 70B Instruct2025.05 | 3.73 | 0.036 | |
| LDLQAlgorithm=LDLQ, Bits=4, Backbone=Llama 3.1 70B Instruct2025.05 | 3.74 | 0.045 | |
| YAQA-BBase Model=Llama 3.1 70B Instruct, Bits=3, Quantizer=QTIP, Recovery Finetuning=true, Incoherence processing=true2025.05 | 3.87 | 0.091 | |
| YAQA-ABase Model=Llama 3.1 70B Instruct, Bits=3, Quantizer=QTIP, Recovery Finetuning=true, Incoherence processing=true2025.05 | 3.88 | 0.098 | |
| LDLQBase Model=Llama 3.1 70B Instruct, Bits=3, Quantizer=QTIP, Recovery Finetuning=true, Incoherence processing=true2025.05 | 3.96 | 0.101 | |
| YAQA-BAlgorithm=YAQA-B, Bits=3, Backbone=Llama 3.1 70B Instruct2025.05 | 4.01 | 0.094 | |
| GuidedQuantAlgorithm=GuidedQuant, Bits=3, Backbone=Llama 3.1 70B Instruct2025.05 | 4.1 | 0.111 | |
| YAQA-AAlgorithm=YAQA-A, Bits=3, Backbone=Llama 3.1 70B Instruct2025.05 | 4.1 | 0.11 | |
| LDLQAlgorithm=LDLQ, Bits=3, Backbone=Llama 3.1 70B Instruct2025.05 | 4.26 | 0.138 | |
| YAQA-BBase Model=Llama 3.1 70B Instruct, Bits=2, Quantizer=QTIP, Recovery Finetuning=true, Incoherence processing=true2025.05 | 4.82 | 0.266 | |
| YAQA-ABase Model=Llama 3.1 70B Instruct, Bits=2, Quantizer=QTIP, Recovery Finetuning=true, Incoherence processing=true2025.05 | 4.92 | 0.279 | |
| LDLQBase Model=Llama 3.1 70B Instruct, Bits=2, Quantizer=QTIP, Recovery Finetuning=true, Incoherence processing=true2025.05 | 5.01 | 0.302 | |
| YAQA-BAlgorithm=YAQA-B, Bits=2, Backbone=Llama 3.1 70B Instruct2025.05 | 5.3 | 0.335 | |
| YAQA-AAlgorithm=YAQA-A, Bits=2, Backbone=Llama 3.1 70B Instruct2025.05 | 5.56 | 0.378 | |
| GuidedQuantAlgorithm=GuidedQuant, Bits=2, Backbone=Llama 3.1 70B Instruct2025.05 | 5.59 | 0.383 | |
| LDLQAlgorithm=LDLQ, Bits=2, Backbone=Llama 3.1 70B Instruct2025.05 | 6.02 | 0.497 | |
| BF16Algorithm=BF16, Bits=16, Backbone=Llama 3.1 8B Instruct2025.05 | 6.5 | 0 | |
| BF16Base Model=Llama 3.1 8B Instruct, Bits=162025.05 | 6.5 | 0 | |
| YAQA-BBase Model=Llama 3.1 8B Instruct, Bits=4, Quantizer=QTIP, Recovery Finetuning=true, Incoherence processing=true2025.05 | 6.56 | 0.012 | |
| YAQA-ABase Model=Llama 3.1 8B Instruct, Bits=4, Quantizer=QTIP, Recovery Finetuning=true, Incoherence processing=true2025.05 | 6.58 | 0.014 | |
| YAQA-BAlgorithm=YAQA-B, Bits=4, Backbone=Llama 3.1 8B Instruct2025.05 | 6.61 | 0.013 | |
| LDLQBase Model=Llama 3.1 8B Instruct, Bits=4, Quantizer=QTIP, Recovery Finetuning=true, Incoherence processing=true2025.05 | 6.61 | 0.016 | |
| YAQA-BBase Model=Llama 3.1 8B Instruct, Bits=3, Quantizer=QTIP, Recovery Finetuning=true, Incoherence processing=true2025.05 | 6.74 | 0.038 | |
| YAQA-ABase Model=Llama 3.1 8B Instruct, Bits=3, Quantizer=QTIP, Recovery Finetuning=true, Incoherence processing=true2025.05 | 6.78 | 0.042 | |
| LDLQBase Model=Llama 3.1 8B Instruct, Bits=3, Quantizer=QTIP, Recovery Finetuning=true, Incoherence processing=true2025.05 | 6.8 | 0.048 | |
| DenseSparsity=0%, Backbone=Llama2-7B2026.06 | 7.33 | — | |
| GOOGLE QATQuantization Type=QAT, Bits=4.52025.05 | 7.56 | 0.089 | |
| YAQA-BBase Model=Llama 3.1 8B Instruct, Bits=2, Quantizer=QTIP, Recovery Finetuning=true, Incoherence processing=true2025.05 | 7.6 | 0.147 | |
| YAQA-ABase Model=Llama 3.1 8B Instruct, Bits=2, Quantizer=QTIP, Recovery Finetuning=true, Incoherence processing=true2025.05 | 7.63 | 0.163 | |
| LDLQBase Model=Llama 3.1 8B Instruct, Bits=2, Quantizer=QTIP, Recovery Finetuning=true, Incoherence processing=true2025.05 | 7.82 | 0.185 | |
| BF16Quantization Type=NONE, Bits=162025.05 | 7.85 | 0 | |
| YAQA-BQuantization Type=PTQ, Bits=42025.05 | 7.94 | 0.056 | |
| YAQA-AQuantization Type=PTQ, Bits=42025.05 | 7.96 | 0.058 | |
| UNQUANTIZEDBITS=16, Context window=4K2025.05 | 9 | 0 | |
| YAQA-BBITS=4, Context window=4K2025.05 | 9.04 | 0.0153 | |
| LDLQBITS=4, Context window=4K2025.05 | 9.07 | 0.0216 | |
| YAQA-BBITS=3, Context window=4K2025.05 | 9.2 | 0.0479 | |
| LDLQBITS=3, Context window=4K2025.05 | 9.31 | 0.0681 | |
| BF16Algorithm=BF16, Bits=16, Backbone=Llama 3.2 3B Instruct2025.05 | 9.58 | 0 | |
| YAQA-BAlgorithm=YAQA-B, Bits=4, Backbone=Llama 3.2 3B Instruct2025.05 | 9.8 | 0.014 | |
| YAQA-BBITS=2, Context window=4K2025.05 | 10.16 | 0.205 | |
| LDLQBITS=2, Context window=4K2025.05 | 10.81 | 0.285 | |
| BF16Algorithm=BF16, Bits=16, Backbone=Llama 3.2 1B Instruct2025.05 | 11.57 | 0 | |
| YAQA-BAlgorithm=YAQA-B, Bits=4, Backbone=Llama 3.2 1B Instruct2025.05 | 11.83 | 0.019 | |
| UniRankSparsity=40%, Backbone=Llama2-7B, Base Method=LoRAP2026.06 | 14.03 | — | |
| SkipGPTSparsity=40%, Backbone=Llama2-7B2026.06 | 14.35 | — | |
| SliceGPTSparsity=40%, Backbone=Llama2-7B2026.06 | 14.87 | — | |
| LoRAPSparsity=40%, Backbone=Llama2-7B2026.06 | 16.08 | — | |
| DRDG/TWEvidence form=Dual-end output aggregation (token-wise sync)2026.06 | 17.016 | — | |
| DRAGONEvidence form=Dual-end output aggregation2026.06 | 17.016 | — | |
| CONCORDEvidence form=Dual-end output aggregation2026.06 | 17.016 | — | |
| DRCGEvidence form=Dual-end evidence merged into single-end context2026.06 | 17.179 | — | |
| Shortened-llamaSparsity=40%, Backbone=Llama2-7B2026.06 | 17.22 | — | |
| ShortGPTSparsity=40%, Backbone=Llama2-7B2026.06 | 18.35 | — | |
| CRCGEvidence form=Single-end local shard2026.06 | 21.351 | — | |
| MoDSparsity=40%, Backbone=Llama2-7B2026.06 | 40.42 | — | |
| D-LLMSparsity=40%, Backbone=Llama2-7B2026.06 | 52.78 | — |