Long-context understanding on LongBench (Difficulty and Length Aggregation)
31.8Overall Average ScoreBLASST
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| BLASSTModel=Llama-3.1-8B-Instruct, Phase=Prefill, Sparsity=~50%2025.12 | 31.8 | 30.7 | 32.5 | 38.3 | 29.8 | 25 | |
| Dense AttentionModel=Llama-3.1-8B-Instruct, Phase=Prefill2025.12 | 31.4 | 29.7 | 32.5 | 38.3 | 28.8 | 25 | |
| MInferenceModel=Llama-3.1-8B-Instruct, Phase=Prefill2025.12 | 31.2 | 28.6 | 32.8 | 36.7 | 30.2 | 24.1 | |
| XAttentionModel=Llama-3.1-8B-Instruct, Phase=Prefill2025.12 | 30.6 | 29.2 | 31.5 | 38.3 | 26 | 26.9 | |
| FlexPrefillModel=Llama-3.1-8B-Instruct, Phase=Prefill2025.12 | 25.7 | 28.8 | 23.8 | 24.4 | 26.5 | 26.2 | |
| GPTAQ + MaCaModel=Qwen3-4B-IT, Quantization=4bit2026.02 | 11.12 | — | — | — | — | — | |
| GPTAQModel=Qwen3-4B-IT, Quantization=4bit2026.02 | 10.81 | — | — | — | — | — | |
| GPTAQ + MaCaModel=Gemma3-4B-IT, Quantization=4bit2026.02 | 9.72 | — | — | — | — | — | |
| GPTAQModel=Gemma3-4B-IT, Quantization=4bit2026.02 | 9.71 | — | — | — | — | — | |
| GPTQ + MaCaModel=Gemma3-4B-IT, Quantization=4bit2026.02 | 9.49 | — | — | — | — | — | |
| GPTQ + MaCaModel=Qwen3-4B-IT, Quantization=4bit2026.02 | 8.31 | — | — | — | — | — | |
| GPTQModel=Gemma3-4B-IT, Quantization=4bit2026.02 | 8.29 | — | — | — | — | — | |
| GPTQModel=Qwen3-4B-IT, Quantization=4bit2026.02 | 6.07 | — | — | — | — | — | |
| GPTAQ + MaCaModel=LLaMA3.2-3B-IT, Quantization=4bit2026.02 | 3.84 | — | — | — | — | — | |
| GPTAQModel=LLaMA3.2-3B-IT, Quantization=4bit2026.02 | 3.64 | — | — | — | — | — | |
| GPTQModel=LLaMA3.2-3B-IT, Quantization=4bit2026.02 | 0.14 | — | — | — | — | — | |
| GPTQ + MaCaModel=LLaMA3.2-3B-IT, Quantization=4bit2026.02 | 0.14 | — | — | — | — | — |