Long-context understanding on LongBench V1 (QA, Summarization, and Retrieval Subset)
59.63HotpotQA AccuracySpec prefill with Llama-3.2-1B-Instruct
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Spec prefill with Llama-3.2-1B-InstructTarget Model=Qwen3-8B, Keep rate=45%, Draft Model=Llama-3.2-1B-Instruct2026.03 | 59.63 | 100 | 24.05 | 52.56 | 46.21 | 49.74 | 59.69 | 99 | 55.98 | |
| Full PromptTarget Model=Qwen3-8B, Keep rate=Full2026.03 | 59.47 | 100 | 24.03 | 53.84 | 50.49 | 51.13 | 62.45 | 98.5 | 57.34 | |
| Spec prefill with Qwen3-1.7BTarget Model=Llama3.1-8B-Instruct, Keep rate=45%, Draft Model=Qwen3-1.7B2026.03 | 55.99 | 100 | 24.9 | 54.32 | 41.09 | 41.33 | 62.4 | 95.25 | 54.29 | |
| Full PromptTarget Model=Llama3.1-8B-Instruct, Keep rate=Full2026.03 | 55.97 | 99.5 | 25.38 | 55.8 | 44.56 | 43.5 | 62.45 | 97.38 | 55.31 |