Long-context Information Extraction on RULER 4K-32K Average
46.39CWE ScoreTraining–Inference Consistent Segmented Execution
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Training–Inference Consistent Segmented ExecutionBackbone=LLaMA2-7B-32K2026.05 | 46.39 | 43.88 | |
| LLaMA2-7B-32K (Vanilla Self-Attention)Backbone=LLaMA2-7B-32K, Attention=Vanilla Self-Attention2026.05 | 32.94 | 41.33 | |
| MInferenceBackbone=LLaMA2-7B-32K2026.05 | 32.94 | 41.75 | |
| DuoAttentionBackbone=LLaMA2-7B-32K2026.05 | 32.33 | 43.42 | |
| StreamingLLMBackbone=LLaMA2-7B-32K2026.05 | 27.78 | 41.37 | |
| CCA-attentionBackbone=LLaMA2-7B-32K2026.05 | 24.9 | 31.96 |