Retrieval on Needle-in-A-Haystack RULER
100Success Rate (4K Context)FullAttn
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| FullAttnTraining Protocol=Length Extrapolation (RoPE Scaling), Inference Mode=Full Attention, Receptive Field=Full2025.11 | 100 | 100 | 48 | 35.4 | 70.9 | |
| MoBATraining Protocol=Length Extrapolation (RoPE Scaling), Inference Mode=Full Attention, Receptive Field=Full2025.11 | 100 | 32.4 | 2.2 | 0 | 33.7 | |
| SSATraining Protocol=Length Extrapolation (RoPE Scaling), Inference Mode=Full Attention, Receptive Field=Full2025.11 | 100 | 100 | 83 | 35.4 | 79.6 | |
| FullAttnTraining Protocol=Continual Trained to 32k, Inference Mode=Full Attention, Receptive Field=Full2025.11 | 100 | 100 | 99.2 | 93.2 | 98.1 | |
| SSATraining Protocol=Continual Trained to 32k, Inference Mode=Full Attention, Receptive Field=Full2025.11 | 100 | 100 | 99.8 | 70.6 | 92.6 | |
| MoBATraining Protocol=Continual Trained to 32k, Inference Mode=Full Attention, Receptive Field=Full2025.11 | 99 | 73.2 | 5.4 | 1.2 | 44.7 | |
| SSATraining Protocol=Continual Trained to 32k, Inference Mode=Sparse Attention, Receptive Field=10242025.11 | 94 | 55.6 | 17.2 | 5.4 | 43.1 | |
| MoBATraining Protocol=Continual Trained to 32k, Inference Mode=Sparse Attention, Receptive Field=10242025.11 | 93.8 | 36.2 | 17.6 | 9.4 | 39.3 | |
| SSATraining Protocol=Continual Trained to 32k, Inference Mode=Sparse Attention, Receptive Field=2562025.11 | 89.2 | 34.8 | 7.8 | 3.2 | 33.8 | |
| MoBATraining Protocol=Continual Trained to 32k, Inference Mode=Sparse Attention, Receptive Field=2562025.11 | 82 | 27.6 | 12.4 | 3 | 31.3 | |
| SSATraining Protocol=Length Extrapolation (RoPE Scaling), Inference Mode=Sparse Attention, Receptive Field=2562025.11 | 81.6 | 33.6 | 5.8 | 2.2 | 30.8 | |
| SSATraining Protocol=Length Extrapolation (RoPE Scaling), Inference Mode=Sparse Attention, Receptive Field=10242025.11 | 77.4 | 47.8 | — | 3.8 | 34.8 | |
| NSATraining Protocol=Continual Trained to 32k, Inference Mode=Sparse Attention, Receptive Field=10242025.11 | 74.4 | 27.8 | 10.6 | 6.6 | 29.9 | |
| MoBATraining Protocol=Length Extrapolation (RoPE Scaling), Inference Mode=Sparse Attention, Receptive Field=2562025.11 | 68.8 | 28.2 | 14.4 | 0 | 27.9 | |
| NSATraining Protocol=Length Extrapolation (RoPE Scaling), Inference Mode=Sparse Attention, Receptive Field=10242025.11 | 64.6 | 25 | 11.4 | 4 | 26.3 | |
| NSATraining Protocol=Continual Trained to 32k, Inference Mode=Sparse Attention, Receptive Field=2562025.11 | 61 | 34.2 | 5.4 | 4.8 | 26.4 | |
| FullAttnTraining Protocol=Continual Trained to 32k, Inference Mode=Sparse Attention, Receptive Field=10242025.11 | 59 | 25 | 9.2 | 5.2 | 24.6 | |
| MoBATraining Protocol=Length Extrapolation (RoPE Scaling), Inference Mode=Sparse Attention, Receptive Field=10242025.11 | 58.4 | 27.8 | 16 | 8 | 27.6 | |
| FullAttnTraining Protocol=Length Extrapolation (RoPE Scaling), Inference Mode=Sparse Attention, Receptive Field=10242025.11 | 58.2 | 21.4 | 7.2 | 4.4 | 22.8 | |
| FullAttnTraining Protocol=Continual Trained to 32k, Inference Mode=Sparse Attention, Receptive Field=2562025.11 | 53 | 11.8 | 5.4 | 3.6 | 18.5 | |
| FullAttnTraining Protocol=Length Extrapolation (RoPE Scaling), Inference Mode=Sparse Attention, Receptive Field=2562025.11 | 48 | 9.2 | 4.2 | 2.8 | 16.1 | |
| NSATraining Protocol=Length Extrapolation (RoPE Scaling), Inference Mode=Sparse Attention, Receptive Field=2562025.11 | 37.4 | 13.8 | 3.6 | 2.4 | 14.3 |