Long-context understanding on RULER
96ScoreQwen3.5-35B-A3B
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3.5-35B-A3B#Token=<1002026.04 | 96 | |
| JoyAI-LLM Flash#Token=<1002026.04 | 95.7 | |
| Qwen2.5-7BSoftmax Ratio=100%2025.12 | 94.45 | |
| Gemini-1.5-ProLength=128k2024.08 | 94.4 | |
| Qwen3-Next-80B-A3B#Token=1002026.04 | 94.2 | |
| Qwen3-30B-A3B#Token=<1002026.04 | 93.7 | |
| GA-S2Teacher Model=Qwen2.5-7B, Softmax Ratio=33%2025.12 | 91.1 | |
| SMARTTeacher Model=Qwen2.5-7B, Softmax Ratio=33%2025.12 | 89.49 | |
| Qwen2.5-1.5BSoftmax Ratio=100%2025.12 | 87.42 | |
| GA-S2Teacher Model=Qwen2.5-7B, Softmax Ratio=25%2025.12 | 85.84 | |
| SMARTTeacher Model=Qwen2.5-7B, Softmax Ratio=25%2025.12 | 81.58 | |
| GPT-4Length=128k2024.08 | 81.2 | |
| GLM4-9B-Chat-1MLength=128k2024.08 | 79.9 | |
| all-attentionSpeedup @32k=1.0x, Speedup @16k=1.0x2026.04 | 79.4 | |
| S1: Distil. Idealized|All–6Speedup @32k=6.13x, Speedup @16k=2.5x2026.04 | 78.7 | |
| Gradient-Llama3-8BLength=128k2024.08 | 78.4 | |
| Llama3.1-8B-InstructLength=128k2024.08 | 77.7 | |
| Yi-34B-200kLength=128k2024.08 | 77.3 | |
| LongRecipeModel=Llama3-8B-I, Length=128k2024.08 | 76 | |
| FLTModel=Llama3-8B-I, Length=128k2024.08 | 75.8 | |
| FLTModel=Llama3-8B-I, Length=80k2024.08 | 75.7 | |
| POSEModel=Llama3-8B-I, Length=128k2024.08 | 75.3 | |
| GLM-4.7-Flash-T#Token=83002026.04 | 74.7 | |
| LongRecipeModel=Llama3-8B-I, Length=80k2024.08 | 74.5 | |
| Reg|Lklhd–26Speedup @32k=2.85x, Speedup @16k=1.5x2026.04 | 74.4 | |
| Gradient-Llama3-70BLength=128k2024.08 | 72.1 | |
| Falcon-H1R 7BSpeedup @32k=4.61x, Speedup @16k=3.4x2026.04 | 72 | |
| RPESModel=Llama3-8B-I, Length=128k2024.08 | 71.5 | |
| RPESModel=Llama3-8B-I, Length=80k2024.08 | 71.4 | |
| Apriel-H1 15BSpeedup @32k=1.97x, Speedup @16k=1.9x2026.04 | 71 | |
| LongRecipeModel=Qwen2-7B, Length=80k2024.08 | 70.8 | |
| POSEModel=Llama3-8B-I, Length=80k2024.08 | 69.9 | |
| Llama3.1-8BLength=128k2024.08 | 69.8 | |
| GA-S2Teacher Model=Qwen2.5-1.5B, Softmax Ratio=33%2025.12 | 69.53 | |
| FLT*Model=Qwen2-7B, Length=80k2024.08 | 69.5 | |
| Apriel-1.6Speedup @32k=1.0x, Speedup @16k=1.0x2026.04 | 69.1 | |
| RPESModel=Qwen2-7B, Length=80k2024.08 | 68.9 | |
| LongRecipeModel=Mistral-7B, Length=80k2024.08 | 67.2 | |
| Idealized|All–18Speedup @32k=1.99x, Speedup @16k=1.1x2026.04 | 67.1 | |
| POSEModel=Qwen2-7B, Length=80k2024.08 | 66.7 | |
| Llama3.1-70B-InstructLength=128k2024.08 | 66.6 | |
| Idealized|Lklhd–6Speedup @32k=6.2x, Speedup @16k=2.4x2026.04 | 66.1 | |
| OLMo-Hybrid-Think 7BSpeedup @32k=2.51x, Speedup @16k=2.1x2026.04 | 65.8 | |
| Nemotron-Nano 12B v2Speedup @32k=5.85x, Speedup @16k=4.3x2026.04 | 65.2 | |
| RPESModel=Mistral-7B, Length=80k2024.08 | 65.1 | |
| POSEModel=Mistral-7B, Length=80k2024.08 | 65 | |
| Nemotron-3-Nano 30BSpeedup @32k=4.09x, Speedup @16k=2.8x2026.04 | 64.9 | |
| LongRecipeModel=Qwen2-7B, Length=128k2024.08 | 64.8 | |
| SMARTTeacher Model=Qwen2.5-1.5B, Softmax Ratio=33%2025.12 | 64.79 | |
| RPESModel=Qwen2-7B, Length=128k2024.08 | 64.6 | |
| Yi-9B-200kLength=128k2024.08 | 62.3 | |
| Idealized|All–6Speedup @32k=6.13x, Speedup @16k=2.5x2026.04 | 61.9 | |
| Reg|Lklhd–18Speedup @32k=4.76x, Speedup @16k=2.2x2026.04 | 60.5 | |
| POSEModel=Qwen2-7B, Length=128k2024.08 | 60.1 | |
| LongRecipeModel=Mistral-7B, Length=128k2024.08 | 58.2 | |
| FLTModel=Mistral-7B, Length=80k2024.08 | 57.4 | |
| Reg|Lklhd–13Speedup @32k=6.9x, Speedup @16k=2.7x2026.04 | 57 | |
| Qwen-3.5 27BSpeedup @32k=0.55x, Speedup @16k=0.5x2026.04 | 54.4 | |
| GA-S2Teacher Model=Qwen2.5-1.5B, Softmax Ratio=25%2025.12 | 54.08 | |
| Qwen2-72B-InstructLength=128k2024.08 | 53.7 | |
| RPESModel=Mistral-7B, Length=128k2024.08 | 52.5 | |
| Qwen2-7B-InstructLength=128k2024.08 | 52.5 | |
| FLT*Model=Qwen2-7B, Length=128k2024.08 | 51.3 | |
| SMARTTeacher Model=Qwen2.5-1.5B, Softmax Ratio=25%2025.12 | 50.98 | |
| Reg|Lklhd–10Speedup @32k=10.69x, Speedup @16k=4.2x2026.04 | 48.6 | |
| POSEModel=Mistral-7B, Length=128k2024.08 | 46.3 |