Long-context Understanding on Open Long Context Benchmarks
55.77Loogle ScoreDeepSeek-V3.1
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| DeepSeek-V3.12026.04 | 55.77 | 52.88 | 50.55 | 56.27 | 42.97 | 46.62 | 50.84 | |
| DeepSeek-R1-distill-32B + Anchor-based Reasoning (AbR) Framework + LoongRLBackbone=DeepSeek-R1-distill-32B, Strategy=Ours+LoongRL2026.04 | 55.59 | 51.29 | 44.45 | 72.27 | 67.01 | 36.12 | 54.46 | |
| Gemini-3-pro2026.04 | 52.86 | 69.38 | 65.43 | 88.07 | 83.01 | 75.3 | 72.34 | |
| Qwen3-235B-A22B-thinking-25072026.04 | 52.77 | 49.3 | 53.7 | 50.76 | 46.81 | 44.61 | 49.66 | |
| Kimi-K2-Thinking2026.04 | 51.5 | 49.3 | 58.01 | 58.1 | 49.98 | 51.77 | 53.11 | |
| DeepSeek-R1-distill-32B + Anchor-based Reasoning (AbR) FrameworkBackbone=DeepSeek-R1-distill-32B, Strategy=Ours2026.04 | 50.59 | 49.7 | 44.68 | 73.09 | 69.38 | 36.74 | 54.03 | |
| QwenLong-L1-32B2026.04 | 49.32 | 43.74 | 44.68 | 69.93 | 47.7 | 27.7 | 47.18 | |
| DeepSeek-R1-distill-32B+LoongRLBackbone=DeepSeek-R1-distill-32B, Strategy=LoongRL2026.04 | 48.41 | 45.73 | 40.29 | 70.74 | 61.84 | 37.23 | 50.71 | |
| DeepSeek-R1-distill-32B+QwenDocqaBackbone=DeepSeek-R1-distill-32B, Strategy=QwenDocqa2026.04 | 47.96 | 46.32 | 39.98 | 71.56 | 67.45 | 34.59 | 51.31 | |
| DeepSeek-R1-distill-32B+LongReasonBackbone=DeepSeek-R1-distill-32B, Strategy=LongReason2026.04 | 43.23 | 44.14 | 38.57 | 57.9 | 58.36 | 32.85 | 45.84 | |
| DeepSeek-R1-distill-32BBackbone=DeepSeek-R1-distill-32B2026.04 | 42.31 | 43.94 | 38.17 | 64.22 | 57.23 | 31.94 | 46.3 |