Long Story Evaluation on LongStoryEval 128K token context
4.4Plot QualityDeepSeek-v2.5
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| DeepSeek-v2.5Evaluation Strategy=One-Pass2025.12 | 4.4 | 3.5 | 4.8 | -0.9 | 3.3 | -1.1 | -1.3 | 9.4 | 4.8 | |
| GPT-4oEvaluation Strategy=One-Pass2025.12 | 3.3 | 4.1 | 7.9 | 0.8 | 3.3 | -1.2 | -3.2 | 8.4 | 5.5 |