ResearchBenchmarksLong-context Question Answering on MultiFieldQA-en LongBench (test)Follow6String MatchM+3.22323.94414.6655.3859Jul 7, 2026Evaluation ResultsMethodMethodLinksString MatchF1 ScoreROUGE-LBERTScoreM+Backbone=Llama-3.1-8B,...Backbone=Llama-3.1-8B, Max Length=16,384 tokens2026.07614.1114.4185.06MemDefragBackbone=Llama-3.1-8B,...Backbone=Llama-3.1-8B, Max Length=16,384 tokens, Top-K=22026.07615.7415.5185.37MemoryLLMBackbone=Llama-3.1-8B,...Backbone=Llama-3.1-8B, Max Length=16,384 tokens2026.073.3314.7614.8385.92