Long Context Benchmarks
Benchmarks
Task NameDataset NameSOTA ResultTrendResults
Long-context Benchmarks 100K context LB-V2 DocMath Frames LB-MQA (test)
66.7DocMath Score
36
Long-context Benchmarks 16K context DocMath Frames LB-MQA V2 (test)
64.1DocMath
36
Long-context benchmarks
52.8Accuracy (8k Context)
21
Long-context benchmarks
38.5Score (8k Context)
21
Long-context benchmarks
50.5Performance (8k Context)
21
Long-context benchmarks
100Synthetic Recall (8k context)
21
Long-context benchmarks
53.7RAG Score (8k Context)
16
Long Context Benchmarks
32.3MDQA-10 Score
5