Long-context Conversational Question Answering
Benchmarks
Dataset NameSOTA methodMetricTrendResultsLast Updated
43.1Multi-Hop F1
59
May 21, 2026
1.94G-EVAL Score
3
May 1, 2026
3.24G-EVAL
3
May 1, 2026
3.08G-EVAL
3
May 1, 2026
2.89G-EVAL
3
May 1, 2026
2.98G-EVAL Score
3
May 1, 2026
2.81G-EVAL
3
May 1, 2026