Multi-hop Question Answering
Benchmarks
Dataset NameSOTA methodMetricTrendResultsLast Updated
34.7F1 Score
4
Apr 10, 2026
0.413F1 Score
4
Apr 10, 2026
75.6F1 Score
4
Apr 10, 2026
71Answer Correctness
4
Mar 17, 2026
90.9Human Evaluation Score
4
Mar 13, 2026
91Human Evaluation Score
4
Mar 13, 2026
33.64EM Score
4
Mar 4, 2026
38.5EM
4
Mar 4, 2026
61.4EM
4
Mar 4, 2026
80Avg F1
4
Feb 26, 2026
98Avg F1 Score
4
Feb 26, 2026
0.885Answer EM
4
Feb 26, 2026
55.6Retrieval
4
Feb 26, 2026
68.1EM
4
Feb 26, 2026
40.2GPT-Acc
3
Jun 30, 2026
72.3GPT-Acc
3
Jun 30, 2026
68.1GPT-Acc
3
Jun 30, 2026
23.56F1 Score
3
Jun 1, 2026
71.7Average Judge Score
3
May 27, 2026
30.2EM
3
May 27, 2026
53.4Exact Match (EM)
3
May 27, 2026
71.3EM
3
May 27, 2026
73.3Score (avg@3)
3
May 22, 2026
90Score (Avg@3)
3
May 22, 2026
0.39EM
3
May 14, 2026