Explanation quality evaluation
Benchmarks
Dataset NameSOTA methodMetricTrendResultsLast Updated
2.2Relevance Rank Accuracy (FeatPerm)
11
May 12, 2026
37.2Rank Acc (FeatPerm)
11
May 12, 2026
37.7RRA (FeatPerm)
11
May 12, 2026
4.42CAC
9
May 11, 2026
2.29Meaningfulness Score
7
Apr 21, 2026
2.07M Score
7
Apr 21, 2026
1.53ChatGPT Meaningfulness Score
7
Apr 21, 2026
87.6Helpfulness
6
Feb 26, 2026
80.8Helpfulness
6
Feb 26, 2026
7.55GPT-4o Score
3
Apr 16, 2026
44Reasoning Soundness Loss (%)
2
Mar 31, 2026
39.7Reasoning Soundness Loss
2
Mar 31, 2026