ResearchBenchmarksUnderstandability on Understandability Experiment Paraphrase strategyFollow7Significance Count (out of 7)Claude2.843.9256.08May 7, 2026Evaluation ResultsMethodMethodLinksSignificance Count (out of 7)Adjusted R2Spearman CorrelationMajority FitClaude2026.0570.63——GPT-4o-m2026.0560.67——Majority Vote (MUM)2026.056———GPT-4o2026.0550.67——Llama2026.0550.98——Grok2026.0550.54——Mistral2026.0531——Qwen2026.0531——