Understandability on Understandability Experiment Unknown word strategy
7Significance CountClaude
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| ClaudeModel Identifier=claude-sonnet-4-5, Evaluation Protocol=Spearman Rank Correlation2026.05 | 7 | -1.03 | — | — | |
| GPT-4o-mModel Identifier=gpt-4o-mini, Evaluation Protocol=Spearman Rank Correlation2026.05 | 6 | -1.94 | — | — | |
| GPT-4oModel Identifier=gpt-4o, Evaluation Protocol=Spearman Rank Correlation2026.05 | 5 | -2.46 | — | — | |
| LlamaModel Identifier=llama-3.1-8b-instant, Evaluation Protocol=Spearman Rank Correlation2026.05 | 5 | 0.34 | — | — | |
| GrokModel Identifier=grok-4.1-fast, Evaluation Protocol=Spearman Rank Correlation2026.05 | 5 | -2.39 | — | — | |
| MistralModel Identifier=Ministral-8B-Instruct-2410, Evaluation Protocol=Spearman Rank Correlation2026.05 | 3 | 0.49 | — | — | |
| QwenModel Identifier=Qwen3-VL-32B-Instruct-FP8, Evaluation Protocol=Spearman Rank Correlation2026.05 | 3 | 0.47 | — | — | |
| Majority Vote (MUM)Evaluation Protocol=Majority fit estimation2026.05 | 1 | — | — | — |