Understandability on Understandability Experiment New word strategy
7Significance CountClaude
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| ClaudeModel Identifier=claude-sonnet-4-52026.05 | 7 | 0.9 | — | — | |
| Majority Vote (MUM)Evaluation Protocol=Majority fit estimation2026.05 | 7 | — | — | — | |
| GPT-4o-mModel Identifier=gpt-4o-mini2026.05 | 6 | 0.46 | — | — | |
| GPT-4oModel Identifier=gpt-4o2026.05 | 5 | 0.67 | — | — | |
| LlamaModel Identifier=llama-3.1-8b-instant2026.05 | 5 | 0.81 | — | — | |
| GrokModel Identifier=grok-4.1-fast2026.05 | 5 | 0.35 | — | — | |
| MistralModel Identifier=Ministral-8B-Instruct-24102026.05 | 3 | 0.87 | — | — | |
| QwenModel Identifier=Qwen3-VL-32B-Instruct-FP82026.05 | 3 | 0.87 | — | — |