Word sense plausibility rating on AmbiStory (test)
0.731Spearman Correlation (ρ)GPT-4o
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPT-4oApproach=Prompting, Prompting Strategy=P2 (structured prompting with decision rules)2026.03 | 0.731 | 79.4 | |
| GPT-4.1Approach=Prompting, Prompting Strategy=P2 (structured prompting with decision rules)2026.03 | 0.722 | 76.7 | |
| GPT-5.2Approach=Prompting, Prompting Strategy=P2 (structured prompting with decision rules)2026.03 | 0.717 | 76 | |
| GPT-5 miniApproach=Prompting, Prompting Strategy=P2 (structured prompting with decision rules)2026.03 | 0.696 | 74.3 | |
| GPT-5.2Approach=Prompting, Prompting Strategy=P1 (few-shot prompting)2026.03 | 0.635 | 71.3 | |
| ELECTRA-large + LoRAApproach=Fine-tune2026.03 | 0.527 | 63.9 | |
| DeBERTa-large + LoRAApproach=Fine-tune2026.03 | 0.492 | 67.6 | |
| ELECTRA-baseApproach=Fine-tune2026.03 | 0.482 | 62.5 | |
| Ministral-3-8BApproach=Prompting, Prompting Strategy=P2 (structured prompting with decision rules)2026.03 | 0.472 | 59.4 | |
| DeBERTa-large + LoRA + uncertainty lossApproach=Fine-tune2026.03 | 0.435 | 65.9 | |
| Llama-3.2-3BApproach=Prompting, Prompting Strategy=P2 (structured prompting with decision rules)2026.03 | 0.134 | 52.2 | |
| RoBERTa + XGBoostApproach=Embedding2026.03 | 0.133 | 52.2 | |
| MPNet + RidgeApproach=Embedding2026.03 | 0.109 | 51.3 |