PMI ranking estimation on Words (500 held-out pairs)
0.89Spearman RhoDecomposed PMI (emp.)
Evaluation Results
| Method | Links | |
|---|---|---|
| Decomposed PMI (emp.)Model=Claude Sonnet 4, Evaluation Type=Empirical marginal (oracle upper bound)2026.05 | 0.89 | |
| InfoNCE (emp.)Model=Claude Sonnet 4, Evaluation Type=Empirical marginal (oracle upper bound)2026.05 | 0.89 | |
| InfoNCE (emp.)Model=GPT-5.2, Evaluation Type=Empirical marginal (oracle upper bound)2026.05 | 0.87 | |
| PromptNCE (emp.)Model=GPT-5.2, Evaluation Type=Empirical marginal (oracle upper bound)2026.05 | 0.85 | |
| PromptNCE (emp.)Model=Claude Sonnet 4, Evaluation Type=Empirical marginal (oracle upper bound)2026.05 | 0.85 | |
| Decomposed PMI (emp.)Model=GPT-5.2, Evaluation Type=Empirical marginal (oracle upper bound)2026.05 | 0.82 | |
| PromptNCEModel=GPT-5.2, Evaluation Type=Zero-shot2026.05 | 0.74 | |
| PromptNCEModel=Claude Sonnet 4, Evaluation Type=Zero-shot2026.05 | 0.69 | |
| MarginalNCEModel=GPT-5.2, Evaluation Type=Zero-shot2026.05 | 0.64 | |
| MarginalNCEModel=Claude Sonnet 4, Evaluation Type=Zero-shot2026.05 | 0.64 | |
| Decomposed PMIModel=Claude Sonnet 4, Evaluation Type=Zero-shot2026.05 | 0.63 | |
| Decomposed PMIModel=GPT-5.2, Evaluation Type=Zero-shot2026.05 | 0.57 | |
| Direct PMIModel=Claude Sonnet 4, Evaluation Type=Zero-shot2026.05 | 0.48 | |
| Direct PMIModel=GPT-5.2, Evaluation Type=Zero-shot2026.05 | 0.36 | |
| InfoNCEModel=GPT-5.2, Evaluation Type=Zero-shot2026.05 | 0.35 | |
| InfoNCEModel=Claude Sonnet 4, Evaluation Type=Zero-shot2026.05 | 0.33 |