Word Similarity on RG-65
0.851Spearman CorrelationOurs (+post-process)
Evaluation Results
| Method | Links | |
|---|---|---|
| Ours (+post-process)Type=Context-to-Vec2022.10 | 0.851 | |
| Ours (preliminary)Type=Context-to-Vec2022.10 | 0.843 | |
| CBOWDim.=768, Source=trained by us2021.06 | 0.8348 | |
| SKIPGRAMDim.=768, Source=trained by us2021.06 | 0.8259 | |
| BERT (avg)Type=Contextualized, Representation level=avg2022.10 | 0.812 | |
| BERT2STATICparaParent Model=BERT-24, Dim.=10242021.06 | 0.8085 | |
| FASTTEXTType=Static2022.10 | 0.808 | |
| ROBERTA2STATICparaParent Model=ROBERTA-12, Dim.=7682021.06 | 0.8057 | |
| BERT2STATICsentParent Model=BERT-24, Dim.=10242021.06 | 0.8031 | |
| ROBERTA2STATICsentParent Model=ROBERTA-12, Dim.=7682021.06 | 0.7999 | |
| ROBERTA2STATICparaParent Model=ROBERTA-24, Dim.=10242021.06 | 0.7939 | |
| GPT22STATICparaParent Model=GPT2-24, Dim.=10242021.06 | 0.7907 | |
| GPT22STATICparaParent Model=GPT2-12, Dim.=7682021.06 | 0.7881 | |
| BERT+Skip-gramType=Context-to-Vec2022.10 | 0.786 | |
| GPT22STATICsentParent Model=GPT2-24, Dim.=10242021.06 | 0.7815 | |
| SENT2VECDim.=768, Source=trained by us2021.06 | 0.7811 | |
| ASE - best layer per taskParent Model=BERT-24, Dim.=1024, Layer=92021.06 | 0.7745 | |
| ASE - best task independent layerParent Model=BERT-24, Dim.=1024, Layer=72021.06 | 0.7677 | |
| ROBERTA2STATICsentParent Model=ROBERTA-24, Dim.=10242021.06 | 0.7677 | |
| FASTTEXTDim.=300, Size of training corpus relative to ours=12x2021.06 | 0.7669 | |
| BERT2STATICparaParent Model=BERT-12, Dim.=7682021.06 | 0.7555 | |
| Skip-gramType=Static2022.10 | 0.752 | |
| GPT22STATICsentParent Model=GPT2-12, Dim.=7682021.06 | 0.7484 | |
| ASE - best layer per taskParent Model=BERT-12, Dim.=768, Layer=12021.06 | 0.7449 | |
| BERT2STATICsentParent Model=BERT-12, Dim.=7682021.06 | 0.7421 | |
| ASE - best layer per taskParent Model=GPT2-12, Dim.=768, Layer=12021.06 | 0.7013 | |
| ASE - best overall layerParent Model=BERT-12, Dim.=768, Layer=32021.06 | 0.6948 | |
| W2GEmbedding Method=W2G, Training Corpus=ukWaC and WaCkypedia2018.10 | 0.69 | |
| ASE - best overall layerParent Model=GPT2-12, Dim.=768, Layer=22021.06 | 0.6833 | |
| ASE - best layer per taskParent Model=ROBERTA-24, Dim.=1024, Layer=82021.06 | 0.6782 | |
| ASE - best task independent layerParent Model=ROBERTA-24, Dim.=1024, Layer=62021.06 | 0.6738 | |
| ASE - best layer per taskParent Model=ROBERTA-12, Dim.=768, Layer=02021.06 | 0.673 | |
| ASE - best overall layerParent Model=ROBERTA-12, Dim.=768, Layer=02021.06 | 0.673 | |
| ASE - best layer per taskParent Model=GPT2-24, Dim.=1024, Layer=12021.06 | 0.6574 | |
| EllEmbedding Method=Ell, Training Corpus=ukWaC and WaCkypedia2018.10 | 0.65 | |
| GLOVEDim.=300, Size of training corpus relative to ours=650x2021.06 | 0.6442 | |
| GloVe+GTEmbedding Method=GloVe, Transform=Gaussian Transform, Refinement Corpus=text82018.10 | 0.62 | |
| GloVeEmbedding Method=GloVe2018.10 | 0.6 | |
| ASE - best task independent layerParent Model=GPT2-24, Dim.=1024, Layer=132021.06 | 0.5773 | |
| W2V*text8Embedding Method=W2V, Training Corpus=text82018.10 | 0.56 | |
| GloVe*text8Embedding Method=GloVe, Training Corpus=text82018.10 | 0.33 |