Sentence Representation Evaluation on SentEval (test)
86.03MR AccuracyBERT-large + SG-OPT
Evaluation Results
| Method | Links | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| BERT-large + SG-OPTBase Model=BERT-large, Pooling Strategy=SG-OPT, Evaluation protocol=linear probing2021.06 | 86.03 | 90.18 | 95.82 | 87.08 | 90.73 | 94.65 | — | — | 73.31 | 88.26 | |
| SBERT-NLI-largeTraining Data=NLI, Model Variant=large2019.08 | 84.88 | 90.07 | 94.52 | 90.33 | 90.66 | 87.4 | — | — | 75.94 | 87.69 | |
| BERT-large + Mean poolingBase Model=BERT-large, Pooling Strategy=Mean pooling, Evaluation protocol=linear probing2021.06 | 84.38 | 89.01 | 95.6 | 86.69 | 89.2 | 90.9 | — | — | 72.79 | 86.94 | |
| SBERT-NLI-baseTraining Data=NLI, Model Variant=base2019.08 | 83.64 | 89.43 | 94.39 | 89.86 | 88.96 | 89.6 | — | — | 76 | 87.41 | |
| SBERT-base + SG-OPTBase Model=SBERT-base, Pooling Strategy=SG-OPT, Evaluation protocol=linear probing2021.06 | 83.34 | 89.45 | 94.68 | 89.78 | 88.57 | 87.3 | — | — | 75.26 | 86.91 | |
| SBERT-base + WK poolingBase Model=SBERT-base, Pooling Strategy=WK pooling, Evaluation protocol=linear probing2021.06 | 82.96 | 89.33 | 95.13 | 90.56 | 88.1 | 91.98 | — | — | 76.66 | 87.82 | |
| SBERT-base + Mean poolingBase Model=SBERT-base, Pooling Strategy=Mean pooling, Evaluation protocol=linear probing2021.06 | 82.8 | 89.03 | 94.07 | 89.79 | 88.08 | 86.93 | — | — | 75.11 | 86.54 | |
| BERT-large + WK poolingBase Model=BERT-large, Pooling Strategy=WK pooling, Evaluation protocol=linear probing2021.06 | 82.68 | 87.92 | 95.32 | 87.25 | 87.81 | 91.18 | — | — | 70.13 | 86.04 | |
| DiscoveryBigN (millions)=3.4, Supervision=Unsupervised2019.03 | 82.6 | 87.4 | 94.5 | 91 | 85.2 | 93.4 | 86.4 | 84.8 | 76.6 | 86.9 | |
| MTLN (millions)=124, Supervision=Supervised2019.03 | 82.5 | 87.7 | 94 | 90.9 | 83.2 | 93 | 88.8 | 87.8 | 78.6 | 87.4 | |
| DiscoveryBaseN (millions)=1.7, Supervision=Unsupervised2019.03 | 82.5 | 86.3 | 94.2 | 90.5 | 85.2 | 91.8 | 85.7 | 84 | 75.8 | 86.2 | |
| BERT-base + SG-OPTBase Model=BERT-base, Pooling Strategy=SG-OPT, Evaluation protocol=linear probing2021.06 | 82.47 | 87.42 | 95.4 | 88.92 | 86.2 | 91.6 | — | — | 74.21 | 86.6 | |
| DiscoveryHardN (millions)=1.7, Supervision=Unsupervised2019.03 | 81.6 | 86.5 | 93.9 | 90.5 | 84.8 | 90 | 85.4 | 83.2 | 76.5 | 85.8 | |
| InferSent - GloVeEmbeddings=GloVe2019.08 | 81.57 | 86.54 | 92.5 | 90.38 | 84.18 | 88.2 | — | — | 75.77 | 85.59 | |
| BERT-base + Mean poolingBase Model=BERT-base, Pooling Strategy=Mean pooling, Evaluation protocol=linear probing2021.06 | 81.46 | 86.71 | 95.37 | 87.9 | 85.83 | 90.3 | — | — | 73.36 | 85.85 | |
| DiscoveryAdvN (millions)=1.4, Supervision=Unsupervised2019.03 | 81.4 | 85.8 | 93.8 | 90.5 | 83.4 | 92 | 86 | 84.3 | 75.7 | 85.9 | |
| DiscoveryShuffledN (millions)=1.7, Supervision=Unsupervised2019.03 | 81.4 | 86.1 | 94.1 | 90.9 | 85.3 | 90.4 | 85.6 | 83.6 | 75.4 | 85.9 | |
| QuickThoughtN (millions)=174, Supervision=Unsupervised2019.03 | 81.3 | 84.5 | 94.6 | 89.5 | — | 92.4 | 87.1 | — | 75.9 | — | |
| Discovery10N (millions)=1.7, Supervision=Unsupervised2019.03 | 81.2 | 85.1 | 93.7 | 90.2 | 83 | 90 | 85.9 | 83.8 | 75.8 | 85.4 | |
| InferSentN (millions)=1, Supervision=Supervised2019.03 | 81.1 | 86.3 | 92.4 | 90.2 | 84.6 | 88.2 | 88.4 | 86.1 | 76.2 | 85.9 | |
| BERT-base + WK poolingBase Model=BERT-base, Pooling Strategy=WK pooling, Evaluation protocol=linear probing2021.06 | 80.64 | 85.53 | 95.27 | 88.63 | 85.03 | 94.03 | — | — | 71.71 | 85.83 | |
| DisSentN (millions)=4.7, Supervision=Unsupervised2019.03 | 80.1 | 84.9 | 93.6 | 90.1 | 84.1 | 93.6 | 84.9 | 83.7 | 75 | 85.6 | |
| Universal Sentence Encoder2019.08 | 80.09 | 85.19 | 93.98 | 86.7 | 86.38 | 93.2 | — | — | 70.14 | 85.1 | |
| BERT CLS-vectorPooling=CLS-token, Embeddings=BERT2019.08 | 78.68 | 84.85 | 94.21 | 88.23 | 84.13 | 91.4 | — | — | 71.13 | 84.66 | |
| Avg. BERT embeddingsPooling=Average, Embeddings=BERT2019.08 | 78.66 | 86.25 | 94.37 | 88.66 | 84.4 | 92.8 | — | — | 69.45 | 84.94 | |
| Avg. fast-text embeddingsPooling=Average, Embeddings=fast-text2019.08 | 77.96 | 79.23 | 91.68 | 87.81 | 82.15 | 83.6 | — | — | 74.49 | 82.42 | |
| Avg. GloVe embeddingsPooling=Average, Embeddings=GloVe2019.08 | 77.25 | 78.3 | 91.17 | 87.85 | 80.18 | 83 | — | — | 72.87 | 81.52 | |
| SkipThoughtN (millions)=74, Supervision=Unsupervised2019.03 | 76.5 | 80.1 | 93.6 | 87.1 | 82 | 92.2 | 85.8 | 82.3 | 73 | 83.6 |