Classification on news20 (test)
1,494.4CPU ScorePG-LL
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| PG-LLSparsity level (s)=0.05m, Stopping threshold (epsilon)=1e-6, Max iterations reached=true2022.11 | 1,494.4 | 20,000 | 0 | 92.2 | — | |
| PGSparsity level (s)=2m2022.11 | 904.5 | 10,000 | 0 | 96.6 | — | |
| PGSparsity level (s)=1.5m2022.11 | 885 | 10,000 | 0 | 96.4 | — | |
| APGSparsity level (s)=0.05m, Stopping threshold (epsilon)=1e-62022.11 | 758.3 | 8,428 | 0 | 92.3 | — | |
| PGSparsity level (s)=0.01m, Stopping threshold (epsilon)=1e-6, Max iterations reached=true2022.11 | 738.7 | 10,000 | 0 | 87.7 | — | |
| PGSparsity level (s)=0.05m, Stopping threshold (epsilon)=1e-6, Max iterations reached=true2022.11 | 728.9 | 10,000 | 0 | 93.5 | — | |
| PG-LLSparsity level (s)=0.01m, Stopping threshold (epsilon)=1e-62022.11 | 366.7 | 4,682 | 0 | 87.3 | — | |
| APGSparsity level (s)=2m2022.11 | 217.1 | 2,170 | 0 | 96.4 | — | |
| APGSparsity level (s)=1.5m2022.11 | 208.2 | 2,072 | 0 | 96.4 | — | |
| PG-LLSparsity level (s)=1.5m2022.11 | 155.8 | 1,736 | 0 | 96.3 | — | |
| PG-LLSparsity level (s)=2m2022.11 | 153.1 | 1,700 | 0 | 96.3 | — | |
| APGSparsity level (s)=0.01m, Stopping threshold (epsilon)=1e-62022.11 | 151.7 | 1,583 | 0 | 87.7 | — | |
| APG+Sparsity level (s)=2m2022.11 | 86.4 | 875 | 18 | 96.7 | — | |
| APG-LL+Sparsity level (s)=2m2022.11 | 80.4 | 846 | 2 | 96.2 | — | |
| APG+Sparsity level (s)=1.5m2022.11 | 78.3 | 826 | 17 | 96.7 | — | |
| APG-LL+Sparsity level (s)=1.5m2022.11 | 64.6 | 690 | 6 | 96.2 | — | |
| APG-LL+Sparsity level (s)=0.05m, Stopping threshold (epsilon)=1e-62022.11 | 29.2 | 417 | 89 | 92 | — | |
| APG+Sparsity level (s)=0.05m, Stopping threshold (epsilon)=1e-62022.11 | 16.1 | 171 | 67 | 92.3 | — | |
| APG-LL+Sparsity level (s)=0.01m, Stopping threshold (epsilon)=1e-62022.11 | 6.6 | 152 | 88 | 85.4 | — | |
| APG+Sparsity level (s)=0.01m, Stopping threshold (epsilon)=1e-62022.11 | 5 | 52 | 63 | 85.3 | — | |
| BERT_BASE AdaptersBackbone=BERT_BASE, Evaluation Protocol=Adapter-tuning, Total number of parameters=1.19x, Trained parameters per task=1.14%2019.02 | — | — | — | — | 96.2 | |
| BERT_BASE Fine-tuneBackbone=BERT_BASE, Evaluation Protocol=Fine-tuning, Total number of parameters=17x, Trained parameters per task=100%2019.02 | — | — | — | — | 96.3 | |
| BERT_BASE Variable FTBackbone=BERT_BASE, Evaluation Protocol=Variable Fine-tuning, Total number of parameters=9.9x, Trained parameters per task=52.9%2019.02 | — | — | — | — | 96.5 | |
| No BERT baseline2019.02 | — | — | — | — | 95.2 |