Speech Processing on SUPERB (test)
97.34KS AccuracyM2D/0.6 T=4.00s
Evaluation Results
| Method | Links | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| M2D/0.6 T=4.00sMasking Ratio=0.6, Input Duration (T)=4.00s, Number of Parameters (Million)=86M, Pre-training Dataset=LS-960+AS, Patch Size=80 x 42024.04 | 97.34 | 8.5 | — | 94.83 | — | — | — | 81.24 | 7.13 | 63.81 | 74.11 | 52.23 | |
| M2D/0.6 T=5.12sMasking Ratio=0.6, Input Duration (T)=5.12s, Number of Parameters (Million)=86M, Pre-training Dataset=LS-960+AS, Patch Size=80 x 42024.04 | 97.17 | 8.1 | — | 94.7 | — | — | — | 78.73 | 7.15 | 65.47 | 74.69 | 51.66 | |
| M2D/0.6 T=6.08sMasking Ratio=0.6, Input Duration (T)=6.08s, Number of Parameters (Million)=86M, Pre-training Dataset=LS-960+AS, Patch Size=80 x 42024.04 | 97.17 | 7.74 | — | 95.5 | — | — | — | 80.48 | 6.84 | 64.06 | 72.42 | 50.64 | |
| M2D-S/0.6 T=6.08sMasking Ratio=0.6, Input Duration (T)=6.08s, Number of Parameters (Million)=86M, Pre-training Dataset=LS-960+AS, Patch Size=80 x 22024.04 | 96.87 | 5.33 | — | 97.65 | — | — | — | 80.69 | 7.07 | 66.13 | 54.77 | 43.75 | |
| M2D-S/0.6 T=4.0sMasking Ratio=0.6, Input Duration (T)=4.0s, Number of Parameters (Million)=86M, Pre-training Dataset=LS-960+AS, Patch Size=80 x 22024.04 | 96.8 | 5.72 | — | 97.63 | — | — | — | 81.74 | 5.97 | 66.36 | 53.22 | 41.71 | |
| WavLMpre-trained SSL model=frozen, weighted-summing all layers=true, iterative offline clustering=true, large batch size=true, hyper-parameter sweep=true2023.05 | 96.79 | 4.84 | 6.31 | 98.63 | 89.38 | 22.86 | 20.74 | — | — | — | — | — | |
| WavLM BaseNumber of Parameters (Million)=95M, Pre-training Dataset=LS-960+DNS2024.04 | 96.79 | 4.84 | — | 98.63 | — | — | — | 84.51 | 4.69 | 65.94 | 54.45 | 40.98 | |
| CCC-wav2vec 2.0pre-trained SSL model=frozen, weighted-summing all layers=true2023.05 | 96.72 | 5.95 | 6.3 | 96.47 | 88.08 | 24.34 | 16.2 | — | — | — | — | — | |
| SSAST-FrameNumber of Parameters (Million)=89M, Pre-training Dataset=LS-960 U AS2024.04 | 96.7 | — | — | 80.8 | — | — | — | — | — | 60.5 | — | — | |
| DinoSRpre-trained SSL model=frozen, weighted-summing all layers=true, runs per task=no more than five2023.05 | 96.69 | 3.21 | 4.71 | 98.02 | 88.83 | 23.57 | 17.68 | — | — | — | — | — | |
| data2vecpre-trained SSL model=frozen, weighted-summing all layers=true2023.05 | 96.56 | 4.69 | 4.94 | 97.63 | 88.59 | 25.27 | 17.42 | — | — | — | — | — | |
| M2D-S/0.6 T=5.12sMasking Ratio=0.6, Input Duration (T)=5.12s, Number of Parameters (Million)=86M, Pre-training Dataset=LS-960+AS, Patch Size=80 x 22024.04 | 96.47 | 5.64 | — | 97.8 | — | — | — | 81.97 | 6.29 | 65.35 | 57.34 | 43.23 | |
| HuBERTpre-trained SSL model=frozen, weighted-summing all layers=true, iterative offline clustering=true2023.05 | 96.3 | 5.41 | 6.42 | 98.34 | 88.53 | 25.2 | 15.53 | — | — | — | — | — | |
| HuBERT Base (Baseline)Number of Parameters (Million)=95M, Pre-training Dataset=LS-9602024.04 | 96.3 | 5.41 | — | 98.34 | — | — | — | 81.42 | 5.11 | 64.92 | 62.76 | 46.26 | |
| wav2vec 2.0pre-trained SSL model=frozen, weighted-summing all layers=true2023.05 | 96.23 | 5.74 | 6.43 | 92.35 | 88.3 | 24.77 | 14.81 | — | — | — | — | — | |
| wav2vec2.0 BaseNumber of Parameters (Million)=95M, Pre-training Dataset=LS-9602024.04 | 96.23 | 5.74 | — | 92.35 | — | — | — | 75.18 | 6.02 | 63.43 | 37.66 | 32.02 | |
| M2D/0.6, T=6.08s, patch size 16 x 16Masking Ratio=0.6, Input Duration (T)=6.08s, Patch Size=16 x 16, Number of Parameters (Million)=86M, Pre-training Dataset=AS2024.04 | 95.65 | 78.3 | — | 76.77 | — | — | — | 80.68 | — | 61.17 | 88.63 | 66.56 | |
| SSAST-PatchNumber of Parameters (Million)=89M, Pre-training Dataset=LS-960 U AS2024.04 | 94.8 | — | — | 57.1 | — | — | — | — | — | 56.8 | — | — |