Speech Separation on Libri2Mix (test)
22SI-SNRi (dB)SepReformer-M
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| SepReformer-MParams. (M)=17.3, MACs (G/s)=81.32024.06 | 22 | 22.2 | — | — | — | — | — | — | — | |
| SepReformer (M)# Params (M)=17.3, GMAC/s (G/s)=81.3, Model scale=M2025.07 | 22 | 22.2 | — | — | — | — | — | — | — | |
| MossFormer2Parameters (M)=55.72023.12 | 21.7 | — | — | — | — | — | — | — | — | |
| MossFormer2 + DM# Params (M)=55.7, GMAC/s (G/s)=84.2, dynamic mixing (+DM)=true2025.07 | 21.7 | — | — | — | — | — | — | — | — | |
| SepReformer-BParams. (M)=14.2, MACs (G/s)=39.82024.06 | 21.6 | 21.9 | — | — | — | — | — | — | — | |
| PRESS-12 @ 12 (M) + FT# Params (M)=22.4, GMAC/s (G/s)=79.7, Model scale=M, Exit point=12, finetuned (+FT)=true2025.07 | 21.29 | 21.68 | — | — | — | — | — | — | — | |
| PRESS-4 @ 4 (S) + FT# Params (M)=3.4, GMAC/s (G/s)=11.3, Model scale=S, Exit point=4, finetuned (+FT)=true2025.07 | 21.01 | 21.36 | — | — | — | — | — | — | — | |
| PRESS-12 @ 8 (M) + FT# Params (M)=15.6, GMAC/s (G/s)=54.4, Model scale=M, Exit point=8, finetuned (+FT)=true2025.07 | 20.92 | 21.33 | — | — | — | — | — | — | — | |
| PRESS-12 @ 12 (M)# Params (M)=22.4, GMAC/s (G/s)=79.7, Model scale=M, Exit point=122025.07 | 20.88 | 21.31 | — | — | — | — | — | — | — | |
| SepReformer-SParams. (M)=4.3, MACs (G/s)=21.32024.06 | 20.6 | 21 | — | — | — | — | — | — | — | |
| SepReformer (S)# Params (M)=4.5, GMAC/s (G/s)=21.3, Model scale=S2025.07 | 20.6 | 21 | — | — | — | — | — | — | — | |
| PRESS-12 @ 8 (M)# Params (M)=15.6, GMAC/s (G/s)=54.4, Model scale=M, Exit point=82025.07 | 20.42 | 20.86 | — | — | — | — | — | — | — | |
| SFSRNetParameters (M)=59.02023.12 | 20.4 | — | — | — | — | — | — | — | — | |
| PRESS-12 @ 4 (M) + FT# Params (M)=8.7, GMAC/s (G/s)=29.1, Model scale=M, Exit point=4, finetuned (+FT)=true2025.07 | 20.31 | 20.72 | — | — | — | — | — | — | — | |
| TF-GridNet(OA)Attention mechanism=Omni-directional Attention (OA)2026.01 | 20.1 | 20.2 | — | — | — | — | — | — | — | |
| PRESS-4 @ 4 (S)# Params (M)=3.4, GMAC/s (G/s)=11.3, Model scale=S, Exit point=42025.07 | 20.04 | 20.41 | — | — | — | — | — | — | — | |
| SPMamba(OA)Attention mechanism=Omni-directional Attention (OA)2026.01 | 20 | 20.2 | — | — | — | — | — | — | — | |
| SPMamba2026.01 | 19.9 | 20.4 | — | — | — | — | — | — | — | |
| TF-GridNetParameters=14.4M2026.01 | 19.8 | 20.1 | — | — | — | — | — | — | — | |
| PRESS-12 @ 4 (M)# Params (M)=8.7, GMAC/s (G/s)=29.1, Model scale=M, Exit point=42025.07 | 19.75 | 19.71 | — | — | — | — | — | — | — | |
| SepReformer-TParams. (M)=3.5, MACs (G/s)=10.42024.06 | 19.7 | 20.2 | — | — | — | — | — | — | — | |
| MossFormerParameters (M)=42.12023.12 | 19.7 | — | — | — | — | — | — | — | — | |
| SepReformer (T)# Params (M)=3.7, GMAC/s (G/s)=10.4, Model scale=T2025.07 | 19.7 | 20.2 | — | — | — | — | — | — | — | |
| WavesplitParameters (M)=292023.12 | 19.5 | — | — | — | — | — | — | — | — | |
| SepFormerParameters (M)=25.72023.12 | 19.2 | — | — | — | — | — | — | — | — | |
| SepFormer# Params (M)=26.0, GMAC/s (G/s)=86.92025.07 | 19.2 | 19.4 | — | — | — | — | — | — | — | |
| USE-BParameters(M)=8.1, Separation under an unknown and variable number of speakers=false2025.12 | 17.8 | — | — | — | — | — | — | — | — | |
| USE-BParameters(M)=8.1, Separation under an unknown and variable number of speakers=true2025.12 | 17.7 | — | — | — | — | — | — | — | — | |
| TDANet-WavParameters(M)=10.8, Separation under an unknown and variable number of speakers=false2025.12 | 17.5 | — | — | — | — | — | — | — | — | |
| TDANetParams. (M)=2.3, MACs (G/s)=9.12024.06 | 17.4 | 17.9 | — | — | — | — | — | — | — | |
| TDANet LargeParams (M)=2.3, kernel length=2ms, stride size=0.5ms2022.09 | 17.4 | 17.9 | — | — | — | — | — | — | — | |
| TDANet2026.01 | 17.4 | 17.9 | — | — | — | — | — | — | — | |
| UME ASR init. (0.33, 0.33, 0.34)Training set=LibriMix (460 hours), mixclean=speech + noise, RWSE=true, ASR initialization=true, weights (λasr, λdiar, λsep)=(0.33, 0.33, 0.34)2025.08 | 17.06 | 17.41 | — | — | — | — | — | — | 95.64 | |
| S4MParams. (M)=3.6, MACs (G/s)=38.42024.06 | 16.9 | 17.4 | — | — | — | — | — | — | — | |
| TDANetParams (M)=2.32022.09 | 16.9 | 17.4 | — | — | — | — | — | — | — | |
| DPTNetParams. (M)=2.7, MACs (G/s)=102.52024.06 | 16.7 | 17.1 | — | — | — | — | — | — | — | |
| A-FRCNNParams. (M)=6.1, MACs (G/s)=125.02024.06 | 16.7 | 17.2 | — | — | — | — | — | — | — | |
| DPTNetParams (M)=2.72022.09 | 16.7 | 17.1 | — | — | — | — | — | — | — | |
| A-FRCNN-16Params (M)=6.12022.09 | 16.7 | 17.2 | — | — | — | — | — | — | — | |
| A-FRCNN2026.01 | 16.7 | 17.2 | — | — | — | — | — | — | — | |
| WaveSplitParams. (M)=29.02024.06 | 16.6 | 17.2 | — | — | — | — | — | — | — | |
| SepformerParams. (M)=26.0, MACs (G/s)=86.92024.06 | 16.5 | 17 | — | — | — | — | — | — | — | |
| SepformerParams (M)=26.02022.09 | 16.5 | 17 | — | — | — | — | — | — | — | |
| DPRNNParams. (M)=2.6, MACs (G/s)=88.52024.06 | 16.1 | 16.6 | — | — | — | — | — | — | — | |
| DualPathRNNParams (M)=2.72022.09 | 16.1 | 16.6 | — | — | — | — | — | — | — | |
| DualPathRNN2026.01 | 16.1 | 16.6 | — | — | — | — | — | — | — | |
| DualPathRNN# Params (M)=2.6, GMAC/s (G/s)=42.52025.07 | 16.1 | 16.6 | — | — | — | — | — | — | — | |
| BSRNN-LargeParameters(M)=21.8, Separation under an unknown and variable number of speakers=false2025.12 | 15.2 | — | — | — | — | — | — | — | — | |
| Conv-TasNetParameters (M)=5.12023.12 | 14.7 | — | — | — | — | — | — | — | — | |
| SuDORM-RFParams. (M)=6.4, MACs (G/s)=10.12024.06 | 14 | 14.4 | — | — | — | — | — | — | — | |
| SuDORM-RF 2.5xParams (M)=6.42022.09 | 14 | 14.4 | — | — | — | — | — | — | — | |
| SudoRM-RF2026.01 | 14 | 14.4 | — | — | — | — | — | — | — | |
| SuDORM-RF 1.0xParams (M)=2.72022.09 | 13.5 | 14 | — | — | — | — | — | — | — | |
| UME (0.1, 0.1, 0.8)Training set=LibriMix (460 hours), mixboth=speech + noise, weighted sum=true, weights (λasr, λdiar, λsep)=(0.1, 0.1, 0.8)2025.08 | 12.84 | 13.39 | — | — | — | — | — | — | 90.82 | |
| TDANet-STFTParameters(M)=7.4, Separation under an unknown and variable number of speakers=false2025.12 | 12.7 | — | — | — | — | — | — | — | — | |
| UME (0.1, 0.1, 0.8)Training set=LibriMix (460 hours), mixboth=speech + noise, weighted sum=false, weights (λasr, λdiar, λsep)=(0.1, 0.1, 0.8)2025.08 | 12.64 | 13.18 | — | — | — | — | — | — | 90.49 | |
| UME (0.33, 0.33, 0.34)Training set=LibriMix (460 hours), mixboth=speech + noise, weighted sum=true, weights (λasr, λdiar, λsep)=(0.33, 0.33, 0.34)2025.08 | 12.51 | 13.05 | — | — | — | — | — | — | 90.29 | |
| UME ASR init. (0.33, 0.33, 0.34)Training set=LibriMix (460 hours), mixboth=speech + noise, RWSE=true, ASR initialization=true, weights (λasr, λdiar, λsep)=(0.33, 0.33, 0.34)2025.08 | 12.22 | 12.76 | — | — | — | — | — | — | 89.95 | |
| Conv-TasNetParams. (M)=5.1, MACs (G/s)=10.52024.06 | 12.2 | 12.7 | — | — | — | — | — | — | — | |
| GALRSize=2.3M, M=8, K=1502021.01 | 12.2 | 12.7 | — | — | — | — | — | — | — | |
| Conv-TasNetParams (M)=5.62022.09 | 12.2 | 12.7 | — | — | — | — | — | — | — | |
| Conv-TasNet2026.01 | 12.2 | 12.7 | — | — | — | — | — | — | — | |
| Conv-TasNet# Params (M)=5.1, GMAC/s (G/s)=10.52025.07 | 12.2 | 12.7 | — | — | — | — | — | — | — | |
| UME + ASR init.Training set=LibriMix (460 hours), mixboth=speech + noise, weighted sum=true, ASR initialization=true2025.08 | 12.12 | 12.68 | — | — | — | — | — | — | 89.82 | |
| DPRNNSize=2.6M, M=8, K=1502021.01 | 12 | 12.5 | — | — | — | — | — | — | — | |
| UME (λsep = 1.0)Training set=LibriMix (460 hours), mixboth=speech + noise, weighted sum=false, weights (λasr, λdiar, λsep)=(0.0, 0.0, 1.0)2025.08 | 11.81 | 12.39 | — | — | — | — | — | — | 89.13 | |
| ConvTasNet (reprod.)Training set=LibriMix (460 hours), mixboth=speech + noise, mode=max2025.08 | 10.93 | 11.48 | — | — | — | — | — | — | 87.63 | |
| BLSTM-TasNetParams (M)=23.62022.09 | 7.9 | 8.7 | — | — | — | — | — | — | — | |
| A-FRCNN-16number of blocks=162021.12 | — | — | — | — | 123.3 | — | — | — | — | |
| A-FRCNN-16 (sum)number of blocks=16, unfolding method=sum2021.12 | — | — | — | — | 22.8 | — | — | — | — | |
| Gated DualPathRNN2021.12 | — | — | — | — | 125.28 | — | — | — | — | |
| LLaSE-G1Type=Generative, Scale=single2025.03 | — | — | — | — | — | 3.48 | 3.83 | 3.11 | — | |
| LLaSE-G1Type=Generative, Scale=multi2025.03 | — | — | — | — | — | 3.5 | 3.9 | 3.17 | — | |
| LLaSE-G12025.12 | — | — | — | — | — | 3.48 | 3.83 | 3.11 | — | |
| LLaSE-G1Type=G2025.10 | — | — | — | — | — | 3.48 | 3.83 | 3.11 | — | |
| Mixture2025.10 | — | — | — | — | — | 2.33 | 1.66 | 1.64 | — | |
| Mossformer2Type=Discriminative2025.03 | — | — | — | — | — | 3.44 | 3.94 | 3.11 | — | |
| Mossformer2Type=D2025.10 | — | — | — | — | — | 3.44 | 3.94 | 3.11 | — | |
| NoisyType=None2025.03 | — | — | — | — | — | 2.33 | 1.66 | 1.64 | — | |
| QuarkAudio2025.12 | — | — | — | — | — | 3.56 | 4.04 | 3.25 | — | |
| Sepformer2021.12 | — | — | — | — | 145.58 | — | — | — | — | |
| SepformerType=Discriminative2025.03 | — | — | — | — | — | 3.33 | 3.88 | 3.02 | — | |
| SepformerType=D2025.10 | — | — | — | — | — | 3.33 | 3.88 | 3.02 | — | |
| SuDORM-RF 2.5xscale=2.5x2021.12 | — | — | — | — | 19.8 | — | — | — | — | |
| UniSEType=G2025.10 | — | — | — | — | — | 3.62 | 4.09 | 3.34 | — | |
| UniSE + PRLType=G, Strategy=PRL2025.10 | — | — | — | — | — | 3.76 | 4.2 | 3.55 | — |