Speaker Recognition on VoxCeleb1 original (vox1-o)
0.24EER (mean)SpeakerRPL V2
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| SpeakerRPL V2Unknown speaker augmentation=true, Adaptive anchor=true, Model Fusion=true2026.04 | 0.24 | — | 0.01 | 99.54 | 99.85 | |
| NEMOParams (Million)=15.88, Need Speaker Labels?=Yes, Need Pre-trained ASR Model?=No2023.10 | 0.74 | — | 0.11 | — | — | |
| NEMOParams (Million)=15.88, Need Speaker Labels?=Yes, Need Pre-trained ASR Model?=Yes2023.10 | 0.88 | — | 0.137 | — | — | |
| Proposed RecXiParams (Million)=7.06, Need Speaker Labels?=Yes, Need Pre-trained ASR Model?=No2023.10 | 0.984 | — | 0.091 | — | — | |
| MCL-DPP-CParams (Million)=10.5, Need Speaker Labels?=Pseudo label, Need Pre-trained ASR Model?=No2023.10 | 1.44 | — | — | — | — | |
| ECAPA-TDNNinput=40-dimensional filterbanks, loss=AAM softmax2021.09 | 1.61 | 0.03 | — | — | — | |
| Direct EnrollmentUnknown speaker augmentation=false, Adaptive anchor=false, Model Fusion=false2026.04 | 1.72 | — | 0.08 | 98.02 | 99.76 | |
| IPAParams (Million)=>60, Need Speaker Labels?=Yes, Need Pre-trained ASR Model?=Yes2023.10 | 1.81 | — | — | — | — | |
| w2v2-aampooling=mean&std, loss=AAM softmax, feature encoder=frozen2021.09 | 1.91 | 0.12 | — | — | — | |
| w2v2-cepooling=mean&std, feature encoder=frozen2021.09 | 2.25 | 0.2 | — | — | — | |
| MCL-DPPParams (Million)=10.5, Need Speaker Labels?=No, Need Pre-trained ASR Model?=No2023.10 | 2.89 | — | — | — | — | |
| x-vectorinput=40-dimensional filterbanks2021.09 | 5.22 | 0.12 | — | — | — | |
| w2v2-bcescoring=computes scores directly, feature encoder=frozen2021.09 | 7.28 | 0.22 | — | — | — |