Speaker Verification on VoxCeleb1 (Vox1-H)
0.55EERLoRA Adapter MFA
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| LoRA Adapter MFAFrontend=w2v-BERT 2.0, Params=580+6.2M, LMFT=true, Score calibration=true2025.10 | 0.55 | 0.051 | — | — | — | — | — | — | |
| ResNet293Frontend=Fbank, Params=98.9M, LMFT=true, Score calibration=true2025.10 | 0.68 | 0.07 | — | — | — | — | — | — | |
| LoRA Adapter MFAFrontend=w2v-BERT 2.0, Params=580+6.2M, LMFT=true, Score calibration=false2025.10 | 0.73 | 0.071 | — | — | — | — | — | — | |
| W2V-BERT 2.0Params=587M, GMACs=57.90, Large-Margin Finetuning=-, Training Data=VoxBlink2 + VoxCeleb22026.03 | 0.73 | — | — | — | — | — | — | — | |
| LoRA Adapter MFAFrontend=w2v-BERT 2.0, Params=580+6.2M, LMFT=false, Score calibration=false2025.10 | 0.81 | 0.082 | — | — | — | — | — | — | |
| SimAM-ResNet100Params=50.2M, GMACs=57.24, Large-Margin Finetuning=-, Training Data=VoxBlink2 + VoxCeleb22026.03 | 0.87 | — | — | — | — | — | — | — | |
| CA-MHFAFrontend=WavLM Large, Params=317+2.3M, LMFT=true, Score calibration=true2025.10 | 0.96 | — | — | — | — | — | — | — | |
| ECAPA-TDNNScoring=Cosine scoring2026.03 | 0.96 | — | — | — | — | — | — | — | |
| WavLM LargeLarge margin fine-tuning=true, Quality-aware score calibration=true2021.10 | 0.986 | — | — | — | — | — | — | — | |
| ECAPA-TDNN(C=512)Frontend=WavLM Large, Params=317+8.8M, LMFT=true, Score calibration=true2025.10 | 0.99 | — | — | — | — | — | — | — | |
| ReDimNet2-B6Params=12.3M, GMACs=13.05, Large-Margin Finetuning=Yes, Training Data=VoxCeleb2-dev2026.03 | 0.99 | — | — | — | — | — | — | — | |
| ReDimNet-B6Frontend=Fbank, Params=15.0M, LMFT=true, Score calibration=true2025.10 | 1 | 0.097 | — | — | — | — | — | — | |
| LAP+ASTPFrontend=WavLM Large, Params=317+2.3M, LMFT=true, Score calibration=true2025.10 | 1.01 | 0.099 | — | — | — | — | — | — | |
| ReDimNet-B6Params=15.0M, GMACs=20.27, Large-Margin Finetuning=Yes, Training Data=VoxCeleb2-dev2026.03 | 1.05 | — | — | — | — | — | — | — | |
| W2V-BERT 2.0Params=587M, GMACs=57.90, Large-Margin Finetuning=Yes, Training Data=VoxCeleb2-dev2026.03 | 1.06 | — | — | — | — | — | — | — | |
| ReDimNet2-B4Params=6.6M, GMACs=4.62, Large-Margin Finetuning=Yes, Training Data=VoxCeleb2-dev2026.03 | 1.07 | — | — | — | — | — | — | — | |
| ReDimNet2-B5Params=8.9M, GMACs=9.62, Large-Margin Finetuning=Yes, Training Data=VoxCeleb2-dev2026.03 | 1.07 | — | — | — | — | — | — | — | |
| ReDimNet-B5Params=9.2M, GMACs=9.87, Large-Margin Finetuning=Yes, Training Data=VoxCeleb2-dev2026.03 | 1.08 | — | — | — | — | — | — | — | |
| SimAM-ResNet34Params=25.2M, GMACs=18.20, Large-Margin Finetuning=-, Training Data=VoxBlink2 + VoxCeleb22026.03 | 1.12 | — | — | — | — | — | — | — | |
| KiwanoBackbone Architecture=fwSE-ResNet-2002026.06 | 1.13 | — | — | — | — | — | — | — | |
| ECAPA-TDNN(C=512)Frontend=Wav2Wec2.0 Large, Params=317+8.8M, LMFT=true, Score calibration=true2025.10 | 1.14 | — | — | — | — | — | — | — | |
| ECAPA2Params=27.1M, GMACs=187.0*, Large-Margin Finetuning=Yes, Training Data=VoxCeleb2-dev2026.03 | 1.15 | — | — | — | — | — | — | — | |
| ECAPA-TDNN(C=512)Frontend=UniSpeech-SAT Large, Params=317+8.8M, LMFT=true, Score calibration=true2025.10 | 1.18 | — | — | — | — | — | — | — | |
| ResNet221Frontend=Fbank, Params=23.8M, LMFT=true, Score calibration=true2025.10 | 1.21 | — | — | — | — | — | — | — | |
| ReDimNet2-B3Params=4.1M, GMACs=2.70, Large-Margin Finetuning=Yes, Training Data=VoxCeleb2-dev2026.03 | 1.22 | — | — | — | — | — | — | — | |
| ECAPA-TDNN(C=512)Frontend=HuBERT Large, Params=317+8.8M, LMFT=true, Score calibration=true2025.10 | 1.23 | — | — | — | — | — | — | — | |
| ReDimNet-B4Params=6.3M, GMACs=4.80, Large-Margin Finetuning=Yes, Training Data=VoxCeleb2-dev2026.03 | 1.26 | — | — | — | — | — | — | — | |
| ResNet293Params=28.6M, GMACs=28.10, Large-Margin Finetuning=Yes, Training Data=VoxCeleb2-dev2026.03 | 1.3 | — | — | — | — | — | — | — | |
| WeSpeakerBackbone Architecture=ResNet-2932026.06 | 1.31 | — | — | — | — | — | — | — | |
| WavLM Large2021.10 | 1.318 | — | — | — | — | — | — | — | |
| ReDimNet-B2Params=4.7M, GMACs=0.90, Large-Margin Finetuning=Yes, Training Data=VoxCeleb2-dev2026.03 | 1.32 | — | — | — | — | — | — | — | |
| ReDimNet-B3Params=3.0M, GMACs=3.00, Large-Margin Finetuning=Yes, Training Data=VoxCeleb2-dev2026.03 | 1.33 | — | — | — | — | — | — | — | |
| WavLMParams=324M, GMACs=26.53, Large-Margin Finetuning=Yes, Training Data=VoxCeleb2-dev2026.03 | 1.34 | — | — | — | — | — | — | — | |
| HuBERT LargeLarge margin fine-tuning=true, Quality-aware score calibration=true2021.10 | 1.342 | — | — | — | — | — | — | — | |
| MFAFrontend=Nemo Large, Params=131M, LMFT=true, Score calibration=true2025.10 | 1.35 | 0.135 | — | — | — | — | — | — | |
| ReDimNet2-B2Params=3.6M, GMACs=0.95, Large-Margin Finetuning=Yes, Training Data=VoxCeleb2-dev2026.03 | 1.41 | — | — | — | — | — | — | — | |
| 3D-SpeakerBackbone Architecture=ERes2Net-large-lm2026.06 | 1.44 | — | — | — | — | — | — | — | |
| ERes2NetV2Frontend=Fbank, Params=17.8M, LMFT=true, Score calibration=false2025.10 | 1.45 | 0.143 | — | — | — | — | — | — | |
| SpeakerCard-1MRegime=audio only2026.06 | 1.58 | — | — | — | — | — | — | — | |
| Gemini-DFResNet114Params=6.5M, GMACs=5.42, Large-Margin Finetuning=No, Training Data=VoxCeleb2-dev2026.03 | 1.6 | — | — | — | — | — | — | — | |
| TRKDTeacher Model=ReDim-B5, Student Model=ReDim-B22026.01 | 1.644 | — | — | — | — | — | — | — | |
| CAM++Params=7.2M, GMACs=1.15, Large-Margin Finetuning=Yes, Training Data=VoxCeleb2-dev2026.03 | 1.66 | — | — | — | — | — | — | — | |
| HuBERT Large2021.10 | 1.678 | — | — | — | — | — | — | — | |
| ReDimNet2-B1Params=2.1M, GMACs=0.56, Large-Margin Finetuning=Yes, Training Data=VoxCeleb2-dev2026.03 | 1.72 | — | — | — | — | — | — | — | |
| ReDimNet-B1Params=2.2M, GMACs=0.54, Large-Margin Finetuning=Yes, Training Data=VoxCeleb2-dev2026.03 | 1.73 | — | — | — | — | — | — | — | |
| WavLM-Base SV2026.06 | 1.75 | — | — | — | — | — | — | — | |
| WavLM Base+2021.10 | 1.758 | — | — | — | — | — | — | — | |
| CAM++Frontend=Fbank, Params=7.2M, LMFT=false, Score calibration=false2025.10 | 1.76 | 0.173 | — | — | — | — | — | — | |
| ECAPA512 + U^3-xi# Param.=6.69 M, Loss=USphereFace2 (1)+(4), Uncertainty-aware cosine score=true2026.07 | 1.771 | 0.178 | — | — | — | — | — | 12.81 | |
| ECAPA512 + U^3-xi# Param.=6.69 M, Loss=UAM-Softmax (1)+(4), Uncertainty-aware cosine score=true2026.07 | 1.794 | 0.178 | — | — | — | — | — | 19.46 | |
| NeXt-TDNN (C=256)Params=7.1M, GMACs=1.35*, Large-Margin Finetuning=No, Training Data=VoxCeleb2-dev2026.03 | 1.82 | — | — | — | — | — | — | — | |
| ECAPA512 + U^3-xi# Param.=6.69 M, Loss=UAAM-Softmax (1)+(4), Uncertainty-aware cosine score=true2026.07 | 1.833 | 0.189 | — | — | — | — | — | 21.22 | |
| Teacher (w/o KD)Teacher Model=ECAPA1024, Student Model=ECAPA4002026.01 | 1.839 | — | — | — | — | — | — | — | |
| NeXt-TDNN-l (C=256)Params=6.0M, GMACs=1.13*, Large-Margin Finetuning=No, Training Data=VoxCeleb2-dev2026.03 | 1.86 | — | — | — | — | — | — | — | |
| CAM++# Param.=7.2 M, Uncertainty-aware cosine score=false2026.07 | 1.863 | 0.179 | — | — | — | — | — | — | |
| ECAPA-TDNN2026.06 | 1.87 | — | — | — | — | — | — | — | |
| ECAPA512 + U^3-xi# Param.=6.69 M, Loss=UAAM-Softmax (1)+(2), Uncertainty-aware cosine score=true2026.07 | 1.879 | 0.19 | — | — | — | — | — | 18.53 | |
| ChebyAAMObjective=ChebyAAM, Backbone=ECAPA-TDNN2026.01 | 1.9008 | 0.196 | — | — | — | — | — | — | |
| ECAPA512 + U^3-xi# Param.=6.69 M, Loss=USphereFace2 (1)+(4), Uncertainty-aware cosine score=false2026.07 | 1.918 | 0.196 | — | — | — | — | — | 5.21 | |
| ECAPA512 + U^3-xi# Param.=6.69 M, Loss=UAAM-Softmax (1) [22], Uncertainty-aware cosine score=true2026.07 | 1.926 | 0.191 | — | — | — | — | — | 14.95 | |
| ResNet34# Param.=6.63 M, Uncertainty-aware cosine score=false2026.07 | 1.96 | 0.192 | — | — | — | — | — | — | |
| ECAPA512# Param.=6.19 M, Loss=SphereFace2 [37], [38], Uncertainty-aware cosine score=false2026.07 | 1.967 | 0.199 | — | — | — | — | — | — | |
| AAM-SoftmaxObjective=AAM-Softmax, Backbone=ECAPA-TDNN2026.01 | 1.9681 | 0.1989 | — | — | — | — | — | — | |
| ReDimNet2-B0Params=1.1M, GMACs=0.33, Large-Margin Finetuning=Yes, Training Data=VoxCeleb2-dev2026.03 | 1.97 | — | — | — | — | — | — | — | |
| ECAPA512 + U^3-xi# Param.=6.69 M, Loss=UAM-Softmax (1)+(4), Uncertainty-aware cosine score=false2026.07 | 1.973 | 0.186 | — | — | — | — | — | 11.46 | |
| Gemini SD-ResNet38# Param.=6.72 M, Uncertainty-aware cosine score=false2026.07 | 1.974 | 0.185 | — | — | — | — | — | — | |
| ECAPA512 + U^3-xi# Param.=6.69 M, Loss=UAAM-Softmax (1)+(4), Uncertainty-aware cosine score=false2026.07 | 1.978 | 0.195 | — | — | — | — | — | 13.4 | |
| NeXt-TDNN (C=128)Params=1.9M, GMACs=0.35*, Large-Margin Finetuning=No, Training Data=VoxCeleb2-dev2026.03 | 1.98 | — | — | — | — | — | — | — | |
| TRKDTeacher Model=ECAPA1024, Student Model=ECAPA4002026.01 | 2.001 | — | — | — | — | — | — | — | |
| ECAPA512 + U^3-xi# Param.=6.69 M, Loss=UAAM-Softmax (1) [22], Uncertainty-aware cosine score=false2026.07 | 2.006 | 0.199 | — | — | — | — | — | 10.6 | |
| TRKDTeacher Model=RN152, Student Model=MNV22026.01 | 2.016 | — | — | — | — | — | — | — | |
| ECAPA512 + U^3-xi# Param.=6.69 M, Loss=UAAM-Softmax (1)+(2), Uncertainty-aware cosine score=false2026.07 | 2.036 | 0.202 | — | — | — | — | — | 9.26 | |
| ECAPA1024# Param.=14.65 M, Uncertainty-aware cosine score=false2026.07 | 2.059 | 0.205 | — | — | — | — | — | — | |
| Teachersource=Wespeaker2026.07 | 2.06 | 0.205 | — | — | — | — | — | — | |
| SpeakerCard-1MRegime=balanced2026.06 | 2.07 | — | — | — | — | — | — | — | |
| ECAPA512 + U^3-xi# Param.=6.69 M, Loss=UAAM-Softmax (3), Uncertainty-aware cosine score=true2026.07 | 2.093 | 0.205 | — | — | — | — | — | 9.05 | |
| ECAPA512 + U^3-xi# Param.=6.69 M, Loss=UAAM-Softmax (3), Uncertainty-aware cosine score=false2026.07 | 2.114 | 0.206 | — | — | — | — | — | 8.32 | |
| ECAPA-TDNN(C=1024)Frontend=Fbank, Params=14.7M, LMFT=false, Score calibration=false2025.10 | 2.12 | 0.21 | — | — | — | — | — | — | |
| NeXt-TDNN-l (C=128)Params=1.6M, GMACs=0.29*, Large-Margin Finetuning=No, Training Data=VoxCeleb2-dev2026.03 | 2.12 | — | — | — | — | — | — | — | |
| ECAPA-TDNNimplementation=reproduction2021.10 | 2.127 | — | — | — | — | — | — | — | |
| AM-SoftmaxObjective=AM-Softmax, Backbone=ECAPA-TDNN2026.01 | 2.1448 | 0.2131 | — | — | — | — | — | — | |
| TRKDTeacher Model=SAM-RN50, Student Model=R2N342026.01 | 2.178 | — | — | — | — | — | — | — | |
| SA-TinyLLaMAScoring=Log-likelihood based scoring2026.03 | 2.2 | — | — | — | — | — | — | — | |
| ReDimNet-B0Params=1.0M, GMACs=0.43, Large-Margin Finetuning=Yes, Training Data=VoxCeleb2-dev2026.03 | 2.2 | — | — | — | — | — | — | — | |
| ECAPA (C=512)Params=6.2M, GMACs=1.04, Large-Margin Finetuning=No, Training Data=VoxCeleb2-dev2026.03 | 2.2 | — | — | — | — | — | — | — | |
| HuBERT Base2021.10 | 2.216 | — | — | — | — | — | — | — | |
| ECAPA512# Param.=6.19 M, Loss=AM-Softmax [35], [36], Uncertainty-aware cosine score=false2026.07 | 2.254 | 0.221 | — | — | — | — | — | — | |
| ECAPA512# Param.=6.19 M, Loss=AAM-Softmax [30], Uncertainty-aware cosine score=false2026.07 | 2.31 | 0.226 | — | — | — | — | — | — | |
| ECAPA-TDNN2021.10 | 2.32 | — | — | — | — | — | — | — | |
| SpeakerCard-1MRegime=retrieval-spec.2026.06 | 2.38 | — | — | — | — | — | — | — | |
| CFKDBitrate=24 kbps, Backbone=ECAPA-TDNN1024, lambda=402026.07 | 2.39 | 0.23 | — | — | — | — | — | — | |
| TRKDTeacher Model=RN34, Student Model=RN182026.01 | 2.461 | — | — | — | — | — | — | — | |
| Student (w/o KD)Teacher Model=ECAPA1024, Student Model=ECAPA4002026.01 | 2.607 | — | — | — | — | — | — | — | |
| CFKDBitrate=12 kbps, Backbone=ECAPA-TDNN1024, lambda=402026.07 | 2.62 | 0.25 | — | — | — | — | — | — | |
| A-SoftmaxObjective=A-Softmax, Backbone=ECAPA-TDNN2026.01 | 2.8732 | 0.2954 | — | — | — | — | — | — | |
| TRKDTeacher Model=CAM++, Student Model=X-vector2026.01 | 3.164 | — | — | — | — | — | — | — | |
| SA-TinyLLaMAScoring=Log-likelihood based scoring, Variant=XS2026.03 | 3.44 | — | — | — | — | — | — | — | |
| CFKDBitrate=6 kbps, Backbone=ECAPA-TDNN1024, lambda=402026.07 | 3.56 | 0.325 | — | — | — | — | — | — | |
| SA-TinyLLaMAScoring=Log-likelihood based scoring, Backbone=Frozen2026.03 | 6.6 | — | — | — | — | — | — | — | |
| CFKDBitrate=3 kbps, Backbone=ECAPA-TDNN1024, lambda=402026.07 | 6.96 | 0.543 | — | — | — | — | — | — |