Speaker Verification on VoxCeleb-E
0.27EERLoRA Adapter MFA
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| LoRA Adapter MFAFrontend=w2v-BERT 2.0, Params=580+6.2M, LMFT=true, Score calibration=true2025.10 | 0.27 | 0.028 | — | — | — | — | |
| LoRA Adapter MFAFrontend=w2v-BERT 2.0, Params=580+6.2M, LMFT=true, Score calibration=false2025.10 | 0.31 | 0.032 | — | — | — | — | |
| W2V-BERT 2.0Params=587M, GMACs=57.90, Large-Margin Finetuning=-, Training Data=VoxBlink2 + VoxCeleb22026.03 | 0.31 | — | — | — | — | — | |
| ResNet293Frontend=Fbank, Params=98.9M, LMFT=true, Score calibration=true2025.10 | 0.37 | 0.037 | — | — | — | — | |
| LoRA Adapter MFAFrontend=w2v-BERT 2.0, Params=580+6.2M, LMFT=false, Score calibration=false2025.10 | 0.38 | 0.04 | — | — | — | — | |
| ECAPA-TDNNScoring=Cosine scoring2026.03 | 0.45 | — | — | — | — | — | |
| SimAM-ResNet100Params=50.2M, GMACs=57.24, Large-Margin Finetuning=-, Training Data=VoxBlink2 + VoxCeleb22026.03 | 0.46 | — | — | — | — | — | |
| ECAPA-TDNN(C=512)Frontend=WavLM Large, Params=317+8.8M, LMFT=true, Score calibration=true2025.10 | 0.48 | — | — | — | — | — | |
| CA-MHFAFrontend=WavLM Large, Params=317+2.3M, LMFT=true, Score calibration=true2025.10 | 0.48 | — | — | — | — | — | |
| LAP+ASTPFrontend=WavLM Large, Params=317+2.3M, LMFT=true, Score calibration=true2025.10 | 0.5 | 0.055 | — | — | — | — | |
| W2V-BERT 2.0Params=587M, GMACs=57.90, Large-Margin Finetuning=Yes, Training Data=VoxCeleb2-dev2026.03 | 0.51 | — | — | — | — | — | |
| ReDimNet2-B6Params=12.3M, GMACs=13.05, Large-Margin Finetuning=Yes, Training Data=VoxCeleb2-dev2026.03 | 0.52 | — | — | — | — | — | |
| ReDimNet-B6Frontend=Fbank, Params=15.0M, LMFT=true, Score calibration=true2025.10 | 0.53 | 0.051 | — | — | — | — | |
| ReDimNet-B6Params=15.0M, GMACs=20.27, Large-Margin Finetuning=Yes, Training Data=VoxCeleb2-dev2026.03 | 0.55 | — | — | — | — | — | |
| ReDimNet2-B5Params=8.9M, GMACs=9.62, Large-Margin Finetuning=Yes, Training Data=VoxCeleb2-dev2026.03 | 0.56 | — | — | — | — | — | |
| ECAPA-TDNN(C=512)Frontend=UniSpeech-SAT Large, Params=317+8.8M, LMFT=true, Score calibration=true2025.10 | 0.57 | — | — | — | — | — | |
| ReDimNet2-B4Params=6.6M, GMACs=4.62, Large-Margin Finetuning=Yes, Training Data=VoxCeleb2-dev2026.03 | 0.58 | — | — | — | — | — | |
| ReDimNet-B5Params=9.2M, GMACs=9.87, Large-Margin Finetuning=Yes, Training Data=VoxCeleb2-dev2026.03 | 0.61 | — | — | — | — | — | |
| ECAPA2Params=27.1M, GMACs=187.0*, Large-Margin Finetuning=Yes, Training Data=VoxCeleb2-dev2026.03 | 0.62 | — | — | — | — | — | |
| SimAM-ResNet34Params=25.2M, GMACs=18.20, Large-Margin Finetuning=-, Training Data=VoxBlink2 + VoxCeleb22026.03 | 0.62 | — | — | — | — | — | |
| ECAPA-TDNN(C=512)Frontend=Wav2Wec2.0 Large, Params=317+8.8M, LMFT=true, Score calibration=true2025.10 | 0.63 | — | — | — | — | — | |
| WavLMParams=324M, GMACs=26.53, Large-Margin Finetuning=Yes, Training Data=VoxCeleb2-dev2026.03 | 0.63 | — | — | — | — | — | |
| ECAPA-TDNN(C=512)Frontend=HuBERT Large, Params=317+8.8M, LMFT=true, Score calibration=true2025.10 | 0.65 | — | — | — | — | — | |
| MFAFrontend=Nemo Large, Params=131M, LMFT=true, Score calibration=true2025.10 | 0.66 | 0.071 | — | — | — | — | |
| ReDimNet2-B3Params=4.1M, GMACs=2.70, Large-Margin Finetuning=Yes, Training Data=VoxCeleb2-dev2026.03 | 0.66 | — | — | — | — | — | |
| ResNet221Frontend=Fbank, Params=23.8M, LMFT=true, Score calibration=true2025.10 | 0.68 | — | — | — | — | — | |
| ReDimNet-B4Params=6.3M, GMACs=4.80, Large-Margin Finetuning=Yes, Training Data=VoxCeleb2-dev2026.03 | 0.68 | — | — | — | — | — | |
| ResNet293Params=28.6M, GMACs=28.10, Large-Margin Finetuning=Yes, Training Data=VoxCeleb2-dev2026.03 | 0.71 | — | — | — | — | — | |
| ReDimNet-B3Params=3.0M, GMACs=3.00, Large-Margin Finetuning=Yes, Training Data=VoxCeleb2-dev2026.03 | 0.73 | — | — | — | — | — | |
| ERes2NetV2Frontend=Fbank, Params=17.8M, LMFT=true, Score calibration=false2025.10 | 0.76 | 0.082 | — | — | — | — | |
| ReDimNet-B2Params=4.7M, GMACs=0.90, Large-Margin Finetuning=Yes, Training Data=VoxCeleb2-dev2026.03 | 0.76 | — | — | — | — | — | |
| ReDimNet2-B2Params=3.6M, GMACs=0.95, Large-Margin Finetuning=Yes, Training Data=VoxCeleb2-dev2026.03 | 0.76 | — | — | — | — | — | |
| SpeakerCard-1MRegime=audio only2026.06 | 0.79 | — | — | — | — | — | |
| CAM++Params=7.2M, GMACs=1.15, Large-Margin Finetuning=Yes, Training Data=VoxCeleb2-dev2026.03 | 0.85 | — | — | — | — | — | |
| TRKDTeacher Model=ReDim-B5, Student Model=ReDim-B22026.01 | 0.885 | — | — | — | — | — | |
| CAM++Frontend=Fbank, Params=7.2M, LMFT=false, Score calibration=false2025.10 | 0.89 | 0.1 | — | — | — | — | |
| Gemini-DFResNet114Params=6.5M, GMACs=5.42, Large-Margin Finetuning=No, Training Data=VoxCeleb2-dev2026.03 | 0.91 | — | — | — | — | — | |
| SpeakerCard-1MRegime=balanced2026.06 | 0.91 | — | — | — | — | — | |
| WavLM-Base SV2026.06 | 0.92 | — | — | — | — | — | |
| CAM++# Param.=7.2 M, Uncertainty-aware cosine score=false2026.07 | 0.931 | 0.109 | — | — | — | — | |
| ReDimNet2-B1Params=2.1M, GMACs=0.56, Large-Margin Finetuning=Yes, Training Data=VoxCeleb2-dev2026.03 | 0.96 | — | — | — | — | — | |
| ECAPA512 + U^3-xi# Param.=6.69 M, Loss=UAAM-Softmax (1)+(4), Uncertainty-aware cosine score=true2026.07 | 0.965 | 0.11 | — | — | — | 21.22 | |
| ECAPA512 + U^3-xi# Param.=6.69 M, Loss=USphereFace2 (1)+(4), Uncertainty-aware cosine score=true2026.07 | 0.965 | 0.108 | — | — | — | 12.81 | |
| ReDimNet-B1Params=2.2M, GMACs=0.54, Large-Margin Finetuning=Yes, Training Data=VoxCeleb2-dev2026.03 | 0.97 | — | — | — | — | — | |
| ChebyAAMObjective=ChebyAAM, Backbone=ECAPA-TDNN2026.01 | 0.9793 | 0.1127 | — | — | — | — | |
| ECAPA512 + U^3-xi# Param.=6.69 M, Loss=UAAM-Softmax (1)+(2), Uncertainty-aware cosine score=true2026.07 | 0.988 | 0.113 | — | — | — | 18.53 | |
| ECAPA-TDNN2026.06 | 0.99 | — | — | — | — | — | |
| ECAPA512 + U^3-xi# Param.=6.69 M, Loss=UAM-Softmax (1)+(4), Uncertainty-aware cosine score=true2026.07 | 0.991 | 0.109 | — | — | — | 19.46 | |
| AAM-SoftmaxObjective=AAM-Softmax, Backbone=ECAPA-TDNN2026.01 | 1.0066 | 0.1132 | — | — | — | — | |
| SA-TinyLLaMAScoring=Log-likelihood based scoring2026.03 | 1.03 | — | — | — | — | — | |
| ECAPA512 + U^3-xi# Param.=6.69 M, Loss=UAAM-Softmax (1) [22], Uncertainty-aware cosine score=true2026.07 | 1.035 | 0.115 | — | — | — | 14.95 | |
| ECAPA512 + U^3-xi# Param.=6.69 M, Loss=USphereFace2 (1)+(4), Uncertainty-aware cosine score=false2026.07 | 1.035 | 0.119 | — | — | — | 5.21 | |
| NeXt-TDNN-l (C=256)Params=6.0M, GMACs=1.13*, Large-Margin Finetuning=No, Training Data=VoxCeleb2-dev2026.03 | 1.04 | — | — | — | — | — | |
| NeXt-TDNN (C=256)Params=7.1M, GMACs=1.35*, Large-Margin Finetuning=No, Training Data=VoxCeleb2-dev2026.03 | 1.04 | — | — | — | — | — | |
| ResNet34# Param.=6.63 M, Uncertainty-aware cosine score=false2026.07 | 1.049 | 0.121 | — | — | — | — | |
| ECAPA512 + U^3-xi# Param.=6.69 M, Loss=UAAM-Softmax (1)+(4), Uncertainty-aware cosine score=false2026.07 | 1.05 | 0.122 | — | — | — | 13.4 | |
| Teacher (w/o KD)Teacher Model=ECAPA1024, Student Model=ECAPA4002026.01 | 1.051 | — | — | — | — | — | |
| TRKDTeacher Model=RN152, Student Model=MNV22026.01 | 1.068 | — | — | — | — | — | |
| ECAPA512 + U^3-xi# Param.=6.69 M, Loss=UAAM-Softmax (1)+(2), Uncertainty-aware cosine score=false2026.07 | 1.069 | 0.121 | — | — | — | 9.26 | |
| SpeakerCard-1MRegime=retrieval-spec.2026.06 | 1.07 | — | — | — | — | — | |
| Teachersource=Wespeaker2026.07 | 1.07 | 0.117 | — | — | — | — | |
| ECAPA1024# Param.=14.65 M, Uncertainty-aware cosine score=false2026.07 | 1.072 | 0.117 | — | — | — | — | |
| ECAPA512 + U^3-xi# Param.=6.69 M, Loss=UAAM-Softmax (1) [22], Uncertainty-aware cosine score=false2026.07 | 1.075 | 0.121 | — | — | — | 10.6 | |
| ECAPA512 + U^3-xi# Param.=6.69 M, Loss=UAM-Softmax (1)+(4), Uncertainty-aware cosine score=false2026.07 | 1.076 | 0.119 | — | — | — | 11.46 | |
| AM-SoftmaxObjective=AM-Softmax, Backbone=ECAPA-TDNN2026.01 | 1.0872 | 0.1301 | — | — | — | — | |
| ECAPA512 + U^3-xi# Param.=6.69 M, Loss=UAAM-Softmax (3), Uncertainty-aware cosine score=true2026.07 | 1.094 | 0.125 | — | — | — | 9.05 | |
| ECAPA512 + U^3-xi# Param.=6.69 M, Loss=UAAM-Softmax (3), Uncertainty-aware cosine score=false2026.07 | 1.102 | 0.127 | — | — | — | 8.32 | |
| TRKDTeacher Model=ECAPA1024, Student Model=ECAPA4002026.01 | 1.115 | — | — | — | — | — | |
| ECAPA-TDNN(C=1024)Frontend=Fbank, Params=14.7M, LMFT=false, Score calibration=false2025.10 | 1.12 | 0.132 | — | — | — | — | |
| ECAPA512# Param.=6.19 M, Loss=SphereFace2 [37], [38], Uncertainty-aware cosine score=false2026.07 | 1.121 | 0.125 | — | — | — | — | |
| Gemini SD-ResNet38# Param.=6.72 M, Uncertainty-aware cosine score=false2026.07 | 1.13 | 0.117 | — | — | — | — | |
| TRKDTeacher Model=SAM-RN50, Student Model=R2N342026.01 | 1.157 | — | — | — | — | — | |
| ReDimNet2-B0Params=1.1M, GMACs=0.33, Large-Margin Finetuning=Yes, Training Data=VoxCeleb2-dev2026.03 | 1.16 | — | — | — | — | — | |
| NeXt-TDNN (C=128)Params=1.9M, GMACs=0.35*, Large-Margin Finetuning=No, Training Data=VoxCeleb2-dev2026.03 | 1.17 | — | — | — | — | — | |
| CFKDBitrate=24 kbps, Backbone=ECAPA-TDNN1024, lambda=402026.07 | 1.2 | 0.139 | — | — | — | — | |
| ECAPA512# Param.=6.19 M, Loss=AM-Softmax [35], [36], Uncertainty-aware cosine score=false2026.07 | 1.206 | 0.133 | — | — | — | — | |
| ECAPA512# Param.=6.19 M, Loss=AAM-Softmax [30], Uncertainty-aware cosine score=false2026.07 | 1.209 | 0.136 | — | — | — | — | |
| ECAPA (C=512)Params=6.2M, GMACs=1.04, Large-Margin Finetuning=No, Training Data=VoxCeleb2-dev2026.03 | 1.21 | — | — | — | — | — | |
| NeXt-TDNN-l (C=128)Params=1.6M, GMACs=0.29*, Large-Margin Finetuning=No, Training Data=VoxCeleb2-dev2026.03 | 1.24 | — | — | — | — | — | |
| ReDimNet-B0Params=1.0M, GMACs=0.43, Large-Margin Finetuning=Yes, Training Data=VoxCeleb2-dev2026.03 | 1.25 | — | — | — | — | — | |
| CFKDBitrate=12 kbps, Backbone=ECAPA-TDNN1024, lambda=402026.07 | 1.3 | 0.15 | — | — | — | — | |
| TRKDTeacher Model=RN34, Student Model=RN182026.01 | 1.322 | — | — | — | — | — | |
| Student (w/o KD)Teacher Model=ECAPA1024, Student Model=ECAPA4002026.01 | 1.395 | — | — | — | — | — | |
| A-SoftmaxObjective=A-Softmax, Backbone=ECAPA-TDNN2026.01 | 1.429 | 0.175 | — | — | — | — | |
| TRKDTeacher Model=CAM++, Student Model=X-vector2026.01 | 1.692 | — | — | — | — | — | |
| CFKDBitrate=6 kbps, Backbone=ECAPA-TDNN1024, lambda=402026.07 | 1.8 | 0.208 | — | — | — | — | |
| SA-TinyLLaMAScoring=Log-likelihood based scoring, Variant=XS2026.03 | 2.21 | — | — | — | — | — | |
| CFKDBitrate=3 kbps, Backbone=ECAPA-TDNN1024, lambda=402026.07 | 3.47 | 0.386 | — | — | — | — | |
| SA-TinyLLaMAScoring=Log-likelihood based scoring, Backbone=Frozen2026.03 | 4.21 | — | — | — | — | — | |
| CFKDBitrate=1.5 kbps, Backbone=ECAPA-TDNN1024, lambda=402026.07 | 9.72 | 0.787 | — | — | — | — | |
| SA-Ministral3Scoring=Log-likelihood based scoring2026.03 | 15.88 | — | — | — | — | — | |
| SA-TinyLLaMAScoring=Log-likelihood based scoring, Backbone=Frozen, Variant=XS2026.03 | 27.82 | — | — | — | — | — | |
| KiwanoBackbone Architecture=fwSE-ResNet-2002026.06 | 64 | — | — | — | — | — | |
| WeSpeakerBackbone Architecture=ResNet-2932026.06 | 71 | — | — | — | — | — | |
| 3D-SpeakerBackbone Architecture=ERes2Net-large-lm2026.06 | 75 | — | — | — | — | — | |
| D-ALMFTEncoder=ResNet342026.01 | — | — | 1.18 | 2.15 | 4.06 | — | |
| D-ALMFTEncoder=ECAPA-TDNN2026.01 | — | — | 1.18 | 2.31 | 4.17 | — | |
| D-ALMFTEncoder=ERes2NetV22026.01 | — | — | 1.14 | 2.16 | 3.89 | — | |
| DAME-FT-HWEncoder=ResNet342026.01 | — | — | 0.97 | 2.11 | 4.03 | — | |
| DAME-FT-HWEncoder=ECAPA-TDNN2026.01 | — | — | 1.14 | 2.3 | 4.22 | — |