Audio Deepfake Detection on ITW In-the-Wild
4.78EERTran et al.
Evaluation Results
| Method | Links | |
|---|---|---|
| Tran et al.Backbone=XLS-R, Backbone status=Fine-tuned, Layers used=25 (gated), Trainable params=318M§, Fusion=Gated2026.06 | 4.78 | |
| Probing-Guided Layer SelectionBackbone=XLS-R, Backbone status=Frozen, Layers used=4, Trainable params=1.34M, Fusion=Concat2026.06 | 4.94 | |
| XLSR-Conformer-TCMLoss Function=QAMO2025.09 | 5.21 | |
| XLSR-AASISTTraining Strategy=DPDA training + PCGrad2025.09 | 5.42 | |
| XLSR-AASISTTraining Strategy=DPDA training2025.09 | 6.2 | |
| XLSR-AASIST-SAM2025.09 | 6.34 | |
| XLSR-MambaTraining Strategy=DPDA training + PCGrad2025.09 | 6.43 | |
| XLSR-Conformer-TCMTraining Strategy=DPDA training + PCGrad2025.09 | 6.48 | |
| XLSR-Conformer-NACLLoss Function=WCE2025.09 | 6.6 | |
| XLSR-Nes2NetXLoss Function=WCE, experimental_setting=Original2025.09 | 6.6 | |
| XLSR-Nes2Net-X2025.09 | 6.6 | |
| W2V2-layer7Layer count=72026.06 | 6.6 | |
| XLSR-MambaLoss Function=WCE2025.09 | 6.7 | |
| XLSR-MambaTraining Strategy=Baseline2025.09 | 6.7 | |
| XLSR-Conformer-TCMLoss Function=OC-Softmax2025.09 | 6.72 | |
| MLDG-LoRABackbone=W2V 2.0, Backbone status=LoRA, Layers used=All, Trainable params=3.59M, Fusion=Feature2026.06 | 6.81 | |
| Wav2DF-TSL2025.09 | 6.83 | |
| Xiao & VuBackbone=XLS-R, Backbone status=Frozen, Layers used=25, Trainable params=~25×cls, Fusion=Decision (sum)2026.06 | 6.9 | |
| XLSR-Conformer-TCMLoss Function=WCE, experimental_setting=Reproduced2025.09 | 7.13 | |
| XLSR-SLS2025.09 | 7.46 | |
| XLSR-MambaTraining Strategy=DPDA training2025.09 | 7.62 | |
| W2V2-layer24Layer count=242026.06 | 7.7 | |
| W2V2-TCMParams=319M, Training Protocol=Protocol 1 (ASV19 train)2026.03 | 7.79 | |
| XLSR-Conformer-TCMLoss Function=WCE, experimental_setting=Original2025.09 | 7.79 | |
| XLSR-Conformer-TCMTraining Strategy=Baseline2025.09 | 7.79 | |
| XLSR-Conformer-TCMTraining Strategy=DPDA training2025.09 | 7.97 | |
| XLSR-Nes2NetXLoss Function=OC-Softmax2025.09 | 8.23 | |
| XLSR-ConformerLoss Function=WCE2025.09 | 8.42 | |
| XLSR-Nes2NetXLoss Function=QAMO2025.09 | 8.83 | |
| HuBERT-BaseParams=100M, Training Protocol=Protocol 1 (ASV19 train)2026.03 | 9.06 | |
| RACL Diffusionreconstruction_method=SemantiCodec, aggregation=multi-layer, regularization=RACL2026.04 | 9.155 | |
| XLSR-MoE2025.09 | 9.17 | |
| WavLM-Base+Params=100M, Training Protocol=Protocol 1 (ASV19 train)2026.03 | 9.27 | |
| XLSR-Nes2NetXLoss Function=WCE, experimental_setting=Reproduced2025.09 | 9.76 | |
| mHuBERT-FinalParams=100M, Training Protocol=Protocol 1 (ASV19 train)2026.03 | 10.08 | |
| XLSR-AASISTTraining Strategy=Baseline2025.09 | 10.46 | |
| Agg Diffusionreconstruction_method=SemantiCodec, aggregation=multi-layer2026.04 | 10.679 | |
| WavLM-BaseParams=100M, Training Protocol=Protocol 1 (ASV19 train)2026.03 | 11.12 | |
| W2V2-AASISTParams=317M, Training Protocol=Protocol 1 (ASV19 train)2026.03 | 11.19 | |
| AASIST2026.06 | 11.2 | |
| mHuBERT-Iter2Params=100M, Training Protocol=Protocol 1 (ASV19 train)2026.03 | 12.24 | |
| mHuBERT-Iter1Params=100M, Training Protocol=Protocol 1 (ASV19 train)2026.03 | 13.58 | |
| Baselineimplementation=authors' implementation2026.04 | 17.949 | |
| Diffusionreconstruction_method=SemantiCodec2026.04 | 18.159 | |
| Encodecreconstruction_method=Encodec2026.04 | 22.964 | |
| Baseline*source=CodecFake paper [18]2026.04 | 23.713 | |
| HiFi-GANreconstruction_method=HiFi-GAN2026.04 | 23.779 | |
| RawNet32026.06 | 29.7 | |
| RawNet22026.06 | 33.7 | |
| RawGAT2026.06 | 33.8 | |
| DACreconstruction_method=DAC2026.04 | 39.477 |