Speech Deepfake Detection on CodecFake (CF) seen setting
99.36AccuracyGARUDA-FT
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GARUDA-FTEvaluation Protocol=Fine-tuning of LM decoder2026.06 | 99.36 | 1.68 | |
| GARUDA-FTTraining Data=75%2026.06 | 98.43 | 4.03 | |
| GARUDA-FTTraining Data=50%2026.06 | 98.24 | 5.47 | |
| GARUDA-FT (Wh+XV CA)Input Components=Whisper + x-vector, Fusion=Cross-Attention, Evaluation Protocol=Fine-tuning of LM decoder2026.06 | 97.26 | 5.49 | |
| GARUDA2026.06 | 97 | 4.19 | |
| GARUDA-FT (Wh+XV KL)Input Components=Whisper + x-vector, Objective=KL Divergence, Evaluation Protocol=Fine-tuning of LM decoder2026.06 | 96.75 | 4.21 | |
| GARUDA-FTTraining Data=25%2026.06 | 96.26 | 6.19 | |
| GARUDATraining Data=75%2026.06 | 96.23 | 5.12 | |
| MiOArchitecture=Pre-Trained Backbone2026.06 | 95.64 | 6.37 | |
| GARUDA-FT (Wh+XV Concat)Input Components=Whisper + x-vector, Fusion=Concatenation, Evaluation Protocol=Fine-tuning of LM decoder2026.06 | 95.4 | 4.38 | |
| GARUDA-FT (Only XV)Input Components=x-vector, Evaluation Protocol=Fine-tuning of LM decoder2026.06 | 95.31 | 5.17 | |
| GARUDATraining Data=50%2026.06 | 95.17 | 6.56 | |
| Wav2vec2-AASISTArchitecture=Pre-Trained Backbone2026.06 | 95.16 | 7.08 | |
| GARUDA (Wh+XV CA)Input Components=Whisper + x-vector, Fusion=Cross-Attention2026.06 | 95.1 | 6.41 | |
| Qwen2-Audio-BaseEvaluation Protocol=Fine-Tuning (FT)2026.06 | 95.06 | 4.21 | |
| GARUDA (Wh+XV Concat)Input Components=Whisper + x-vector, Fusion=Concatenation2026.06 | 94.88 | 7.16 | |
| Wh-LCCNArchitecture=Pre-Trained Backbone2026.06 | 94.41 | 7.63 | |
| GARUDATraining Data=25%2026.06 | 93.82 | 7.31 | |
| GARUDA (Wh+XV KL)Input Components=Whisper + x-vector, Objective=KL Divergence2026.06 | 93.46 | 6.68 | |
| Qwen2-Audio-Base-FTTraining Data=75%2026.06 | 93.24 | 5.92 | |
| GARUDA (Only XV)Input Components=x-vector2026.06 | 93.15 | 7.28 | |
| AASISTEvaluation Protocol=End-to-end2026.06 | 93.09 | 8.16 | |
| GARUDA-FT (Only Wh)Input Components=Whisper, Evaluation Protocol=Fine-tuning of LM decoder2026.06 | 92.8 | 6.43 | |
| SeaLLMs-Audio-7BEvaluation Protocol=Fine-Tuning (FT)2026.06 | 90.75 | 6.96 | |
| Qwen2-Audio-Base-FTTraining Data=50%2026.06 | 90.23 | 7.42 | |
| GARUDA (Only Wh)Input Components=Whisper2026.06 | 90.12 | 7.53 | |
| Qwen2-Audio-Base-FTTraining Data=25%2026.06 | 88.07 | 8.76 | |
| Qwen2-Audio-BaseEvaluation Protocol=Zero-shot2026.06 | 19.48 | 80.47 | |
| SeaLLMs-Audio-7BEvaluation Protocol=Zero-shot2026.06 | 18.35 | 80.25 | |
| Qwen2-Audio-ChatEvaluation Protocol=Zero-shot2026.06 | 17.23 | 81.17 | |
| Qwen-Audio-BaseEvaluation Protocol=Zero-shot2026.06 | 16.07 | 84.36 | |
| Qwen-Audio-ChatEvaluation Protocol=Zero-shot2026.06 | 14.1 | 85.8 |