Speech Deepfake Detection on SeaCF (seen setting)
98.41AccuracyGARUDA-FT
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GARUDA-FTEvaluation Protocol=Fine-tuning of LM decoder2026.06 | 98.41 | 2.78 | |
| GARUDA-FT (Wh+XV CA)Input Components=Whisper + x-vector, Fusion=Cross-Attention, Evaluation Protocol=Fine-tuning of LM decoder2026.06 | 96.07 | 5.82 | |
| GARUDA-FTTraining Data=75%2026.06 | 95.12 | 6.34 | |
| GARUDA-FT (Wh+XV KL)Input Components=Whisper + x-vector, Objective=KL Divergence, Evaluation Protocol=Fine-tuning of LM decoder2026.06 | 94.38 | 8.11 | |
| GARUDA2026.06 | 94.37 | 6.26 | |
| Qwen2-Audio-BaseEvaluation Protocol=Fine-Tuning (FT)2026.06 | 93.88 | 6.95 | |
| GARUDA-FT (Wh+XV Concat)Input Components=Whisper + x-vector, Fusion=Concatenation, Evaluation Protocol=Fine-tuning of LM decoder2026.06 | 93.78 | 8.24 | |
| GARUDA (Wh+XV CA)Input Components=Whisper + x-vector, Fusion=Cross-Attention2026.06 | 93.62 | 7.04 | |
| GARUDA-FTTraining Data=50%2026.06 | 93.56 | 6.89 | |
| GARUDA-FT (Only XV)Input Components=x-vector, Evaluation Protocol=Fine-tuning of LM decoder2026.06 | 92.69 | 9.62 | |
| GARUDATraining Data=75%2026.06 | 92.21 | 7.48 | |
| GARUDA-FTTraining Data=25%2026.06 | 92.04 | 8.29 | |
| GARUDA-FT (Only Wh)Input Components=Whisper, Evaluation Protocol=Fine-tuning of LM decoder2026.06 | 91.87 | 10.31 | |
| GARUDA (Wh+XV Concat)Input Components=Whisper + x-vector, Fusion=Concatenation2026.06 | 91.56 | 9.17 | |
| GARUDATraining Data=50%2026.06 | 91.48 | 8.27 | |
| GARUDA (Only XV)Input Components=x-vector2026.06 | 90.99 | 10.43 | |
| GARUDA (Wh+XV KL)Input Components=Whisper + x-vector, Objective=KL Divergence2026.06 | 90.83 | 9.31 | |
| GARUDATraining Data=25%2026.06 | 90.04 | 9.67 | |
| GARUDA (Only Wh)Input Components=Whisper2026.06 | 89.58 | 11.72 | |
| Qwen2-Audio-Base-FTTraining Data=75%2026.06 | 89.51 | 8.67 | |
| MiOArchitecture=Pre-Trained Backbone2026.06 | 88.76 | 12.51 | |
| SeaLLMs-Audio-7BEvaluation Protocol=Fine-Tuning (FT)2026.06 | 88.74 | 9.64 | |
| Wav2vec2-AASISTArchitecture=Pre-Trained Backbone2026.06 | 88.71 | 13.01 | |
| Wh-LCCNArchitecture=Pre-Trained Backbone2026.06 | 87.69 | 15.22 | |
| AASISTEvaluation Protocol=End-to-end2026.06 | 86.98 | 15.74 | |
| Qwen2-Audio-Base-FTTraining Data=50%2026.06 | 86.02 | 9.48 | |
| Qwen2-Audio-Base-FTTraining Data=25%2026.06 | 82.93 | 10.73 | |
| Qwen2-Audio-BaseEvaluation Protocol=Zero-shot2026.06 | 8.41 | 91.53 | |
| SeaLLMs-Audio-7BEvaluation Protocol=Zero-shot2026.06 | 6.23 | 91.64 | |
| Qwen2-Audio-ChatEvaluation Protocol=Zero-shot2026.06 | 5.96 | 91.71 | |
| Qwen-Audio-BaseEvaluation Protocol=Zero-shot2026.06 | 5.41 | 94.67 | |
| Qwen-Audio-ChatEvaluation Protocol=Zero-shot2026.06 | 3.72 | 94.39 |