Synthetic Speech Detection on ASVspoof Track 1 5 (dev)
12.52EER (%)TFPARN (full model, attention pooling)
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| TFPARN (full model, attention pooling)ID=6, Parameter Count=4.85M, Memory Usage (GB)=1.4, Latency per Utterance (ms/utt)=0.7896 ± 0.0054, Pooling=attention pooling2026.06 | 12.52 | 0.243 | 0.9243 | 0.2897 | |
| TFPARN (mean-pooling variant)ID=5, Parameter Count=4.81M, Memory Usage (GB)=1.4, Latency per Utterance (ms/utt)=0.7893 ± 0.0040, Pooling=mean-pooling2026.06 | 12.7 | 0.2499 | 0.7232 | 0.3325 | |
| TFPARN (base, cross-entropy, mean pooling)ID=3, Parameter Count=4.81M, Memory Usage (GB)=1.4, Latency per Utterance (ms/utt)=0.8073 ± 0.0542, Pooling=mean-pooling2026.06 | 12.91 | 0.2662 | 1.8796 | 0.3547 | |
| TFPARN (mean-pooling variant)ID=4, Parameter Count=4.81M, Memory Usage (GB)=1.4, Latency per Utterance (ms/utt)=0.7873 ± 0.0097, Pooling=mean-pooling2026.06 | 12.92 | 0.2561 | 1.6786 | 0.316 | |
| AASISTID=1, Parameter Count=0.30M, Memory Usage (GB)=56.7, Latency per Utterance (ms/utt)=10.4805 ± 0.00342026.06 | 18.58 | 0.2911 | 2.6545 | 0.4966 | |
| RawNet2ID=2, Parameter Count=17.62M, Memory Usage (GB)=4.9, Latency per Utterance (ms/utt)=0.7802 ± 0.02032026.06 | 27.23 | 0.5375 | 2.8672 | 0.7214 |