Speech Emotion Recognition on CAFE (ASR)
87.9ASR AccuracyTTS-enabled backdoor framework
Evaluation Results
| Method | Links | |
|---|---|---|
| TTS-enabled backdoor frameworkModel Architecture=UniSpeech, Poisoning Ratio (ρ)=0.62026.06 | 87.9 | |
| TTS-enabled backdoor frameworkModel Architecture=Wav2Vec2, Poisoning Ratio (ρ)=0.62026.06 | 75.9 | |
| TTS-enabled backdoor frameworkModel Architecture=data2vec, Poisoning Ratio (ρ)=0.62026.06 | 74.1 | |
| TTS-enabled backdoor frameworkModel Architecture=WavLM, Poisoning Ratio (ρ)=0.62026.06 | 62.1 |