Speech-to-SQL Parsing on MASpider in-domain (test)
67.4SELECT AccuracyS2SQL-TTS
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| S2SQL-TTSSynthesized data upper bound=FastSpeech 22023.05 | 67.4 | 50 | 66.3 | 67.5 | 97.1 | 78.1 | 38 | |
| Wav2SQLSpeech feature extractor=Hubert, Language model=GloVe2023.05 | 63.6 | 47.3 | 56.3 | 60.2 | 96.6 | 71.1 | 34.1 | |
| Wav2SQL (w/o Gradient reversal classifier)Gradient reversal classifier=false2023.05 | 62.1 | 37.9 | 52.6 | 60.7 | 95.8 | 63.2 | 27.5 | |
| Wav2SQL (w/o Speech reprogramming)Speech reprogramming=false2023.05 | 59.6 | 39.7 | 48 | 53.6 | 95.8 | 63.4 | 26.7 | |
| CascadedASR=wav2vec 2.0, Text-to-SQL parser=RAT-SQL2023.05 | 58.6 | 44.2 | 56.4 | 59.4 | 96.2 | 72.4 | 31.6 | |
| DeepSpeechSQLSpeech encoder=DeepSpeech2023.05 | 43.3 | 25.8 | 26.2 | 17.5 | 93.9 | 35.1 | 14.4 |