Dialog State Tracking on SpokenWoz (test)
45.52JGAFull Spoken Context + Gemma2-9B-Instruct
Evaluation Results
| Method | Links | |
|---|---|---|
| Full Spoken Context + Gemma2-9B-InstructSetup=potentially data-contaminated models2025.10 | 45.52 | |
| Compressed Spoken Context + Gemma2-9B-InstructSetup=potentially data-contaminated models2025.10 | 43.16 | |
| Joint Speech and Text E2E DSTTraining Speech=SW, LLM=Gemma-3-12B-it, Training Text=MW2025.11 | 43 | |
| Joint Speech and Text E2E DSTTraining Speech=SW, LLM=Gemma-3-12B-it, Training Text=✗2025.11 | 42.6 | |
| Joint Speech and Text E2E DSTTraining Speech=MW, LLM=Gemma-3-12B-it, Training Text=SW2025.11 | 42.2 | |
| WavLM + conn. + Gemma2-9B-InstructSetup=potentially data-contaminated models2025.10 | 42.17 | |
| Joint Speech and Text E2E DSTTraining Speech=SW, LLM=Gemma-3-4B-it, Training Text=✗2025.11 | 41 | |
| Joint Speech and Text E2E DSTTraining Speech=SW, LLM=Gemma-3-4B-it, Training Text=MW2025.11 | 40.3 | |
| Full Spoken ContextSetup=open data models2025.10 | 39.32 | |
| Joint Speech and Text E2E DSTTraining Speech=MW, LLM=Gemma-3-4B-it, Training Text=SW2025.11 | 37.4 | |
| Compressed Spoken ContextSetup=open data models2025.10 | 36.49 | |
| Joint Speech and Text E2E DSTTraining Speech=SW, LLM=Gemma-3-1B-it, Training Text=MW2025.11 | 36.3 | |
| Joint Speech and Text E2E DSTTraining Speech=SW, LLM=Gemma-3-1B-it, Training Text=✗2025.11 | 35.6 | |
| Joint Speech and Text E2E DSTTraining Speech=SW, LLM=OLMo-1B, Training Text=MW2025.11 | 34.9 | |
| WavLM + conn. + OLMo-1BSetup=open data models2025.10 | 34.66 | |
| Joint Speech and Text E2E DSTTraining Speech=SW, LLM=OLMo-1B, Training Text=✗2025.11 | 34.1 | |
| OLMo-1B [7]Training Speech=SW, LLM=OLMo-1B, Training Text=✗2025.11 | 32.1 | |
| Joint Speech and Text E2E DSTTraining Speech=MW, LLM=Gemma-3-1B-it, Training Text=SW2025.11 | 30.9 | |
| Joint Speech and Text E2E DSTTraining Speech=MW, LLM=OLMo-1B, Training Text=SW2025.11 | 29.3 | |
| UBAR + GenWOZSetup=open data models2025.10 | 25.9 | |
| SPACE+WavLMalignSetup=open data models2025.10 | 25.65 | |
| E2E (Whisper+T5)Setup=open data models2025.10 | 24.1 | |
| SPECTRAmode=Fine-tuning2023.05 | 21.96 | |
| SPACE+WavLM+TripPymode=Fine-tuning2023.05 | 20.9 | |
| Joint Speech and Text E2E DSTTraining Speech=MW, LLM=Gemma-3-12B-it, Training Text=✗2025.11 | 20.7 | |
| Joint Speech and Text E2E DSTTraining Speech=MW, LLM=Gemma-3-4B-it, Training Text=✗2025.11 | 19.8 | |
| Joint Speech and Text E2E DSTTraining Speech=MW, LLM=Gemma-3-1B-it, Training Text=✗2025.11 | 18.9 | |
| Joint Speech and Text E2E DSTTraining Speech=MW, LLM=OLMo-1B, Training Text=✗2025.11 | 18.7 |