Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Speech interaction S2T | 1 | 1 | Apr 21, 2026 | |
| Speech-to-text Instruction Following | 1 | 1 | Apr 21, 2026 | |
| Zero-shot speaker adaptation | 1 | 1 | Feb 18, 2026 | |
| Emotion Recognition (Valence) | 1 | 1 | Feb 21, 2026 | |
| Omnimodal Audio & Video Question Answering | 1 | 1 | Apr 21, 2026 | |
| Emotion Recognition (Arousal) | 1 | 1 | Feb 21, 2026 | |
| Audio-visual alignment | 1 | 1 | Apr 21, 2026 | |
| Audio-visual video parsing (Segment-level) | 1 | 2 | Feb 18, 2026 | |
| 1 |
| 1 |
| Feb 21, 2026 |
| Token-level Predictability | 1 | 1 | Apr 21, 2026 |
|---|
| Speech Probing | 1 | 1 | Apr 21, 2026 |
|---|
| Audio-visual video parsing (Event-level) | 1 | 2 | Feb 18, 2026 |
|---|
| Multi-modal to Audio Generation Latency | 1 | 1 | Feb 21, 2026 |
|---|
| Text-to-Audio Retrieval (T2A) | 1 | 1 | Apr 21, 2026 |
|---|
| IMU-based Human Activity Recognition | 1 | 1 | Apr 23, 2026 |
|---|
| Audio-referring Video Object Segmentation | 1 | 1 | Apr 23, 2026 |
|---|
| Four-class Emotion Recognition | 1 | 1 | Feb 21, 2026 |
|---|
| Self-noise measurement | 1 | 1 | Apr 23, 2026 |
|---|
| Pitch accent classification | 1 | 1 | Apr 23, 2026 |
|---|
| Acoustic Scene Fake Detection | 1 | 1 | Apr 23, 2026 |
|---|
| Acoustic Event Fake Detection | 1 | 1 | Apr 23, 2026 |
|---|
| Singing voice separation metric correlation analysis | 1 | 1 | Apr 23, 2026 |
|---|
| Speech pretraining | 1 | 1 | Feb 21, 2026 |
|---|
| Musical Source Separation Quality Assessment | 1 | 1 | Apr 23, 2026 |
|---|
| Dropping | 1 | 2 | Feb 18, 2026 |
|---|