Video-Text-to-Audio (VT2A) on VGGSound Omini (off-screen track)
0.97FADS1 → S2
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| S1 → S2Training Stages=Stage 1 and Stage 2, Off-screen Synthesis augmentation=false2026.01 | 0.97 | 1.46 | 0.31 | 0.468 | |
| S1 → S2 → S3Training Stages=Full Three-Stage Schedule, Off-screen Synthesis augmentation=true2026.01 | 0.85 | 1.39 | 0.32 | 0.532 |