Audio Visual Question Answering on AVQA (Robustness Evaluation)
95.6AVQA Clean AccuracyNegative Language Modeling Loss
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Negative Language Modeling LossObjective=L_negLM2026.01 | 95.6 | 85 | 10.6 | |
| Encoder-Based Cosine Similarity LossObjective=L^(cos)2026.01 | 95.6 | 11 | 84.6 | |
| Vision Attention Suppression LossObjective=L^(visionatt)2026.01 | 95.6 | 92 | 3.6 | |
| Audio Attention Amplification LossObjective=L^(audioatt)2026.01 | 95.6 | 41 | 54.6 | |
| Attention Randomization LossObjective=L^(randatt)2026.01 | 95.6 | 91 | 4.6 | |
| Hidden-State Similarity LossObjective=L^(hidden-cos)2026.01 | 95.6 | 79.5 | 16.1 | |
| Combined Loss (SOUNDBREAK)Objective=L^(combined)2026.01 | 95.6 | 3.9 | 91.8 |