Utterance-level Audio Aesthetics Prediction on AMC 2025 (eval)
1.041PQ MSEProposed (VERSA)
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Proposed (VERSA)components=VERSA-based predictor2025.12 | 1.041 | 0.796 | 0.475 | 0.908 | 2.33 | 0.835 | 1.923 | 0.811 | |
| B02 (Baseline)2025.12 | 1.184 | 0.78 | 0.562 | 0.902 | 1.893 | 0.811 | 2.255 | 0.774 | |
| T12 (AESCA)components=Ensemble of KAN #1–#4 & VERSA2025.12 | 1.184 | 0.832 | 0.719 | 0.911 | 2.472 | 0.855 | 1.853 | 0.852 | |
| Proposed (KAN #1)components=Single KAN-based predictor2025.12 | 1.303 | 0.814 | 0.751 | 0.906 | 2.703 | 0.835 | 2.051 | 0.842 | |
| Proposed (KAN #1–#4)components=Ensemble of four KAN-based predictors2025.12 | 1.312 | 0.818 | 0.746 | 0.909 | 2.681 | 0.839 | 2.066 | 0.842 |