Audio Classification on VocalSound (test)
85.52AccuracyMUKA
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| MUKATraining Approach=Training-free, Shots=16, Backbone=PENGI2026.02 | 85.52 | — | |
| PaLMTraining Approach=Training-based, Shots=16, Backbone=PENGI2026.02 | 80.78 | — | |
| transductive GMM-EMPrompt=ensemble, Batch size=full, Evaluation protocol=Transductive GMM-EM, Iterations=32026.06 | 79.84 | 4.57 | |
| transductive GMM-EMPrompt=ensemble, Batch size=256, Evaluation protocol=Transductive GMM-EM, Iterations=32026.06 | 79.62 | — | |
| transductive GMM-EMPrompt=ensemble, Batch size=64, Evaluation protocol=Transductive GMM-EM, Iterations=32026.06 | 79.19 | — | |
| Linear ProbingTraining Approach=Training-based, Shots=16, Backbone=PENGI2026.02 | 78.1 | — | |
| CoCoOpTraining Approach=Training-based, Shots=16, Backbone=PENGI2026.02 | 77.9 | — | |
| CLAPPrompt=ensemble, Evaluation protocol=Zero-shot2026.06 | 75.27 | — | |
| transductive GMM-EMPrompt=single, Batch size=full, Evaluation protocol=Transductive GMM-EM, Iterations=32026.06 | 73.99 | 8.27 | |
| Treff-AdapterTraining Approach=Training-free, Shots=16, Backbone=PENGI2026.02 | 73.94 | — | |
| transductive GMM-EMPrompt=single, Batch size=256, Evaluation protocol=Transductive GMM-EM, Iterations=32026.06 | 73.71 | — | |
| transductive GMM-EMPrompt=single, Batch size=64, Evaluation protocol=Transductive GMM-EM, Iterations=32026.06 | 73.67 | — | |
| CoOpTraining Approach=Training-based, Shots=16, Backbone=PENGI2026.02 | 70.96 | — | |
| CLAPPrompt=single, Evaluation protocol=Zero-shot2026.06 | 65.72 | — | |
| Zero-ShotTraining Approach=Training-free, Shots=0, Backbone=PENGI2026.02 | 41.97 | — |