Speech Classification on VF
99.2AccuracyWav2Vec2
Evaluation Results
| Method | Links | |
|---|---|---|
| Wav2Vec2Evaluation protocol=Linear evaluation2022.10 | 99.2 | |
| ATST-FrameAudio enc. fine-tuning on AudioSet=false2025.03 | 98.8 | |
| M2D-AS# Params=86M, Masking ratio=0.7, Evaluation protocol=Linear evaluation2024.04 | 98.4 | |
| M2DAudio enc. fine-tuning on AudioSet=false2025.03 | 98.3 | |
| M2D-CLAP stage1 2024Audio enc. fine-tuning on AudioSet=false2025.03 | 98.3 | |
| M2D# Params=86M, Masking ratio=0.7, Evaluation protocol=Linear evaluation2024.04 | 98.3 | |
| M2DNumber of Parameters=86M, Masking ratio=0.7, Target Feeding=masked patches, Linear Evaluation=true2024.04 | 98.3 | |
| M2DNumber of Parameters=86M, Masking ratio=0.7, Target Feeding=all patches, Linear Evaluation=true2024.04 | 98.2 | |
| M2DNumber of Parameters=86M, Masking ratio=0.6, Target Feeding=masked patches, Linear Evaluation=true2024.04 | 98.2 | |
| M2D-CLAP stage1 2025Audio enc. fine-tuning on AudioSet=false2025.03 | 98.1 | |
| M2DEvaluation protocol=Linear evaluation, Masking ratio=0.62022.10 | 97.9 | |
| MSM-MAENumber of Parameters=86M, Masking ratio=0.75, Linear Evaluation=true2024.04 | 97.8 | |
| M2DNumber of Parameters=86M, Masking ratio=0.6, Target Feeding=all patches, Linear Evaluation=true2024.04 | 97.8 | |
| M2DEvaluation protocol=Linear evaluation, Masking ratio=0.72022.10 | 97.7 | |
| ATST-ClipAudio enc. fine-tuning on AudioSet=false2025.03 | 97.6 | |
| ATST BaseNumber of Parameters=86M, Linear Evaluation=true2024.04 | 97.6 | |
| MSM-MAEEvaluation protocol=Linear evaluation2022.10 | 97.5 | |
| ATST BaseEvaluation protocol=Linear evaluation, Variant=Base2022.10 | 97.4 | |
| M2D-CLAP2025Audio enc. fine-tuning on AudioSet=true2025.03 | 97.1 | |
| M2DASAudio enc. fine-tuning on AudioSet=true2025.03 | 96.9 | |
| SF NFNet-F0Evaluation protocol=Linear evaluation2022.10 | 95.4 | |
| CEDAudio enc. fine-tuning on AudioSet=false2025.03 | 94.8 | |
| BEATsiter3Audio enc. fine-tuning on AudioSet=false2025.03 | 94.1 | |
| BEATSiter3Number of Parameters=90M, Linear Evaluation=true2024.04 | 94.1 | |
| BYOL-AEvaluation protocol=Linear evaluation2022.10 | 93.3 | |
| BYOL-ANumber of Parameters=5.3M, Linear Evaluation=true2024.04 | 93.3 | |
| BEATsiter3+Audio enc. fine-tuning on AudioSet=true2025.03 | 92.5 | |
| BEATSiter3+# Params=90M, Evaluation protocol=Linear evaluation2024.04 | 92.5 | |
| BEATSiter3+Number of Parameters=90M, Linear Evaluation=true2024.04 | 92.5 | |
| SF NFNet-F0Number of Parameters=63M, Linear Evaluation=true2024.04 | 90.4 | |
| CLAP2023Audio enc. fine-tuning on AudioSet=true2025.03 | 89.6 | |
| ConformerXL-P Non-RAEvaluation protocol=Linear evaluation, Variant=Non-RA2022.10 | 88.2 | |
| DeLoRes-MEvaluation protocol=Linear evaluation2022.10 | 88 | |
| DeLoRes-MNumber of Parameters=5.3M, Linear Evaluation=true2024.04 | 88 | |
| AST-Fusion#5#12Evaluation protocol=Linear evaluation, Variant=#5#122022.10 | 87.6 | |
| AST-Fusion#5#12Number of Parameters=86M, Linear Evaluation=true2024.04 | 87.6 | |
| HTS-ATAudio enc. fine-tuning on AudioSet=true2025.03 | 82.3 | |
| HTS-AT# Params=31M, Evaluation protocol=Linear evaluation2024.04 | 82.3 | |
| HTS-ATNumber of Parameters=31M, Linear Evaluation=true2024.04 | 82.3 | |
| ASTAudio enc. fine-tuning on AudioSet=true2025.03 | 81.2 | |
| AST# Params=86M, Evaluation protocol=Linear evaluation2024.04 | 81.2 | |
| ASTNumber of Parameters=86M, Linear Evaluation=true2024.04 | 81.2 | |
| LAION-CLAPAudio enc. fine-tuning on AudioSet=true2025.03 | 80.3 | |
| WavCapsAudio enc. fine-tuning on AudioSet=true2025.03 | 80 | |
| BYOL-SEvaluation protocol=Linear evaluation2022.10 | 76.9 | |
| CLAP2022Audio enc. fine-tuning on AudioSet=true2025.03 | 75.8 | |
| PANNs CNN14Audio enc. fine-tuning on AudioSet=true2025.03 | 75.1 |