Malware Classification on CICMalDroid 2020
97.47AccuracyXGBoost
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| XGBoostPreprocessing=Original (raw vectors), Aggregation=Best2026.05 | 97.47 | — | 97.03 | 97.31 | 97.16 | |
| XGBoostApproach=current2026.05 | 95.6 | — | 95.61 | 95.6 | 95.6 | |
| XGBoostPreprocessing=Synthetic data augmentation2026.05 | 95.6 | — | 95.61 | 95.6 | 95.6 | |
| Random forestApproach=current2026.05 | 95.3 | — | 95.34 | 95.3 | 95.29 | |
| Random forestPreprocessing=Synthetic data augmentation2026.05 | 95.3 | — | 95.34 | 95.3 | 95.29 | |
| XGBoostApproach=from [7]2026.05 | 95 | — | 95 | 95 | 95 | |
| MARD2026.04 | 94.35 | — | 89.29 | 98.04 | 93.46 | |
| Abedin and MehrubPreprocessing=Original (raw vectors), Aggregation=Average2026.05 | 93.97 | — | 93.16 | 93.05 | 93.05 | |
| ExtraTreesPreprocessing=LDA (projected onto class separating axes), Aggregation=Best2026.05 | 93.25 | — | 92.19 | 92.01 | 92.09 | |
| CL-Malware2026.04 | 92.74 | — | 85 | 100 | 91.89 | |
| ExtraTreesPreprocessing=PCA (retained components), Aggregation=Best2026.05 | 92.16 | — | 90.33 | 90.53 | 90.37 | |
| Abedin and MehrubPreprocessing=LDA (projected onto class separating axes), Aggregation=Average2026.05 | 92 | — | 90.95 | 90.56 | 90.73 | |
| Abedin and MehrubPreprocessing=PCA (retained components), Aggregation=Average2026.05 | 89.48 | — | 87.68 | 87.37 | 87.46 | |
| Sequential modelApproach=current2026.05 | 88.58 | — | 88.66 | 88.58 | 88.54 | |
| Sequential modelPreprocessing=Synthetic data augmentation2026.05 | 88.58 | — | 88.66 | 88.58 | 88.54 | |
| Sequential modelApproach=from [7]2026.05 | 86 | — | 86 | 86 | 86 | |
| MaMaDroidgranularity=family2026.04 | 60.16 | — | 51.04 | 96.08 | 66.67 | |
| Malscaconfig=knn-12026.04 | 58 | — | 52.94 | 60 | 56.25 | |
| MaMaDroidgranularity=package2026.04 | 48.78 | — | 44.44 | 94.12 | 60.38 | |
| Malscanconfig=random2026.04 | 46 | — | 44.71 | 84.44 | 58.46 | |
| Malscanconfig=knn-32026.04 | 44 | — | 42.03 | 64.44 | 50.88 | |
| DroidEvolver2026.04 | 40.32 | — | 40.65 | 98.04 | 57.47 | |
| HuBERT ⊞ ViT (FOCA)Modality=Multi-modal (Audio ⊞ Image), Audio Encoder=HuBERT, Image Encoder=ViT, Fusion=Hyperbolic cross-attention (FOCA ⊞)2026.01 | 0.991 | 0.9885 | — | — | — | |
| Yang et al.2026.01 | 0.9852 | 0.9831 | — | — | — | |
| Samaneh et al.2026.01 | 0.9673 | 0.9784 | — | — | — | |
| Vasan et al.2026.01 | 0.9571 | 0.9646 | — | — | — | |
| Scott et al.2026.01 | 0.9374 | 0.9181 | — | — | — | |
| Wav2Vec2 ⊗ ViTModality=Multi-modal (Audio ⊗ Image), Audio Encoder=Wav2Vec2, Image Encoder=ViT, Fusion=Euclidean cross-modal attention (⊗)2026.01 | 0.9321 | 0.919 | — | — | — | |
| Devnath et al.2026.01 | 0.9274 | 0.0094 | — | — | — | |
| HuBERT ⊗ ViTModality=Multi-modal (Audio ⊗ Image), Audio Encoder=HuBERT, Image Encoder=ViT, Fusion=Euclidean cross-modal attention (⊗)2026.01 | 0.9221 | 0.9189 | — | — | — | |
| Wav2Vec2 ⊞ ViT (FOCA)Modality=Multi-modal (Audio ⊞ Image), Audio Encoder=Wav2Vec2, Image Encoder=ViT, Fusion=Hyperbolic cross-attention (FOCA ⊞)2026.01 | 0.9121 | 0.899 | — | — | — | |
| Wav2Vec2 ⊗ VGG-19Modality=Multi-modal (Audio ⊗ Image), Audio Encoder=Wav2Vec2, Image Encoder=VGG-19, Fusion=Euclidean cross-modal attention (⊗)2026.01 | 0.8868 | 0.8437 | — | — | — | |
| Wav2Vec2 ⊞ ResNet-50 (FOCA)Modality=Multi-modal (Audio ⊞ Image), Audio Encoder=Wav2Vec2, Image Encoder=ResNet-50, Fusion=Hyperbolic cross-attention (FOCA ⊞)2026.01 | 0.8862 | 0.8735 | — | — | — | |
| WavLM ⊞ ViT (FOCA)Modality=Multi-modal (Audio ⊞ Image), Audio Encoder=WavLM, Image Encoder=ViT, Fusion=Hyperbolic cross-attention (FOCA ⊞)2026.01 | 0.8847 | 0.8723 | — | — | — | |
| Wav2Vec2 ⊞ VGG-19 (FOCA)Modality=Multi-modal (Audio ⊞ Image), Audio Encoder=Wav2Vec2, Image Encoder=VGG-19, Fusion=Hyperbolic cross-attention (FOCA ⊞)2026.01 | 0.8768 | 0.8643 | — | — | — | |
| WavLM ⊞ ResNet-50 (FOCA)Modality=Multi-modal (Audio ⊞ Image), Audio Encoder=WavLM, Image Encoder=ResNet-50, Fusion=Hyperbolic cross-attention (FOCA ⊞)2026.01 | 0.8756 | 0.8634 | — | — | — | |
| HuBERT ⊞ ResNet-50 (FOCA)Modality=Multi-modal (Audio ⊞ Image), Audio Encoder=HuBERT, Image Encoder=ResNet-50, Fusion=Hyperbolic cross-attention (FOCA ⊞)2026.01 | 0.8756 | 0.8631 | — | — | — | |
| Wav2Vec2 ⊗ ResNet-50Modality=Multi-modal (Audio ⊗ Image), Audio Encoder=Wav2Vec2, Image Encoder=ResNet-50, Fusion=Euclidean cross-modal attention (⊗)2026.01 | 0.8752 | 0.8521 | — | — | — | |
| WavLM ⊞ VGG-19 (FOCA)Modality=Multi-modal (Audio ⊞ Image), Audio Encoder=WavLM, Image Encoder=VGG-19, Fusion=Hyperbolic cross-attention (FOCA ⊞)2026.01 | 0.8665 | 0.8541 | — | — | — | |
| HuBERT ⊞ VGG-19 (FOCA)Modality=Multi-modal (Audio ⊞ Image), Audio Encoder=HuBERT, Image Encoder=VGG-19, Fusion=Hyperbolic cross-attention (FOCA ⊞)2026.01 | 0.8665 | 0.8542 | — | — | — | |
| WavLM ⊗ ViTModality=Multi-modal (Audio ⊗ Image), Audio Encoder=WavLM, Image Encoder=ViT, Fusion=Euclidean cross-modal attention (⊗)2026.01 | 0.8556 | 0.8525 | — | — | — | |
| HuBERT ⊗ VGG-19Modality=Multi-modal (Audio ⊗ Image), Audio Encoder=HuBERT, Image Encoder=VGG-19, Fusion=Euclidean cross-modal attention (⊗)2026.01 | 0.8475 | 0.8442 | — | — | — | |
| WavLM ⊗ VGG-19Modality=Multi-modal (Audio ⊗ Image), Audio Encoder=WavLM, Image Encoder=VGG-19, Fusion=Euclidean cross-modal attention (⊗)2026.01 | 0.8423 | 0.8394 | — | — | — | |
| HuBERT ⊗ ResNet-50Modality=Multi-modal (Audio ⊗ Image), Audio Encoder=HuBERT, Image Encoder=ResNet-50, Fusion=Euclidean cross-modal attention (⊗)2026.01 | 0.8394 | 0.8363 | — | — | — | |
| WavLM ⊗ ResNet-50Modality=Multi-modal (Audio ⊗ Image), Audio Encoder=WavLM, Image Encoder=ResNet-50, Fusion=Euclidean cross-modal attention (⊗)2026.01 | 0.8381 | 0.8351 | — | — | — | |
| Wav2Vec2 + ViTModality=Multi-modal (Audio + Image), Audio Encoder=Wav2Vec2, Image Encoder=ViT, Fusion=Concatenation (+)2026.01 | 0.8221 | 0.819 | — | — | — | |
| HuBERTModality=Audio, Audio Encoder=HuBERT2026.01 | 0.8098 | 0.788 | — | — | — | |
| HuBERT + ViTModality=Multi-modal (Audio + Image), Audio Encoder=HuBERT, Image Encoder=ViT, Fusion=Concatenation (+)2026.01 | 0.8016 | 0.7985 | — | — | — | |
| WavLM + ViTModality=Multi-modal (Audio + Image), Audio Encoder=WavLM, Image Encoder=ViT, Fusion=Concatenation (+)2026.01 | 0.7974 | 0.7944 | — | — | — | |
| Wav2Vec2 + ResNet-50Modality=Multi-modal (Audio + Image), Audio Encoder=Wav2Vec2, Image Encoder=ResNet-50, Fusion=Concatenation (+)2026.01 | 0.7974 | 0.7945 | — | — | — | |
| Wav2Vec2 + VGG-19Modality=Multi-modal (Audio + Image), Audio Encoder=Wav2Vec2, Image Encoder=VGG-19, Fusion=Concatenation (+)2026.01 | 0.7892 | 0.7863 | — | — | — | |
| HuBERT + VGG-19Modality=Multi-modal (Audio + Image), Audio Encoder=HuBERT, Image Encoder=VGG-19, Fusion=Concatenation (+)2026.01 | 0.7892 | 0.7865 | — | — | — | |
| WavLM + VGG-19Modality=Multi-modal (Audio + Image), Audio Encoder=WavLM, Image Encoder=VGG-19, Fusion=Concatenation (+)2026.01 | 0.7853 | 0.7823 | — | — | — | |
| WavLM + ResNet-50Modality=Multi-modal (Audio + Image), Audio Encoder=WavLM, Image Encoder=ResNet-50, Fusion=Concatenation (+)2026.01 | 0.781 | 0.7783 | — | — | — | |
| HuBERT + ResNet-50Modality=Multi-modal (Audio + Image), Audio Encoder=HuBERT, Image Encoder=ResNet-50, Fusion=Concatenation (+)2026.01 | 0.781 | 0.7785 | — | — | — | |
| Wav2Vec2Modality=Audio, Audio Encoder=Wav2Vec22026.01 | 0.7612 | 0.7407 | — | — | — | |
| ViTModality=Image, Image Encoder=ViT2026.01 | 0.749 | 0.7448 | — | — | — | |
| WavLMModality=Audio, Audio Encoder=WavLM2026.01 | 0.7369 | 0.7171 | — | — | — | |
| VGG-19Modality=Image, Image Encoder=VGG-192026.01 | 0.7265 | 0.7225 | — | — | — | |
| ResNet-50Modality=Image, Image Encoder=ResNet-502026.01 | 0.7118 | 0.7076 | — | — | — |