Multimodal Multilabel Classification on MM-IMDB (test)
63.2Macro F1MMBT-Large
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| MMBT-LargeModel Scale=Large2019.09 | 63.2 | 68 | — | — | — | — | — | |
| ViLBertFine-tuning Source=Standard pre-trained2019.09 | 63 | 68.6 | — | — | — | — | — | |
| PMF-large (M=4, Lf=22)Updated Param. (Million)=4.5, Train Memory Usage (GB)=18.44, Inference Memory Usage (GB)=6.42, Backbone=bert-large / vit-large, M (prompt length)=4, Lf (starting fusion layer)=222023.04 | 61.66 | 66.72 | — | — | — | — | — | |
| MMBT2019.09 | 61.6 | 66.8 | — | — | — | — | — | |
| MMBTModel Scale=Base2019.09 | 61.6 | 66.8 | — | — | — | — | — | |
| ViLBert-VCRFine-tuning Source=VCR2019.09 | 61.6 | 67.6 | — | — | — | — | — | |
| ViLBert-RefcocoFine-tuning Source=Refcoco2019.09 | 61.4 | 67.7 | — | — | — | — | — | |
| ViLBert-Flickr30kFine-tuning Source=Flickr30k2019.09 | 61.4 | 67.8 | — | — | — | — | — | |
| MMBT*Updated Param. (Million)=196.5, Train Memory Usage (GB)=37.87, Inference Memory Usage (GB)=3.48, Backbone=bert-base / vit-base2023.04 | 60.8 | 66.1 | — | — | — | — | — | |
| ConcatBertDescription=Concatenation of Bert and Img baselines2019.09 | 60.5 | 65.9 | — | — | — | — | — | |
| ViLBert-VQAFine-tuning Source=VQA2019.09 | 60 | 66.4 | — | — | — | — | — | |
| BertModality=Text-only, Backbone=BERT base-uncased2019.09 | 59.9 | 65.4 | — | — | — | — | — | |
| FiLMBertDescription=FiLM + BERT, Backbone=ResNet-152 features2019.09 | 59.7 | 65.1 | — | — | — | — | — | |
| MBT*Updated Param. (Million)=196.0, Train Memory Usage (GB)=38, Inference Memory Usage (GB)=4.06, Backbone=bert-base / vit-base2023.04 | 59.6 | 64.81 | — | — | — | — | — | |
| LateConcatUpdated Param. (Million)=196.0, Train Memory Usage (GB)=38.54, Inference Memory Usage (GB)=3.36, Backbone=bert-base / vit-base2023.04 | 59.56 | 64.92 | — | — | — | — | — | |
| Late FusionDescription=Averages scores of Bert and Img classifiers2019.09 | 59.4 | 66.2 | — | — | — | — | — | |
| BERTUpdated Param. (Million)=109.0, Train Memory Usage (GB)=30.82, Inference Memory Usage (GB)=2.79, Backbone=bert-base2023.04 | 58.91 | 64.31 | — | — | — | — | — | |
| PMF (M=4, Lf=10)Updated Param. (Million)=2.5, Train Memory Usage (GB)=12.84, Inference Memory Usage (GB)=4.08, Backbone=bert-base / vit-base, M (prompt length)=4, Lf (starting fusion layer)=102023.04 | 58.77 | 64.51 | — | — | — | — | — | |
| Scalable Diffusion ModelMissing Type=Image, Zero-Shot=true, Backbone=CLIP ViT-B/32, Layers=202026.02 | 58.22 | — | — | — | — | — | — | |
| Scalable Diffusion ModelMissing Type=Both, Zero-Shot=true, Backbone=CLIP ViT-B/32, Layers=202026.02 | 57.49 | — | — | — | — | — | — | |
| Scalable Diffusion ModelMissing Type=Text, Zero-Shot=true, Backbone=CLIP ViT-B/32, Layers=202026.02 | 57.24 | — | — | — | — | — | — | |
| REDEEMMissing Type=Text, Zero-Shot=true2026.02 | 56.32 | — | — | — | — | — | — | |
| CentralNet2019.09 | 56.1 | 63.9 | — | — | — | — | — | |
| CentralNetImage Modality=true, Text Modality=true2022.04 | 56.1 | 63.9 | 63.1 | 63.9 | — | — | — | |
| MFASImage Modality=true, Text Modality=true2022.04 | 55.7 | — | 62.5 | — | — | — | — | |
| ViLTImage Modality=true, Text Modality=true2022.04 | 55.3 | 64.7 | 64.4 | 64.6 | — | — | — | |
| RAGPTMissing Type=Text, Zero-Shot=true2026.02 | 54.33 | — | — | — | — | — | — | |
| P-LateConcatUpdated Param. (Million)=0.3, Train Memory Usage (GB)=30.82, Inference Memory Usage (GB)=3.43, Backbone=bert-base / vit-base, Prompt Length=102023.04 | 53.91 | 59.93 | — | — | — | — | — | |
| CL-DMDFModel Category=Graph-Based2026.06 | 53.28 | 63.25 | — | — | — | — | — | |
| DCPMissing rate η=70%, Image presence (%)=30%, Text presence (%)=100%2024.10 | 53.14 | — | — | — | — | — | — | |
| P-MMBTUpdated Param. (Million)=0.9, Train Memory Usage (GB)=30.9, Inference Memory Usage (GB)=3.48, Backbone=bert-base / vit-base, Prompt Length=102023.04 | 52.95 | 59.3 | — | — | — | — | — | |
| ViLTImage Modality=false, Text Modality=true2022.04 | 52.5 | 63.3 | 62 | 62.9 | — | — | — | |
| DePTMissing rate η=70%, Image presence (%)=30%, Text presence (%)=100%2024.10 | 52.13 | — | — | — | — | — | — | |
| DynMM-dModality=I+T, Computation Cost (MAdds)=12.1M2022.03 | 51.6 | 60.35 | — | — | — | — | — | |
| DynMMModel Category=Attention-Based2026.06 | 51.6 | 60.35 | — | — | — | — | — | |
| ReFNetModel Category=Encoder-Decoder2026.06 | 51.51 | 59.45 | — | — | — | — | — | |
| GMU2019.09 | 51.4 | 63 | — | — | — | — | — | |
| DynMM-cModality=I+T, Computation Cost (MAdds)=9.8M2022.03 | 51.2 | 59.72 | — | — | — | — | — | |
| DCPMissing Type=Image, Zero-Shot=true2026.02 | 51.15 | — | — | — | — | — | — | |
| Late Fusion (E2)Modality=I+T, Computation Cost (MAdds)=10.3M2022.03 | 50.94 | 59.55 | — | — | — | — | — | |
| LRMFModel Category=Encoder-Decoder2026.06 | 50.73 | 58.95 | — | — | — | — | — | |
| MaPLeMissing rate η=70%, Image presence (%)=30%, Text presence (%)=100%2024.10 | 50.64 | — | — | — | — | — | — | |
| MMPMissing rate η=70%, Image presence (%)=30%, Text presence (%)=100%2024.10 | 50.52 | — | — | — | — | — | — | |
| REDEEMMissing Type=Both, Zero-Shot=true2026.02 | 50.52 | — | — | — | — | — | — | |
| CCAModel Category=Encoder-Decoder2026.06 | 50.45 | 60.31 | — | — | — | — | — | |
| DynMM-bModality=I+T, Computation Cost (MAdds)=7.8M2022.03 | 50.42 | 59.59 | — | — | — | — | — | |
| BlindPromptUpdated Param. (Million)=< 0.1, Train Memory Usage (GB)=29.57, Inference Memory Usage (GB)=3.65, Backbone=bert-base / vit-base, Prompt Length=202023.04 | 50.18 | 56.46 | — | — | — | — | — | |
| SyPMissing Type=Image, Zero-Shot=true2026.02 | 50.16 | — | — | — | — | — | — | |
| RMFEModel Category=Attention-Based2026.06 | 49.82 | 58.67 | — | — | — | — | — | |
| LinearUpdated Param. (Million)=< 0.1, Train Memory Usage (GB)=3.76, Inference Memory Usage (GB)=3.23, Backbone=bert-base / vit-base2023.04 | 49.76 | 56.83 | — | — | — | — | — | |
| RAGPTMissing Type=Both, Zero-Shot=true2026.02 | 49.55 | — | — | — | — | — | — | |
| LRTFModality=I+T, Computation Cost (MAdds)=10.3M2022.03 | 49.26 | 59.18 | — | — | — | — | — | |
| MFASImage Modality=false, Text Modality=true2022.04 | 48.9 | 60.2 | 58.5 | 60.6 | — | — | — | |
| DynMM-aModality=I+T, Computation Cost (MAdds)=1.6M2022.03 | 48.84 | 59.57 | — | — | — | — | — | |
| CoOpMissing rate η=70%, Image presence (%)=30%, Text presence (%)=100%2024.10 | 48.82 | — | — | — | — | — | — | |
| P-BERTUpdated Param. (Million)=< 0.1, Train Memory Usage (GB)=28.13, Inference Memory Usage (GB)=2.99, Backbone=bert-base, Prompt Length=102023.04 | 48.67 | 54.58 | — | — | — | — | — | |
| MDPMissing Type=Image, Zero-Shot=true2026.02 | 48.61 | — | — | — | — | — | — | |
| PromptFuseUpdated Param. (Million)=< 0.1, Train Memory Usage (GB)=29.57, Inference Memory Usage (GB)=3.55, Backbone=bert-base / vit-base, Prompt Length=202023.04 | 48.59 | 54.49 | — | — | — | — | — | |
| MFMModel Category=Encoder-Decoder2026.06 | 48.53 | 56.44 | — | — | — | — | — | |
| MI-MatrixModality=I+T, Computation Cost (MAdds)=10.3M2022.03 | 48.36 | 58.45 | — | — | — | — | — | |
| SyPMissing Type=Both, Zero-Shot=true2026.02 | 48.02 | — | — | — | — | — | — | |
| Unimodal TextModel Category=Unimodal2026.06 | 47.59 | 59.37 | — | — | — | — | — | |
| DCPMissing Type=Text, Zero-Shot=true2026.02 | 47.58 | — | — | — | — | — | — | |
| DCPMissing Type=Both, Zero-Shot=true2026.02 | 47.47 | — | — | — | — | — | — | |
| MAPMissing Type=Text, Zero-Shot=true2026.02 | 47.44 | — | — | — | — | — | — | |
| Text Network (E1)Modality=T, Computation Cost (MAdds)=0.7M2022.03 | 47.21 | 59.16 | — | — | — | — | — | |
| REDEEMMissing Type=Image, Zero-Shot=true2026.02 | 47.2 | — | — | — | — | — | — | |
| MI-MatrixModel Category=Graph-Based2026.06 | 46.77 | 55.87 | — | — | — | — | — | |
| OursTraining Text Ratio=30%, Testing Text Ratio=30%, Training Image Ratio=100%, Testing Image Ratio=100%2022.04 | 46.6 | — | — | — | — | — | — | |
| SyPMissing Type=Text, Zero-Shot=true2026.02 | 46.52 | — | — | — | — | — | — | |
| Input-level promptsMissing rate η=70%, Training Image=30%, Training Text=100%, Testing Image=30%, Testing Text=100%2023.03 | 46.3 | — | — | — | — | — | — | |
| CentralNetImage Modality=false, Text Modality=true2022.04 | 45.9 | — | 57.5 | — | — | — | — | |
| Attention-level promptsMissing rate η=70%, Training Image=30%, Training Text=100%, Testing Image=30%, Testing Text=100%2023.03 | 44.74 | — | — | — | — | — | — | |
| ConcatBowDescription=Concatenation of Bow and Img baselines2019.09 | 43.8 | 53.6 | — | — | — | — | — | |
| Input-level promptsMissing rate η=70%, Training Image=65%, Training Text=65%, Testing Image=65%, Testing Text=65%2023.03 | 42.66 | — | — | — | — | — | — | |
| MDPMissing Type=Both, Zero-Shot=true2026.02 | 41.62 | — | — | — | — | — | — | |
| Attention-level promptsMissing rate η=70%, Training Image=65%, Training Text=65%, Testing Image=65%, Testing Text=65%2023.03 | 41.56 | — | — | — | — | — | — | |
| New baselineTraining Text Ratio=30%, Testing Text Ratio=30%, Training Image Ratio=100%, Testing Image Ratio=100%2022.04 | 40.4 | — | — | — | — | — | — | |
| MDPMissing Type=Text, Zero-Shot=true2026.02 | 40.36 | — | — | — | — | — | — | |
| AOEPTBackbone=ViLT, Missing Rate=70%2026.05 | 39.86 | — | — | — | 37.46 | 42.23 | 39.89 | |
| Input-level promptsMissing rate η=70%, Training Image=100%, Training Text=30%, Testing Image=100%, Testing Text=30%2023.03 | 39.22 | — | — | — | — | — | — | |
| MAPMissing Type=Both, Zero-Shot=true2026.02 | 39.13 | — | — | — | — | — | — | |
| RAGPTMissing Type=Image, Zero-Shot=true2026.02 | 38.82 | — | — | — | — | — | — | |
| MAPMissing Type=Image, Zero-Shot=true2026.02 | 38.74 | — | — | — | — | — | — | |
| ViTUpdated Param. (Million)=86.5, Train Memory Usage (GB)=9.36, Inference Memory Usage (GB)=1.99, Backbone=vit-base2023.04 | 38.39 | 49.88 | — | — | — | — | — | |
| Attention-level promptsMissing rate η=70%, Training Image=100%, Training Text=30%, Testing Image=100%, Testing Text=30%2023.03 | 38.16 | — | — | — | — | — | — | |
| BowModality=Text-only2019.09 | 38.1 | 45.6 | — | — | — | — | — | |
| MemPromptBackbone=ViLT, Missing Rate=70%2026.05 | 38.07 | — | — | — | 35.4 | 40.58 | 38.23 | |
| BaselineMissing rate η=70%, Training Image=30%, Training Text=100%, Testing Image=30%, Testing Text=100%2023.03 | 37.73 | — | — | — | — | — | — | |
| RAGPTBackbone=ViLT, Missing Rate=70%2026.05 | 37.61 | — | — | — | 36.19 | 39.9 | 36.74 | |
| SyPBackbone=ViLT, Missing Rate=70%2026.05 | 36.34 | — | — | — | 34.55 | 39.66 | 34.81 | |
| BaselineMissing rate η=70%, Training Image=65%, Training Text=65%, Testing Image=65%, Testing Text=65%2023.03 | 36.26 | — | — | — | — | — | — | |
| DCPBackbone=ViLT, Missing Rate=70%2026.05 | 36.06 | — | — | — | 34.15 | 38.18 | 35.86 | |
| MAPsBackbone=ViLT, Missing Rate=70%2026.05 | 35.83 | — | — | — | 35.29 | 36.92 | 35.28 | |
| VPTUpdated Param. (Million)=< 0.1, Train Memory Usage (GB)=6.12, Inference Memory Usage (GB)=2.01, Backbone=vit-base, Prompt Length=102023.04 | 35.22 | 44.49 | — | — | — | — | — | |
| BaselineMissing rate η=70%, Training Image=100%, Training Text=30%, Testing Image=100%, Testing Text=30%2023.03 | 35.13 | — | — | — | — | — | — | |
| ViLTImage Modality=true, Text Modality=false2022.04 | 35 | 51.8 | 48 | 51.1 | — | — | — | |
| CentralNetImage Modality=true, Text Modality=false2022.04 | 33.5 | — | 49.2 | — | — | — | — | |
| ImgModality=Image-only, Backbone=ResNet-1522019.09 | 32.5 | 44.4 | — | — | — | — | — | |
| Image onlyModality=Image-only, Testing Text Ratio=0%2022.04 | 31.2 | — | — | — | — | — | — |