Gender Recognition on CelebA (test)
98.598AccuracyImage-Text Fusion
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Image-Text FusionText generation source=T2 (Attributes)2025.12 | 98.598 | 97.343 | 99.388 | 1.021 | 1.446 | |
| CLIPText generation source=T2 (Attributes)2025.12 | 98.572 | 98.0946 | 98.873 | 1.008 | 0.55 | |
| Image-Text FusionText generation source=T1 (BLIP Image-to-Text Model)2025.12 | 98.572 | 97.446 | 99.281 | 1.019 | 1.298 | |
| ITM GuidedText generation source=T2 (Attributes)2025.12 | 98.567 | 98.185 | 98.808 | 1.006 | 0.441 | |
| ITM GuidedText generation source=T1 (BLIP Image-to-Text Model)2025.12 | 98.467 | 97.369 | 99.159 | 1.018 | 1.266 | |
| Image OnlyInput Modal=Image Only2025.12 | 98.447 | 97.33 | 99.151 | 1.019 | 1.288 | |
| TandemNetText generation source=T1 (BLIP Image-to-Text Model)2025.12 | 98.437 | 97.99 | 98.718 | 1.007 | 0.515 | |
| TandemNetText generation source=T2 (Attributes)2025.12 | 98.402 | 97.874 | 98.734 | 1.009 | 0.608 | |
| CLIPText generation source=T1 (BLIP Image-to-Text Model)2025.12 | 98.294 | 97.174 | 99 | 1.019 | 1.291 | |
| LNets+ANet2016.03 | 98 | — | — | — | — | |
| HF-ResNetBackbone=ResNet2016.03 | 98 | — | — | — | — | |
| PANDA-l2016.03 | 97 | — | — | — | — | |
| Multitask_Face2016.03 | 97 | — | — | — | — | |
| HyperFace2016.03 | 97 | — | — | — | — | |
| [34]+ANet2016.03 | 95 | — | — | — | — | |
| R-CNN_Gender2016.03 | 95 | — | — | — | — | |
| PANDA-w2016.03 | 93 | — | — | — | — | |
| FaceTracer2016.03 | 91 | — | — | — | — |