Image Clustering on DTD
69.7NMIPRCut*
Evaluation Results
| Method | Links | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| PRCut*Embedding Model=DINOv3-B, K=47, Q=0.472025.11 | 69.7 | 63.3 | — | — | — | — | — | — | — | — | — | 832 | 20.5 | |
| Spectral ClusteringEmbedding Model=DINOv3-B, K=47, Q=0.472025.11 | 68.3 | 61 | — | — | — | — | — | — | — | — | — | 692 | 18 | |
| H-NCutEmbedding Model=DINOv3-B, K=47, Q=0.472025.11 | 67.6 | 59.7 | — | — | — | — | — | — | — | — | — | 870 | 21 | |
| H-RCutEmbedding Model=DINOv3-B, K=47, Q=0.472025.11 | 67.5 | 60 | — | — | — | — | — | — | — | — | — | 868 | 20.9 | |
| DCRBackbone=SigLIP ViT-SO@224 (S-1), Zero-shot=true2026.03 | 67 | 59 | 44 | — | — | — | — | — | — | — | — | — | — | |
| Original CLIPBackbone=SigLIP ViT-SO@224 (S-1), Zero-shot=true2026.03 | 66 | 55 | 40 | — | — | — | — | — | — | — | — | — | — | |
| DIVABackbone=SigLIP ViT-SO@224 (S-1), Zero-shot=true2026.03 | 66 | 55 | 41 | — | — | — | — | — | — | — | — | — | — | |
| H-NCutEmbedding=DINOv2, Graph type=50-NN RBF, K=47, Q=0.472025.11 | 65.8 | 57.8 | — | — | — | — | — | — | — | — | — | 776 | 17.6 | |
| PRCut*Embedding=DINOv2, Graph type=50-NN RBF, K=47, Q=0.472025.11 | 65.7 | 57.9 | — | — | — | — | — | — | — | — | — | 760 | 18.3 | |
| H-RCutEmbedding=DINOv2, Graph type=50-NN RBF, K=47, Q=0.472025.11 | 64.3 | 55.1 | — | — | — | — | — | — | — | — | — | 843 | 18.1 | |
| Spectral ClusteringEmbedding=DINOv2, Graph type=50-NN RBF, K=47, Q=0.472025.11 | 64.1 | 54.9 | — | — | — | — | — | — | — | — | — | 575 | 15 | |
| TAC_TURTLE2026.04 | 63.3 | 52.9 | 36.8 | — | — | — | — | — | — | — | — | — | — | |
| GradNorm2025.10 | 63.1 | 50.9 | 34.2 | — | — | — | — | — | — | — | — | — | — | |
| KEC_TURTLE2026.04 | 63.1 | 52.8 | 36.7 | — | — | — | — | — | — | — | — | — | — | |
| TURTLE (1-space)Training Protocol=1-space2026.04 | 62.9 | 52.9 | 36.7 | — | — | — | — | — | — | — | — | — | — | |
| KEC_TAC2026.04 | 62.5 | 51.3 | 36 | — | — | — | — | — | — | — | — | — | — | |
| TACClustering setting=default2023.10 | 62.1 | 50.1 | 34.4 | — | — | — | — | — | — | — | — | — | — | |
| GenHancerBackbone=SigLIP ViT-SO@224 (S-1), Zero-shot=true2026.03 | 62 | 55 | 42 | — | — | — | — | — | — | — | — | — | — | |
| H-NCutK=47, Q=0.38, Embeddings=CLIP ViT-L/142025.11 | 61.2 | 53.9 | — | — | — | — | — | — | — | — | — | 995 | 24.2 | |
| PRCut*K=47, Q=0.38, Embeddings=CLIP ViT-L/142025.11 | 61.1 | 52.7 | — | — | — | — | — | — | — | — | — | 938 | 25.5 | |
| H-RCutK=47, Q=0.38, Embeddings=CLIP ViT-L/142025.11 | 61.1 | 54.1 | — | — | — | — | — | — | — | — | — | 964 | 24.5 | |
| Spectral ClusteringK=47, Q=0.38, Embeddings=CLIP ViT-L/142025.11 | 60.9 | 54.3 | — | — | — | — | — | — | — | — | — | 860 | 21.6 | |
| TAC2026.04 | 60.8 | 47.8 | 32.4 | — | — | — | — | — | — | — | — | — | — | |
| KEC (no train)Training Protocol=no train2026.04 | 60.7 | 47.4 | 31.5 | — | — | — | — | — | — | — | — | — | — | |
| TAC (no train)Training Protocol=no train2026.04 | 60.3 | 48.2 | 31.4 | — | — | — | — | — | — | — | — | — | — | |
| TACClustering setting=no train2023.10 | 60.1 | 45.9 | 29 | — | — | — | — | — | — | — | — | — | — | |
| TAC2025.10 | 60.1 | 45.9 | 29 | — | — | — | — | — | — | — | — | — | — | |
| DIVABackbone=MetaCLIP ViT-L@224 (M-1), Zero-shot=true2026.03 | 60 | 52 | 36 | — | — | — | — | — | — | — | — | — | — | |
| DCRBackbone=MetaCLIP ViT-L@224 (M-1), Zero-shot=true2026.03 | 60 | 54 | 37 | — | — | — | — | — | — | — | — | — | — | |
| SICClustering setting=default2023.10 | 59.6 | 45.9 | 30.5 | — | — | — | — | — | — | — | — | — | — | |
| SIC2025.10 | 59.6 | 45.9 | 30.5 | — | — | — | — | — | — | — | — | — | — | |
| SIC2026.04 | 59.6 | 45.9 | 30.5 | — | — | — | — | — | — | — | — | — | — | |
| SCANClustering setting=default2023.10 | 59.4 | 46.4 | 31.7 | — | — | — | — | — | — | — | — | — | — | |
| SCAN2025.10 | 59.4 | 46.4 | 31.7 | — | — | — | — | — | — | — | — | — | — | |
| GenHancerBackbone=MetaCLIP ViT-L@224 (M-1), Zero-shot=true2026.03 | 59 | 49 | 34 | — | — | — | — | — | — | — | — | — | — | |
| CLIP (k-means)Clustering Strategy=k-means2026.04 | 58.6 | 45.4 | 29 | — | — | — | — | — | — | — | — | — | — | |
| Finsler t-SNEEmbedding=Finsler t-SNE, Clustering Algorithm=kMeans2026.03 | 58.5 | — | 29.4 | 50.4 | 58 | 59 | 58.5 | 31 | 40.8 | 0.859 | — | — | — | |
| t-SNEEmbedding=t-SNE, Clustering Algorithm=kMeans2026.03 | 58.1 | — | 28.8 | 49.8 | 57.8 | 58.4 | 58.1 | 30.3 | 40.2 | 0.804 | — | — | — | |
| DCRBackbone=OpenAI CLIP ViT-L@224 (O-1), Zero-shot=true2026.03 | 58 | 48 | 33 | — | — | — | — | — | — | — | — | — | — | |
| CLIPClustering setting=k-means2023.10 | 57.3 | 42.6 | 27.4 | — | — | — | — | — | — | — | — | — | — | |
| CLIPEvaluation Protocol=k-means2025.10 | 57.3 | 42.6 | 27.4 | — | — | — | — | — | — | — | — | — | — | |
| CLIPClustering setting=zero-shot2023.10 | 56.5 | 43.1 | 26.9 | — | — | — | — | — | — | — | — | — | — | |
| CLIPEvaluation Protocol=zero-shot2025.10 | 56.5 | 43.1 | 26.9 | — | — | — | — | — | — | — | — | — | — | |
| CLIP (zero-shot)Training Protocol=zero-shot2026.04 | 56.1 | 42.6 | 26.6 | — | — | — | — | — | — | — | — | — | — | |
| GenHancerBackbone=OpenAI CLIP ViT-L@224 (O-1), Zero-shot=true2026.03 | 56 | 45 | 31 | — | — | — | — | — | — | — | — | — | — | |
| Original CLIPBackbone=MetaCLIP ViT-L@224 (M-1), Zero-shot=true2026.03 | 56 | 51 | 36 | — | — | — | — | — | — | — | — | — | — | |
| Original CLIPBackbone=OpenAI CLIP ViT-L@224 (O-1), Zero-shot=true2026.03 | 55 | 45 | 30 | — | — | — | — | — | — | — | — | — | — | |
| un2CLIPBackbone=OpenAI CLIP ViT-L@224 (O-1), Zero-shot=true2026.03 | 55 | 46 | 31 | — | — | — | — | — | — | — | — | — | — | |
| DIVABackbone=OpenAI CLIP ViT-L@224 (O-1), Zero-shot=true2026.03 | 51 | 41 | 26 | — | — | — | — | — | — | — | — | — | — |