Composed Image Retrieval on CIRCO 1.0 (test)
42.2mAP@5MMRet-MLLM
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| MMRet-MLLMBackbone=LLaVA-1.6, # Params=7.57B, Zero-shot evaluation=true2024.12 | 42.2 | — | — | |
| MMRet-LargeBackbone=CLIP-L, # Params=428M, Zero-shot evaluation=true2024.12 | 39.2 | — | — | |
| MMRet-BaseBackbone=CLIP-B, # Params=149M, Zero-shot evaluation=true2024.12 | 34.3 | — | — | |
| MagicLens-L*Backbone=CoCa-L, # Params=613M, Zero-shot evaluation=true2024.12 | 34.1 | — | — | |
| IP-CIRBackbone=CLIP-G, # Params=43.8B+, Zero-shot evaluation=true2024.12 | 32.8 | — | — | |
| MM-EmbedBackbone=LLaVA-1.6, # Params=7.57B, Zero-shot evaluation=true2024.12 | 32.3 | — | — | |
| LDREBackbone=CLIP-G, # Params=10.3B+, Zero-shot evaluation=true2024.12 | 31.1 | — | — | |
| MagicLens-B*Backbone=CoCa-B, # Params=267M, Zero-shot evaluation=true2024.12 | 30.8 | — | — | |
| MagicLens-LBackbone=CLIP-L, # Params=465M, Zero-shot evaluation=true2024.12 | 29.6 | — | — | |
| CIREVLBackbone=CLIP-G, # Params=14.6B+, Zero-shot evaluation=true2024.12 | 26.8 | — | — | |
| LDREBackbone=CLIP-L, # Params=8.2B+, Zero-shot evaluation=true2024.12 | 23.4 | — | — | |
| MagicLens-BBackbone=CLIP-B, # Params=166M, Zero-shot evaluation=true2024.12 | 23.1 | — | — | |
| CoLLMVision Encoder=OpenAI CLIP-L/14, Evaluation Protocol=Zero-shot, Incorporates LLM=true2025.03 | 20.3 | 20.8 | 23.4 | |
| CoLLMVision Encoder=BLIP-L/16, Evaluation Protocol=Zero-shot, Incorporates LLM=true2025.03 | 19.7 | 20.4 | 23.1 | |
| E5-VBackbone=LLaVA-1.6, # Params=8.35B, Zero-shot evaluation=true2024.12 | 19.1 | — | — | |
| CIREVLBackbone=CLIP-L, # Params=12.5B+, Zero-shot evaluation=true2024.12 | 18.6 | — | — | |
| CIREVLVision Encoder=OpenAI CLIP-L/14, Evaluation Protocol=Zero-shot, Incorporates LLM=true2025.03 | 18.6 | 19 | 21.8 | |
| Slerp-TATVision Encoder=OpenAI CLIP-L/14, Evaluation Protocol=Zero-shot2025.03 | 18.5 | 19.4 | 21.4 | |
| LDREBackbone=CLIP-B, # Params=7.9B+, Zero-shot evaluation=true2024.12 | 18 | — | — | |
| MCLVision Encoder=OpenAI CLIP-L/14, Evaluation Protocol=Zero-shot, Incorporates LLM=true2025.03 | 17.8 | 18.4 | 21.8 | |
| CIREVLBackbone=CLIP-B, # Params=12.3B+, Zero-shot evaluation=true2024.12 | 14.9 | — | — | |
| CIREVLVision Encoder=OpenAI CLIP-B/32, Evaluation Protocol=Zero-shot, Incorporates LLM=true2025.03 | 14.9 | 15.4 | 17.8 | |
| Slerp-TATVision Encoder=BLIP-L/16, Evaluation Protocol=Zero-shot2025.03 | 13.8 | 18.4 | 21.1 | |
| LinCIRVision Encoder=OpenAI CLIP-L/14, Evaluation Protocol=Zero-shot, Reproduced Results=true2025.03 | 13 | 13.9 | 16.2 | |
| ContextI2WVision Encoder=OpenAI CLIP-L/14, Evaluation Protocol=Zero-shot2025.03 | 13 | 13.8 | 16 | |
| CoLLMVision Encoder=OpenAI CLIP-B/32, Evaluation Protocol=Zero-shot, Incorporates LLM=true2025.03 | 12.9 | 13.2 | 15 | |
| CompoDiffBackbone=CLIP-L, # Params=568M, Zero-shot evaluation=true2024.12 | 12.6 | — | — | |
| SEARLEBackbone=CLIP-L, # Params=442M, Zero-shot evaluation=true2024.12 | 11.7 | — | — | |
| SEARLEVision Encoder=OpenAI CLIP-L/14, Evaluation Protocol=Zero-shot2025.03 | 11.7 | 12.7 | 15.1 | |
| PLIBackbone=CLIP-L, # Params=428M, Zero-shot evaluation=true2024.12 | 10.4 | — | — | |
| SEARLEBackbone=CLIP-B, # Params=165M, Zero-shot evaluation=true2024.12 | 9.4 | — | — | |
| SEARLEVision Encoder=OpenAI CLIP-B/32, Evaluation Protocol=Zero-shot2025.03 | 9.4 | 9.9 | 11.8 | |
| Slerp-TATVision Encoder=OpenAI CLIP-B/32, Evaluation Protocol=Zero-shot2025.03 | 9.3 | 10.3 | 12.3 | |
| Pic2WordBackbone=CLIP-L, # Params=429M, Zero-shot evaluation=true2024.12 | 8.7 | — | — | |
| Pic2WorldVision Encoder=OpenAI CLIP-L/14, Evaluation Protocol=Zero-shot2025.03 | 8.7 | 9.5 | 11.3 | |
| PALAVRAVision Encoder=OpenAI CLIP-B/32, Evaluation Protocol=Zero-shot2025.03 | 4.6 | 5.3 | 6.8 |