Text-to-Image Retrieval on internal-8M
68AccuracyOurs-FT
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Ours-FTComparison Baseline=Ours-PT + Rerank, Evaluation Judge=GPT-4V, LLM Rephrasing=false, System A Stages=1, System B Stages=22024.06 | 68 | 70 | |
| Ours-FTComparison Baseline=Getty search, Evaluation Judge=Human labeler, LLM Rephrasing=true, System A Stages=1, System B Stages=>22024.06 | 63.3 | 61.3 | |
| Ours-FTComparison Baseline=Getty search, Evaluation Judge=GPT-4V, LLM Rephrasing=true, System A Stages=1, System B Stages=>22024.06 | 62.7 | 62.7 | |
| Ours-FTComparison Baseline=Bing search, Evaluation Judge=GPT-4V, LLM Rephrasing=true, System A Stages=1, System B Stages=>22024.06 | 45.3 | 57.3 |