Ranking on Human-labeled query set
0.849NDCG@1ExpModel
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| ExpModel2026.03 | 0.849 | 0.854 | 0.93 | |
| GPT-4o (VL-P)Ranking Strategy=Pointwise, Modality=Vision-Language, Zero-shot=true2026.03 | 0.811 | 0.837 | 0.921 | |
| GPT-4o (T-P)Ranking Strategy=Pointwise, Modality=Text-only, Zero-shot=true2026.03 | 0.793 | 0.824 | 0.905 | |
| RankGPTRanking Strategy=Listwise, Zero-shot=true2026.03 | 0.759 | 0.801 | 0.904 | |
| GPT-4o (T-L)Ranking Strategy=Listwise, Modality=Text-only, Zero-shot=true2026.03 | 0.687 | 0.752 | 0.885 | |
| BGE-m3Retrieval Paradigm=Dual-Encoder, Mode=dense2026.03 | 0.612 | 0.721 | 0.864 |