Video-to-Text Retrieval on MSRVTT (V2T)
52.6V2T ScoreQwen3-VL-emb-8B + DS2M
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3-VL-emb-8B + DS2MModel Type=Embedding Model, Fine-tuning Data=DS2M2026.04 | 52.6 | |
| Qwen3-VL-emb-8BModel Type=Embedding Model2026.04 | 52 | |
| Qwen3-VL-8B + DS2MModel Type=Instruct Model, Fine-tuning Data=DS2M2026.04 | 40 | |
| Qwen3-VL-8B + Ego4D + DS2MModel Type=Instruct Model, Fine-tuning Data=Ego4D + DS2M2026.04 | 37.1 | |
| Qwen3-VL-8B + Ego4DModel Type=Instruct Model, Fine-tuning Data=Ego4D2026.04 | 30.5 | |
| Qwen3-VL-8BModel Type=Instruct Model2026.04 | 30.3 |