Video Segment Retrieval on CourseTimeQA (LOOCV micro-average)
53R@1CrossFusion-RAG
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| CrossFusion-RAGLearning protocol=learned2025.11 | 53 | 74 | 88 | 64 | 84 | |
| Late-fusion gatingLearning protocol=learned2025.11 | 51 | 70 | 85 | 62 | 79 | |
| CLIP + cross-encoder reranker + MMRLearning protocol=learned, Reranking and diversification=cross-encoder, MMR2025.11 | 50 | 70 | 85 | 61 | 80 | |
| Caption-aug text retrievalLearning protocol=learned2025.11 | 49 | 68 | 83 | 60 | 76 | |
| Text-hybrid + MMRLearning protocol=learned, Reranking and diversification=no reranker, MMR2025.11 | 48 | 67 | 81 | 57 | 75 | |
| CLIP poolingLearning protocol=zero-shot, Frames (N)=42025.11 | 48 | 67 | 82 | 59 | 76 | |
| Text-hybrid + MiniLMLearning protocol=learned2025.11 | 47 | 66 | 80 | 58 | 73 | |
| Na¨ıve fusion + smoothingLearning protocol=non-learned, Temporal and fusion strategy=smoothing2025.11 | 45 | 63 | 79 | 55 | 74 |