Sequential Multimodal Retrieval on OBELICS Hybrid6 1.0 (1K document sample)
22.7Seq-IVisInContext
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| VisInContextVisual Input=Raw Image, Text Input=Raw Text, Surrounding Text Input=Rendered Text Image2024.06 | 22.7 | 66.5 | |
| VisInContextVisual Input=Raw Image, Text Input=Raw Text, Surrounding Text Input=Raw Text2024.06 | 18.9 | 67.5 | |
| VisInContextVisual Input=Raw Image, Text Input=Raw Text, Surrounding Text Input=None2024.06 | 16.3 | 64.8 |