Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Multimodal AI research – Page 25 · SOTA2 Research
Back to Research
Research domain
Multimodal
Search tasks
Search
2 benchmarks · 1 papers
Short-caption Text-to-Image Retrieval
2 benchmarks · 1 papers
Video-to-Paragraph Retrieval
2 benchmarks · 1 papers
Long-caption Text-to-Image Retrieval
2 benchmarks · 3 papers
General Task
2 benchmarks · 1 papers
Grounded Text-to-Image Generation
2 benchmarks · 1 papers
Driving Video Multi-Choice Question Answering
2 benchmarks · 2 papers
Fine-grained Grounding
2 benchmarks · 2 papers
Multimodal Vision-Language Evaluation
2 benchmarks · 1 papers
Long-caption Image-to-Text Retrieval
2 benchmarks · 1 papers
Crossmodal retrieval
2 benchmarks · 1 papers
Multimodal Backdoor Attack
2 benchmarks · 2 papers
General Downstream Evaluation
2 benchmarks · 2 papers
Long-context Multimodal Understanding
2 benchmarks · 2 papers
Visual Relation Tagging
2 benchmarks · 4 papers
Multi-discipline Understanding
2 benchmarks · 1 papers
Multimodal Reasoning and Mathematics
2 benchmarks · 1 papers
General Multimodal Question Answering
2 benchmarks · 1 papers
Vision-Language Captioning
2 benchmarks · 2 papers
Multimodal Video Retrieval
2 benchmarks · 1 papers
Audio-driven Digital Human Generation
2 benchmarks · 1 papers
Continual audio-visual sound separation
2 benchmarks · 1 papers
Sketch-to-Real Person Re-identification
2 benchmarks · 1 papers
QA performance by Gemini-2.5-Pro based on captions
2 benchmarks · 1 papers
Context assisted image captioning
2 benchmarks · 1 papers
Video-grounded dialogue generation
2 benchmarks · 1 papers
Multi-modal Video Fusion
2 benchmarks · 1 papers
Multilingual Image Captioning
2 benchmarks · 1 papers
Closed-ended Visual Question Answering
2 benchmarks · 1 papers
Medical Multi-discipline Multimodal Understanding
2 benchmarks · 1 papers
Outdoor Vision-and-Language Navigation
2 benchmarks · 1 papers
Cross-lingual Image Captioning
2 benchmarks · 1 papers
Universal multi-modal reasoning
Page 25 of 83
Previous
Next