Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Multimodal AI research – Page 41 · SOTA2 Research
Back to Research
Research domain
Multimodal
Search tasks
Search
1 benchmarks · 2 papers
Visual Query 2D localization
1 benchmarks · 1 papers
Video object referring question answering
1 benchmarks · 1 papers
Referring Motion Expression Generation
1 benchmarks · 1 papers
Vision-Language Task Evaluation
1 benchmarks · 1 papers
Multi-task Image-Text Understanding
1 benchmarks · 1 papers
Video-Commentary Alignment
1 benchmarks · 1 papers
Video-and-Language Inference
1 benchmarks · 1 papers
Talking to Me (TTM)
1 benchmarks · 1 papers
Video Object Grounding
1 benchmarks · 1 papers
Visual Cognition Evaluation
1 benchmarks · 1 papers
Multi-modal MRI Synthesis (T2+T1ce+FLAIR to T1)
1 benchmarks · 1 papers
Visual Concept Composition
1 benchmarks · 1 papers
Open-Ended Medical Visual Chat
1 benchmarks · 2 papers
Multimodal Multilingual Reasoning
1 benchmarks · 1 papers
Image-Response Generation
1 benchmarks · 1 papers
Multi-modal MRI Synthesis (T2+T1+FLAIR to T1ce)
1 benchmarks · 1 papers
Text-centric Visual Question Answering (Open-Ended)
1 benchmarks · 1 papers
Multimodal Understanding Inference Efficiency
1 benchmarks · 1 papers
Multimodal Oil Painting Generation
1 benchmarks · 1 papers
Multi-image Understanding and Reasoning
1 benchmarks · 2 papers
Financial Multimodal Reasoning
1 benchmarks · 1 papers
Multi-modal MRI Synthesis (T2+T1ce+T1 to FLAIR)
1 benchmarks · 2 papers
General Multi-image Reasoning and Generalization
1 benchmarks · 1 papers
Multimodal Multilingual Evaluation
1 benchmarks · 1 papers
Multimodal Evaluation Collection
1 benchmarks · 1 papers
Aerial Visual Reasoning
1 benchmarks · 1 papers
Aerial Visual Grounding
1 benchmarks · 1 papers
Privacy-aware Image Editing (Multimodal Modality)
1 benchmarks · 1 papers
Vision-Language
1 benchmarks · 1 papers
Fine-grained Image-Text Alignment
1 benchmarks · 3 papers
General Vision-Language
1 benchmarks · 1 papers
text-to-video+audio retrieval
Page 41 of 83
Previous
Next