Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Multimodal AI research – Page 10 · SOTA2 Research
Back to Research
Research domain
Multimodal
Search tasks
Search
6 benchmarks · 2 papers
Missing Modality Imputation
6 benchmarks · 1 papers
3D/4D Video Question Answering
6 benchmarks · 1 papers
Visual World Modelling
6 benchmarks · 4 papers
Visual Explanation Generation
5 benchmarks · 18 papers
Visual Perception and Reasoning
5 benchmarks · 2 papers
Action Grounding
5 benchmarks · 2 papers
Emotional video captioning
5 benchmarks · 10 papers
Multimodal Emotion Recognition in Conversation
5 benchmarks · 3 papers
Zero-shot Video Object Segmentation
5 benchmarks · 1 papers
Multimodal Survival Prediction
5 benchmarks · 1 papers
Audio-Visual Automatic Speech Recognition
5 benchmarks · 3 papers
Multimodal Agent Task
5 benchmarks · 5 papers
Video-Language Understanding
5 benchmarks · 1 papers
Unified Multimodal Parsing
5 benchmarks · 5 papers
Long Video Reasoning
5 benchmarks · 7 papers
Video spatial reasoning
5 benchmarks · 5 papers
Multi-image Spatial Reasoning
5 benchmarks · 2 papers
Short-answer Visual Question Answering
5 benchmarks · 2 papers
Interleaved Generation
5 benchmarks · 6 papers
Recipe-to-image retrieval
5 benchmarks · 3 papers
Audio-driven Avatar Generation
5 benchmarks · 5 papers
Avatar Generation
5 benchmarks · 1 papers
Humorous image captioning
5 benchmarks · 4 papers
Multimodal Embedding Evaluation
5 benchmarks · 3 papers
Vision-Language Compositionality
5 benchmarks · 3 papers
Visual Narrative Generation
5 benchmarks · 2 papers
Multi-modal Manipulation Detection and Grounding
5 benchmarks · 1 papers
Cross-modal binding
5 benchmarks · 3 papers
Zero-shot Image-Text Retrieval
5 benchmarks · 1 papers
Multimodal Document QA
5 benchmarks · 7 papers
Interleaved Image-Text Generation
5 benchmarks · 87 papers
Multi-discipline Multimodal Understanding
Page 10 of 83
Previous
Next