Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Multimodal AI research – Page 55 · SOTA2 Research
Back to Research
Research domain
Multimodal
Search tasks
Search
1 benchmarks · 5 papers
Embodied Visual Question Answering
1 benchmarks · 1 papers
Text-to-World
1 benchmarks · 1 papers
Multi-modal Visual Question Answering
1 benchmarks · 1 papers
Three-way gesture classification
1 benchmarks · 1 papers
Cross-modal and cross-lingual retrieval
1 benchmarks · 1 papers
Text-to-Binaural Audio Generation
1 benchmarks · 3 papers
Movie Fill-in-the-Blank
1 benchmarks · 2 papers
Transition Video Question Answering
1 benchmarks · 1 papers
Bimodal regression
1 benchmarks · 1 papers
Audio-based Referring Video Object Segmentation
1 benchmarks · 1 papers
General MLLM Evaluation
1 benchmarks · 1 papers
Multimodal Conversation Generation
1 benchmarks · 1 papers
Referring and Reasoning Video Object Segmentation
1 benchmarks · 1 papers
Vision-Language Conversation and Reasoning
1 benchmarks · 1 papers
Offline Promptable Video Segmentation
1 benchmarks · 1 papers
Multimodal Retrieval (text query to multimodal candidate)
1 benchmarks · 1 papers
Cross-modal Retrieval (Audio to Text)
1 benchmarks · 1 papers
Region-based text-to-image generation
1 benchmarks · 1 papers
Cross-modal Retrieval (Text to Audio)
1 benchmarks · 1 papers
Cross-modal Retrieval (Audio to Image)
1 benchmarks · 1 papers
Cross-modal Retrieval (Image to Audio)
1 benchmarks · 1 papers
Cross-modal Retrieval (Image to Text)
1 benchmarks · 1 papers
Contrastive Alignment
1 benchmarks · 1 papers
Cross-modal Retrieval (Text to Image)
1 benchmarks · 1 papers
Clothing Style Editing
1 benchmarks · 1 papers
Fill-in-the-blank Visual Question Answering
1 benchmarks · 1 papers
Icon Grounding
1 benchmarks · 1 papers
Spatial-Temporal Video Grounding
1 benchmarks · 1 papers
Multimodal Healthcare Agent Performance
1 benchmarks · 1 papers
Visual Generative Model Evaluation
1 benchmarks · 1 papers
Pseudo-captioning
1 benchmarks · 1 papers
Image-to-sound generation
Page 55 of 83
Previous
Next