Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Multimodal AI research – Page 44 · SOTA2 Research
Back to Research
Research domain
Multimodal
Search tasks
Search
1 benchmarks · 1 papers
Visual Reasoning Robustness
1 benchmarks · 1 papers
Continual Multimodal Instruction Tuning
1 benchmarks · 1 papers
Multimodal app-use reasoning
1 benchmarks · 1 papers
Generalized Zero-Shot Retrieval (Text-to-Video)
1 benchmarks · 1 papers
single-shot video captioning
1 benchmarks · 3 papers
Multi-image visual perception
1 benchmarks · 1 papers
Aggregated Performance Benchmarking
1 benchmarks · 1 papers
Aggregate performance across 10 tasks
1 benchmarks · 2 papers
Visuo-tactile manipulation
1 benchmarks · 7 papers
Compositional Vision-Language Reasoning
1 benchmarks · 1 papers
Image Detail Description
1 benchmarks · 1 papers
Interactive Visual Grounding
1 benchmarks · 1 papers
Internet Content Understanding
1 benchmarks · 1 papers
Facial Expression Imitation
1 benchmarks · 1 papers
Generalized Zero-Shot Retrieval (Text-to-Audio-Video)
1 benchmarks · 3 papers
Music-to-Dance Synthesis
1 benchmarks · 1 papers
Multi-image in the Wild
1 benchmarks · 1 papers
VLA Inference
1 benchmarks · 1 papers
Multimodal Multi-label Classification
1 benchmarks · 1 papers
Audio+caption guided speech generation
1 benchmarks · 1 papers
Audio Caption & Generation
1 benchmarks · 1 papers
Egocentric Video Question Answering
1 benchmarks · 1 papers
Music Caption & Generation
1 benchmarks · 2 papers
Visual Reasoning and Instruction Following
1 benchmarks · 1 papers
Video-and-Text-to-Audio Generation
1 benchmarks · 1 papers
Video-Text-to-Audio (VT2A)
1 benchmarks · 1 papers
Unsupervised Multimodal Machine Translation
1 benchmarks · 1 papers
Foundation Feature Reconstruction
1 benchmarks · 2 papers
Multimodal Reasoning and Tool-use
1 benchmarks · 1 papers
Audio-Visual Audio Remixing
1 benchmarks · 1 papers
Audio-Visual Sound Separation
1 benchmarks · 1 papers
Face-voice cross-modal verification
Page 44 of 83
Previous
Next