Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Multimodal AI research – Page 61 · SOTA2 Research
Back to Research
Research domain
Multimodal
Search tasks
Search
1 benchmarks · 1 papers
Video Physical Reasoning
1 benchmarks · 1 papers
Differences Caption Generation
1 benchmarks · 1 papers
Vision Language Model Inference
1 benchmarks · 1 papers
Visual Decoding
1 benchmarks · 1 papers
Segmented Multimodal Named Entity Recognition
1 benchmarks · 1 papers
Video-to-Music Retrieval
1 benchmarks · 1 papers
Video-motion retrieval
1 benchmarks · 2 papers
Visual Encoding
1 benchmarks · 1 papers
Visual Commonsense
1 benchmarks · 1 papers
Multimodal Depression Detection
1 benchmarks · 1 papers
Spatial and Temporal Reasoning
1 benchmarks · 1 papers
Referring Single-Object Tracking
1 benchmarks · 1 papers
Music-to-Video Retrieval
1 benchmarks · 1 papers
Visual Brain Decoding
1 benchmarks · 1 papers
Video-to-Music
1 benchmarks · 1 papers
Multimodal Hallucination Detection
1 benchmarks · 1 papers
Iconic Gesture Intensity Regression
1 benchmarks · 1 papers
Iconic Gesture Placement
1 benchmarks · 1 papers
Multimodal Ensemble Retrieval
1 benchmarks · 1 papers
Paper-to-slides generation
1 benchmarks · 1 papers
Text+Image-to-Vector Animation Generation
1 benchmarks · 1 papers
Multilingual Vision-Language Reasoning
1 benchmarks · 1 papers
Multi-modal Medical Image Synthesis
1 benchmarks · 1 papers
Multilingual Multimodal Reasoning
1 benchmarks · 1 papers
Hierarchy Question Answering
1 benchmarks · 1 papers
VLM Pairwise Preference
1 benchmarks · 1 papers
Nonverbal Visual Reasoning
1 benchmarks · 1 papers
Slide Deck Generation Quality Evaluation
1 benchmarks · 1 papers
Open Visual Question Answering
1 benchmarks · 1 papers
Multimodal Action Labeling
1 benchmarks · 1 papers
Multimodal Image-to-Image Translation (Shade+Texture to RGB)
1 benchmarks · 1 papers
Video-to-Text temporal grounding
Page 61 of 83
Previous
Next