Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Multimodal AI research – Page 6 · SOTA2 Research
Back to Research
Research domain
Multimodal
Search tasks
Search
9 benchmarks · 4 papers
Multi-Modal Image Fusion
9 benchmarks · 10 papers
Audio-to-Video Retrieval
9 benchmarks · 1 papers
Multi-modal Forgery Detection and Grounding
9 benchmarks · 13 papers
Video-Text Retrieval
9 benchmarks · 7 papers
Vision-Language-Action
9 benchmarks · 12 papers
Multimodal Document Question Answering
9 benchmarks · 10 papers
Vision-Language Compositional Reasoning
9 benchmarks · 1 papers
Cross-Task Visual In-Context Learning
9 benchmarks · 1 papers
Multimodal Safety
9 benchmarks · 7 papers
Large Vision-Language Model Evaluation
9 benchmarks · 5 papers
General multimodal reasoning
9 benchmarks · 25 papers
Visual Hallucination Evaluation
9 benchmarks · 1 papers
Multimodal Manipulation Detection
9 benchmarks · 14 papers
Multi-choice Video Question Answering
9 benchmarks · 6 papers
Video-to-Speech Synthesis
8 benchmarks · 1 papers
Image2Description
8 benchmarks · 9 papers
Video Reasoning Segmentation
8 benchmarks · 5 papers
Grounded Conversation Generation
8 benchmarks · 2 papers
Vision-to-Text
8 benchmarks · 3 papers
Multimodal Entity Linking
8 benchmarks · 14 papers
Multimodal Conversation
8 benchmarks · 5 papers
Object Grounding
8 benchmarks · 5 papers
Audio-visual generation
8 benchmarks · 5 papers
Text-Guided Image Manipulation
8 benchmarks · 1 papers
Pointwise Scoring
8 benchmarks · 6 papers
Multimodal misinformation detection
8 benchmarks · 3 papers
Prompt-image Alignment
8 benchmarks · 1 papers
Multi-modal Stance Detection
8 benchmarks · 3 papers
Multimodal Deep Search
8 benchmarks · 1 papers
Image-text similarity
8 benchmarks · 3 papers
Cross-modal registration
8 benchmarks · 9 papers
Vision-centric Reasoning
Page 6 of 83
Previous
Next