Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Multimodal AI research – Page 34 · SOTA2 Research
Back to Research
Research domain
Multimodal
Search tasks
Search
2 benchmarks · 1 papers
Region-level Visual Grounding
2 benchmarks · 2 papers
Omnimodal Understanding
2 benchmarks · 1 papers
Medical Multi-discipline Multimodal Understanding
2 benchmarks · 2 papers
VQA Hallucination
2 benchmarks · 1 papers
Closed-ended Visual Question Answering
2 benchmarks · 1 papers
QA performance by Gemini-2.5-Pro based on captions
2 benchmarks · 1 papers
Object sketch-based scene retrieval
2 benchmarks · 1 papers
7-class emotion recognition
2 benchmarks · 3 papers
Multi-task vision and language evaluation
2 benchmarks · 1 papers
Caption Matching and Retrieval
2 benchmarks · 1 papers
Audio-driven Digital Human Generation
2 benchmarks · 1 papers
Modality Preference Steering
2 benchmarks · 1 papers
Universal multi-modal reasoning
2 benchmarks · 2 papers
Visual Spatial Intelligence Reasoning
2 benchmarks · 2 papers
Medical Image-Text Classification
2 benchmarks · 1 papers
AV-ASR
2 benchmarks · 2 papers
General Multi-modal Understanding
2 benchmarks · 2 papers
Iterative Vision-and-Language Navigation
2 benchmarks · 3 papers
Visual Trace Generation
2 benchmarks · 1 papers
Multimodal Event Classification
2 benchmarks · 2 papers
EEG-to-Image Retrieval
2 benchmarks · 1 papers
Multi-modal Reconstruction
2 benchmarks · 2 papers
Multi-modal Machine Translation
2 benchmarks · 1 papers
target object grounding
2 benchmarks · 1 papers
Multimodal Coding
2 benchmarks · 1 papers
Text-Rich VQA
2 benchmarks · 1 papers
Multimodal Rumor Detection
2 benchmarks · 2 papers
Radiology VQA
2 benchmarks · 1 papers
Emergent modality binding (au -> te -> vi)
2 benchmarks · 1 papers
Emergent modality binding (vi -> te -> au)
2 benchmarks · 2 papers
Online Video Question Answering
2 benchmarks · 1 papers
Audio-Video Alignment
Page 34 of 83
Previous
Next