Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Multimodal AI research – Page 46 · SOTA2 Research
Back to Research
Research domain
Multimodal
Search tasks
Search
1 benchmarks · 1 papers
Multiview text-to-image inpainting
1 benchmarks · 1 papers
Image Query Sound Separation
1 benchmarks · 1 papers
Grounded Visual Reasoning
1 benchmarks · 1 papers
Image-text fact-checking
1 benchmarks · 2 papers
Multi-modal dialogue retrieval
1 benchmarks · 1 papers
General Multi-modal Assistant Task
1 benchmarks · 1 papers
Zero-shot Cross-modal Robustness
1 benchmarks · 1 papers
Spatial Audio Question Answering (Open-Ended)
1 benchmarks · 1 papers
Spatial Audio Question Answering
1 benchmarks · 1 papers
Multimodal Fact-checking
1 benchmarks · 1 papers
Visual Multiple Choice Question Answering
1 benchmarks · 1 papers
Multi-modal Intent Prediction
1 benchmarks · 2 papers
Multimodal Dialogue Generation
1 benchmarks · 1 papers
Motion Amplitude Assessment
1 benchmarks · 1 papers
Medical Multimodal Reasoning and Understanding
1 benchmarks · 1 papers
Audio localization from visual segment query
1 benchmarks · 5 papers
Visual Question Answering (Multi-choice)
1 benchmarks · 1 papers
Temporal Commonsense Reasoning
1 benchmarks · 1 papers
Multi-modal Continual Learning
1 benchmarks · 1 papers
3D Medical Visual Question Answering (Overall)
1 benchmarks · 1 papers
Vision-Language Interaction Evaluation
1 benchmarks · 1 papers
Image Classification and Multimodal Retrieval
1 benchmarks · 1 papers
Video-audio synchrony classification
1 benchmarks · 1 papers
Cross-modal Image Classification
1 benchmarks · 1 papers
Free-format Multimodal Hallucination Assessment
1 benchmarks · 1 papers
High-quality Vision-Language Evaluation
1 benchmarks · 1 papers
Image-centric Multimodal Understanding
1 benchmarks · 2 papers
Video Referring Expression Segmentation
1 benchmarks · 1 papers
Idiom-based vision-language matching (Subtask A)
1 benchmarks · 1 papers
Cross-lingual Vision-Language Understanding and Retrieval
1 benchmarks · 1 papers
Idiom-based vision-language matching (Subtask B)
1 benchmarks · 1 papers
Multimodal Reasoning Efficiency
Page 46 of 83
Previous
Next