Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Multimodal AI research – Page 77 · SOTA2 Research
Back to Research
Research domain
Multimodal
Search tasks
Search
1 benchmarks · 1 papers
Music Textual Alignment
1 benchmarks · 1 papers
Video Narration
1 benchmarks · 1 papers
Multi-modal to Audio Generation Latency
1 benchmarks · 2 papers
General Robust Image Task (GRIT) multi-task evaluation
1 benchmarks · 1 papers
Video-grounded Role-playing
1 benchmarks · 1 papers
Role-Playing Evaluation (Visual-Element-Groundedness)
1 benchmarks · 1 papers
Multi-skill Compositional Reasoning
1 benchmarks · 1 papers
Multimedia Event Argument Extraction
1 benchmarks · 1 papers
Large Multimodal Model Inference Efficiency
1 benchmarks · 1 papers
Conversational VQA
1 benchmarks · 1 papers
Visual Grounding and Reasoning
1 benchmarks · 1 papers
Text+Image Editing
1 benchmarks · 1 papers
Visual Query 2D
1 benchmarks · 1 papers
Text-and-Visual-to-Image Generation (Style Transfer)
1 benchmarks · 1 papers
Text-driven Image-to-Image Translation
1 benchmarks · 1 papers
Text-to-Image with Visual condition
1 benchmarks · 1 papers
Video-Language Event Prediction
1 benchmarks · 1 papers
Visual Explanation Localization
1 benchmarks · 1 papers
Faithfulness evaluation of image explanation
1 benchmarks · 4 papers
Multimodal Movie Genre Classification
1 benchmarks · 1 papers
Active Modality Acquisition (Text imputed by Audio)
1 benchmarks · 1 papers
Active Modality Acquisition (Image imputed by Audio)
1 benchmarks · 2 papers
Multi-image Dialogue Understanding
1 benchmarks · 1 papers
Image Reconstruction from fMRI
1 benchmarks · 1 papers
Hallucination and Visual Illusion Assessment
1 benchmarks · 1 papers
Fill-in-the-blank Video Question Answering
1 benchmarks · 1 papers
Multimodal Dialogue
1 benchmarks · 1 papers
Vision-centric Visual Question Answering
1 benchmarks · 1 papers
Vague Query-based Affective Video Understanding
1 benchmarks · 1 papers
Multimodal Video Question Answering
1 benchmarks · 1 papers
Multimodal AD vs CN Classification
1 benchmarks · 1 papers
instance-aware caption generation
Page 77 of 83
Previous
Next