Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Multimodal AI research – Page 27 · SOTA2 Research
Back to Research
Research domain
Multimodal
Search tasks
Search
2 benchmarks · 4 papers
text-conditioned human interaction generation
2 benchmarks · 2 papers
Fine-grained Grounding
2 benchmarks · 1 papers
Multi-modal preference alignment
2 benchmarks · 2 papers
Music Conditioned Dance Generation
2 benchmarks · 1 papers
Compositional Visual Reasoning
2 benchmarks · 1 papers
Short-caption Text-to-Image Retrieval
2 benchmarks · 1 papers
Driving Video Multi-Choice Question Answering
2 benchmarks · 2 papers
Image-based Question Answering
2 benchmarks · 2 papers
Egocentric Question Answering
2 benchmarks · 1 papers
Vision-Language Captioning
2 benchmarks · 2 papers
Video QA
2 benchmarks · 2 papers
Video-to-Text
2 benchmarks · 1 papers
Image-to-Sound Retrieval
2 benchmarks · 1 papers
Long-Horizon Vision-Language Navigation
2 benchmarks · 1 papers
General Multimodal Question Answering
2 benchmarks · 1 papers
Sound-to-Image Retrieval
2 benchmarks · 2 papers
Long-context Multimodal Understanding
2 benchmarks · 1 papers
Multimodal Reasoning and Mathematics
2 benchmarks · 1 papers
Video Scene Graph Classification (SGCLS)
2 benchmarks · 1 papers
Closed-ended Visual Question Answering
2 benchmarks · 1 papers
Visual Instruction Generation
2 benchmarks · 1 papers
Audio-driven Digital Human Generation
2 benchmarks · 1 papers
Video Dialog
2 benchmarks · 1 papers
QA performance by Gemini-2.5-Pro based on captions
2 benchmarks · 1 papers
Sketch-to-Real Person Re-identification
2 benchmarks · 1 papers
Fine-grained attribute binding
2 benchmarks · 1 papers
Cross-Object Reenactment
2 benchmarks · 2 papers
Multimodal Video Retrieval
2 benchmarks · 1 papers
Medical Multi-discipline Multimodal Understanding
2 benchmarks · 1 papers
Modality Preference Steering
2 benchmarks · 1 papers
Counterfactual audio-image recognition
2 benchmarks · 1 papers
Knowledge-Based Visual Question Answering (Direct Answer)
Page 27 of 83
Previous
Next