Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Multimodal AI research – Page 32 · SOTA2 Research
Back to Research
Research domain
Multimodal
Search tasks
Search
2 benchmarks · 2 papers
Accident Reason Answering
2 benchmarks · 1 papers
Multimodal Temporal Point Process prediction
2 benchmarks · 1 papers
Multimodal Sarcasm Explanation
2 benchmarks · 2 papers
Long-context Multimodal Understanding
2 benchmarks · 3 papers
Multimodal Translation
2 benchmarks · 1 papers
Segment-level Video-to-Music Retrieval
2 benchmarks · 4 papers
Audiovisual Video Captioning
2 benchmarks · 2 papers
Video-based Facial Expression Recognition
2 benchmarks · 2 papers
Image Captioning Hallucination Detection
2 benchmarks · 8 papers
Visual Spatial Intelligence
2 benchmarks · 1 papers
Vision-Language Captioning
2 benchmarks · 1 papers
Segment-level Music-to-Video Retrieval
2 benchmarks · 4 papers
Multi-modal Instruction Following
2 benchmarks · 1 papers
Multimodal Dialogue Response Generation
2 benchmarks · 1 papers
Multimodal Autocomplete
2 benchmarks · 1 papers
Bi-directional Retrieval
2 benchmarks · 4 papers
Multimodal Retrieval and Understanding
2 benchmarks · 2 papers
Multimodal Video Retrieval
2 benchmarks · 2 papers
Multimodal Chart Reasoning
2 benchmarks · 1 papers
Closed-ended Visual Question Answering
2 benchmarks · 1 papers
Audio-driven Digital Human Generation
2 benchmarks · 1 papers
Sketch-to-Real Person Re-identification
2 benchmarks · 1 papers
Gaze following in video
2 benchmarks · 1 papers
QA performance by Gemini-2.5-Pro based on captions
2 benchmarks · 1 papers
Vision-Language Reward Model Evaluation
2 benchmarks · 1 papers
Audio-Visual Speech-to-Speech Translation
2 benchmarks · 1 papers
7-class emotion recognition
2 benchmarks · 1 papers
Audio-Visual Speech-to-Audio Translation
2 benchmarks · 1 papers
Modality Preference Steering
2 benchmarks · 1 papers
Universal multi-modal reasoning
2 benchmarks · 1 papers
Multimodal Safety Defense
2 benchmarks · 2 papers
General Multi-modal Understanding
Page 32 of 83
Previous
Next