Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Multimodal AI research – Page 13 · SOTA2 Research
Back to Research
Research domain
Multimodal
Search tasks
Search
4 benchmarks · 1 papers
Omni-modal reasoning
4 benchmarks · 3 papers
Image-text-to-text retrieval
4 benchmarks · 4 papers
Multimodal Emotion Reasoning
4 benchmarks · 1 papers
Basic Exocentric Understanding
4 benchmarks · 3 papers
Audio-Visual Recognition
4 benchmarks · 3 papers
Multimodal Knowledge Editing
4 benchmarks · 1 papers
Product Grounding
4 benchmarks · 1 papers
Multi-sentence video grounding
4 benchmarks · 1 papers
Multimodal Humor Explanation
4 benchmarks · 1 papers
Multi-Choice Q&A
4 benchmarks · 1 papers
Visual Equivalence Reward Modeling
4 benchmarks · 1 papers
Streaming Social Task Detection
4 benchmarks · 1 papers
Referring Image Object Grounding
4 benchmarks · 3 papers
Audio-Visual Target Speaker Extraction
4 benchmarks · 3 papers
Multimodal Audio Understanding
4 benchmarks · 1 papers
element-level text-to-image alignment evaluation
4 benchmarks · 3 papers
Multiple Choice Video-QA
4 benchmarks · 1 papers
Video Comment Generation
4 benchmarks · 4 papers
Vision Question Answering
4 benchmarks · 1 papers
Vision-centric Jailbreak Attack
4 benchmarks · 2 papers
Human Motion Transfer
4 benchmarks · 6 papers
Multimodal Video Understanding
4 benchmarks · 11 papers
Multimodal Benchmark
4 benchmarks · 1 papers
Closed Visual Question Answering
4 benchmarks · 1 papers
Multimodal ECG Reasoning
4 benchmarks · 1 papers
Audio-visual source separation
4 benchmarks · 1 papers
Black-Box LVLM Attack
4 benchmarks · 1 papers
Video/Caption Retrieval
4 benchmarks · 49 papers
Multimodal Capability Evaluation
4 benchmarks · 9 papers
Audio-Visual Video Parsing
4 benchmarks · 10 papers
Visual Question Answering (Multiple-choice)
4 benchmarks · 2 papers
Audio-visual source localization
Page 13 of 83
Previous
Next