Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Multimodal AI research – Page 5 · SOTA2 Research
Back to Research
Research domain
Multimodal
Search tasks
Search
11 benchmarks · 1 papers
Active Modality Acquisition
11 benchmarks · 5 papers
Visual Instruction Tuning
11 benchmarks · 12 papers
General Multimodal Understanding
11 benchmarks · 1 papers
Visual Language Model IP Protection
11 benchmarks · 1 papers
Open-vocabulary video instance segmentation
11 benchmarks · 4 papers
Cross-modal Geo-localization
11 benchmarks · 10 papers
Text-guided Video Editing
11 benchmarks · 15 papers
Multiple-choice Video Question Answering
11 benchmarks · 2 papers
Referring Localization
11 benchmarks · 8 papers
Text-Image Retrieval
11 benchmarks · 9 papers
Multimodal Understanding and Reasoning
11 benchmarks · 12 papers
Multimodal Named Entity Recognition
11 benchmarks · 8 papers
Remote Sensing Visual Question Answering
10 benchmarks · 4 papers
Referring Audio-Visual Segmentation
10 benchmarks · 3 papers
Multi-view Multi-label Classification
10 benchmarks · 5 papers
Image Reasoning
10 benchmarks · 1 papers
Multimodal Recommendation Unlearning
10 benchmarks · 4 papers
Element Grounding
10 benchmarks · 8 papers
Semantic Consistency
10 benchmarks · 2 papers
Multimodal Forecasting
10 benchmarks · 14 papers
Composed Video Retrieval
10 benchmarks · 23 papers
Visual Dialog
10 benchmarks · 7 papers
Image-text alignment
10 benchmarks · 2 papers
Multimodal Preference Evaluation
10 benchmarks · 1 papers
Multimodal Autoformalization
10 benchmarks · 18 papers
Audio-Visual Classification
9 benchmarks · 7 papers
Large Vision-Language Model Evaluation
9 benchmarks · 1 papers
Grounding segmentation
9 benchmarks · 28 papers
Referring Video Segmentation
9 benchmarks · 10 papers
Audio-to-Video Retrieval
9 benchmarks · 6 papers
Video-to-Speech Synthesis
9 benchmarks · 5 papers
General multimodal reasoning
Page 5 of 83
Previous
Next