Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Multimodal AI research – Page 12 · SOTA2 Research
Back to Research
Research domain
Multimodal
Search tasks
Search
5 benchmarks · 3 papers
Personalized Image Aesthetic Assessment
5 benchmarks · 1 papers
Dental Visual Question Answering
5 benchmarks · 5 papers
Avatar Generation
5 benchmarks · 1 papers
Humorous image captioning
5 benchmarks · 2 papers
Emotional video captioning
5 benchmarks · 2 papers
Short-answer Visual Question Answering
5 benchmarks · 2 papers
Interleaved Generation
5 benchmarks · 3 papers
Zero-shot Image-Text Retrieval
5 benchmarks · 3 papers
Audio-driven Avatar Generation
5 benchmarks · 4 papers
Text-driven Image Editing
5 benchmarks · 2 papers
Vision Reasoning
5 benchmarks · 7 papers
Interactive Video Object Segmentation
5 benchmarks · 3 papers
Multimodal Agent Task
5 benchmarks · 3 papers
Personalized Video Generation
5 benchmarks · 4 papers
Multimodal regression
4 benchmarks · 6 papers
Open-ended VQA
4 benchmarks · 3 papers
Text-Video Retrieval
4 benchmarks · 3 papers
Audio-Visual Target Speaker Extraction
4 benchmarks · 3 papers
Multimodal Audio Understanding
4 benchmarks · 1 papers
Closed Visual Question Answering
4 benchmarks · 1 papers
Vision-centric Jailbreak Attack
4 benchmarks · 1 papers
Lip to Speech
4 benchmarks · 2 papers
Fine-Grained Sketch-Based Image Retrieval
4 benchmarks · 6 papers
Multimodal Video Understanding
4 benchmarks · 1 papers
Video Comment Generation
4 benchmarks · 3 papers
Multiple Choice Video-QA
4 benchmarks · 4 papers
Multilingual Reading Comprehension
4 benchmarks · 11 papers
Multimodal Benchmark
4 benchmarks · 3 papers
Audio-Visual Recognition
4 benchmarks · 1 papers
Multi-modal Visualization Generation
4 benchmarks · 1 papers
Product Grounding
4 benchmarks · 1 papers
Multi-sentence video grounding
Page 12 of 83
Previous
Next