Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Multimodal AI research – Page 71 · SOTA2 Research
Back to Research
Research domain
Multimodal
Search tasks
Search
1 benchmarks · 1 papers
Perceptual Similarity Assessment
1 benchmarks · 1 papers
Multimodal Reasoning and Conversation
1 benchmarks · 1 papers
Multimodal Comment Generation
1 benchmarks · 1 papers
Note-to-Image Retrieval
1 benchmarks · 1 papers
Multimodal Multi-hop Visual Question Answering
1 benchmarks · 2 papers
Figure Visual Question Answering
1 benchmarks · 1 papers
Visual Reconstruction (B→I)
1 benchmarks · 1 papers
Audio-guided human animation
1 benchmarks · 1 papers
Audio-Visual Fact-checking
1 benchmarks · 1 papers
Referring Expression Understanding
1 benchmarks · 1 papers
Visual Metaphor Transfer
1 benchmarks · 1 papers
Vision-Language Model Inference Efficiency
1 benchmarks · 1 papers
Multi-frame visual story generation
1 benchmarks · 1 papers
Multimodal tasks
1 benchmarks · 1 papers
Multi-task multimodal understanding
1 benchmarks · 1 papers
General VLM Understanding
1 benchmarks · 1 papers
Multi-turn Multi-image Dialog
1 benchmarks · 1 papers
Audio-Visual Target-Speaker ASR
1 benchmarks · 1 papers
Medical Multi-task Visual Reasoning
1 benchmarks · 1 papers
Layout-based scene generation
1 benchmarks · 1 papers
Scene Text Image Captioning
1 benchmarks · 1 papers
Multimodal Large Language Model Inference Efficiency
1 benchmarks · 1 papers
VQAgeneral
1 benchmarks · 7 papers
Real-world Multimodal Reasoning
1 benchmarks · 1 papers
VQAspecific
1 benchmarks · 1 papers
VQA (General)
1 benchmarks · 1 papers
VQA (Specific)
1 benchmarks · 1 papers
Overall Vision-Language Performance
1 benchmarks · 1 papers
Chinese Culture Multimodal Evaluation
1 benchmarks · 2 papers
Video Detailed Captioning
1 benchmarks · 1 papers
Visual Question Answering (general)
1 benchmarks · 1 papers
Visual Question Answering (specific)
Page 71 of 83
Previous
Next