Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Multimodal AI research – Page 75 · SOTA2 Research
Back to Research
Research domain
Multimodal
Search tasks
Search
1 benchmarks · 1 papers
Sign language reconstruction
1 benchmarks · 1 papers
Explanatory Visual Question Answering
1 benchmarks · 1 papers
Natural Language Feedback Generation
1 benchmarks · 1 papers
Vision-Language (Image Description)
1 benchmarks · 1 papers
Multimodal Conversational Question Answering
1 benchmarks · 1 papers
MRI-to-Audio Synthesis
1 benchmarks · 1 papers
Language-guided Image Editing
1 benchmarks · 2 papers
Vision-Language Understanding and Reasoning
1 benchmarks · 1 papers
Multimodal Chart Extraction
1 benchmarks · 1 papers
Cross-modal Retrieval (ta → vi → te)
1 benchmarks · 1 papers
Adversarial Attack Transfer (Target: Qwen2-VL)
1 benchmarks · 2 papers
Multi-disciplinary Visual Question Answering
1 benchmarks · 1 papers
Multi-task aerial image reasoning
1 benchmarks · 1 papers
Aerial image reasoning
1 benchmarks · 1 papers
Foundation Model Training
1 benchmarks · 1 papers
Pure text-guided image editing
1 benchmarks · 1 papers
Cross-Modal Conflict Resolution and Scene Consistency
1 benchmarks · 2 papers
Driving Visual Question Answering
1 benchmarks · 1 papers
Vision-Language-Action Execution
1 benchmarks · 1 papers
Vision-Language Multi-task Performance
1 benchmarks · 1 papers
Music Textual Alignment
1 benchmarks · 2 papers
General Robust Image Task (GRIT) multi-task evaluation
1 benchmarks · 1 papers
Video-grounded Role-playing
1 benchmarks · 1 papers
Role-Playing Evaluation (Visual-Element-Groundedness)
1 benchmarks · 1 papers
Large Multimodal Model Inference Efficiency
1 benchmarks · 1 papers
Conversational VQA
1 benchmarks · 1 papers
Visual Grounding and Reasoning
1 benchmarks · 1 papers
Text+Image Editing
1 benchmarks · 1 papers
Text-and-Visual-to-Image Generation (Style Transfer)
1 benchmarks · 1 papers
Text-driven Image-to-Image Translation
1 benchmarks · 1 papers
Text-to-Image with Visual condition
1 benchmarks · 1 papers
Video-Language Event Prediction
Page 75 of 83
Previous
Next