Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Multimodal AI research – Page 51 · SOTA2 Research
Back to Research
Research domain
Multimodal
Search tasks
Search
1 benchmarks · 1 papers
Synthetic Image-to-Recipe Retrieval
1 benchmarks · 1 papers
Multimodal MRI Segmentation
1 benchmarks · 1 papers
HOI Interpolation
1 benchmarks · 1 papers
Visual Discrepancy Detection
1 benchmarks · 1 papers
Satellite-to-Street Localization
1 benchmarks · 1 papers
Future Video Question Answering
1 benchmarks · 2 papers
Multimodal Compositional Understanding
1 benchmarks · 1 papers
Multimodal Interleaved Generation
1 benchmarks · 2 papers
Cross-Modal Rank-1 Identification
1 benchmarks · 1 papers
Recipe-to-Synthetic Image Retrieval
1 benchmarks · 1 papers
Audio-Visual Action Recognition
1 benchmarks · 1 papers
Multimodal Audio Reasoning
1 benchmarks · 1 papers
Panoramic Visual Question Answering
1 benchmarks · 13 papers
Comprehensive Multi-modal Evaluation
1 benchmarks · 1 papers
Joint Tri-Modal Medical Image Fusion and Super-Resolution
1 benchmarks · 1 papers
Model Merging Performance Aggregation
1 benchmarks · 1 papers
High-level Multimodal Reasoning
1 benchmarks · 1 papers
Spatiotemporal Video Question Answering
1 benchmarks · 1 papers
Multimodal Federated Learning
1 benchmarks · 2 papers
Video/Paragraph Retrieval (Video-to-Text)
1 benchmarks · 1 papers
Multimedia Retrieval
1 benchmarks · 1 papers
Conversational Image Retrieval (tChatSearch)
1 benchmarks · 1 papers
Agentic Multimodal Tool-use
1 benchmarks · 1 papers
Video Camouflaged Object Segmentation
1 benchmarks · 1 papers
Visual Realism Assessment
1 benchmarks · 1 papers
Video-guided Machine Translation
1 benchmarks · 1 papers
VLA Training Throughput (pi_0)
1 benchmarks · 1 papers
Multimodal Task Orchestration and Question Answering
1 benchmarks · 1 papers
Modality-specific Medical Visual Question Answering
1 benchmarks · 1 papers
Heterogeneous Medical VQA and Disease Classification
1 benchmarks · 1 papers
Vision-language robotic manipulation
1 benchmarks · 1 papers
Dynamic State Grounding
Page 51 of 83
Previous
Next