Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Multimodal AI research – Page 52 · SOTA2 Research
Back to Research
Research domain
Multimodal
Search tasks
Search
1 benchmarks · 1 papers
Visual Tool Reasoning
1 benchmarks · 1 papers
Scientific Multimodal Document Reasoning
1 benchmarks · 1 papers
RGB-T Semantic Segmentation
1 benchmarks · 1 papers
Audio-visual speech enhancement
1 benchmarks · 1 papers
Cross-Domain Few-Shot Action Recognition
1 benchmarks · 1 papers
Visual Question Answering, Information Extraction, and Element Localization
1 benchmarks · 1 papers
Audiovisual Deepfake Detection
1 benchmarks · 1 papers
Text-Guided Object Alignment
1 benchmarks · 1 papers
Egocentric Visual Question Answering
1 benchmarks · 1 papers
Place Categorization
1 benchmarks · 1 papers
Augmented Knowledge-based Visual Question Answering
1 benchmarks · 1 papers
VLM Grounding
1 benchmarks · 1 papers
Cross-identity face animation
1 benchmarks · 1 papers
Image Reconstruction Similarity
1 benchmarks · 1 papers
CUA grounding
1 benchmarks · 1 papers
Image Forgery Explanation Quality
1 benchmarks · 1 papers
Video-to-audio captioning
1 benchmarks · 1 papers
Descriptive Visual Question Answering
1 benchmarks · 1 papers
Multimodal Reasoning Evaluation
1 benchmarks · 1 papers
Multimodal AutoML
1 benchmarks · 1 papers
Cross-Modal Scene Retrieval (Image to Description)
1 benchmarks · 1 papers
Cross-Modal Scene Retrieval (Image to Text)
1 benchmarks · 1 papers
Tennis Commentary Generation
1 benchmarks · 1 papers
Audio-visual video forgery detection
1 benchmarks · 1 papers
Cross-Modal Scene Retrieval (Point Cloud to Text)
1 benchmarks · 1 papers
VLM High-level Planning
1 benchmarks · 1 papers
Cross-Modal Scene Retrieval (Image to Floorplan)
1 benchmarks · 1 papers
Multi-Modal Domain Generalization
1 benchmarks · 1 papers
Open-vocabulary language grounding
1 benchmarks · 1 papers
Cross-Modal Scene Retrieval (Text to Floorplan)
1 benchmarks · 1 papers
Expressive Dubbing
1 benchmarks · 1 papers
Cross-Modal Coarse Visual Localization
Page 52 of 83
Previous
Next