Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Multimodal AI research – Page 79 · SOTA2 Research
Back to Research
Research domain
Multimodal
Search tasks
Search
1 benchmarks · 1 papers
Audio-Visual Sound Separation
1 benchmarks · 1 papers
Multimodal Question Answering and Reasoning
1 benchmarks · 1 papers
Image Reconstruction from fMRI
1 benchmarks · 1 papers
Hallucination and Visual Illusion Assessment
1 benchmarks · 1 papers
Pathology Multimodal Understanding
1 benchmarks · 1 papers
Multimodal Brain Image Registration
1 benchmarks · 1 papers
Fill-in-the-blank Video Question Answering
1 benchmarks · 1 papers
Multimodal Dialogue
1 benchmarks · 1 papers
Vision-centric Visual Question Answering
1 benchmarks · 1 papers
Vague Query-based Affective Video Understanding
1 benchmarks · 1 papers
Multimodal voice phishing and synthetic voice detection
1 benchmarks · 1 papers
Multimodal Video Question Answering
1 benchmarks · 1 papers
Multimodal multi-span medical question answering
1 benchmarks · 1 papers
Multimodal AD vs CN Classification
1 benchmarks · 1 papers
instance-aware caption generation
1 benchmarks · 1 papers
Multimodal classification (AD vs CN)
1 benchmarks · 1 papers
Visual Question Answering (Perception)
1 benchmarks · 1 papers
Multi-stage Vision-and-Language Navigation
1 benchmarks · 1 papers
Image & Text to Text (IT2T)
1 benchmarks · 1 papers
Amodal Grounding
1 benchmarks · 1 papers
Multimodal Safety Auditing
1 benchmarks · 1 papers
Image-to-Text Cross-lingual Retrieval
1 benchmarks · 1 papers
Multi-entity Grounding
1 benchmarks · 1 papers
Downstream Generation Impact (LLaVA-1.5)
1 benchmarks · 1 papers
Video Generation Alignment
1 benchmarks · 1 papers
Downstream Generation Impact (LLaVA-1.6)
1 benchmarks · 1 papers
Downstream Generation Impact (InstructBLIP)
1 benchmarks · 1 papers
Commonsense knowledge probing
1 benchmarks · 1 papers
Masked multi-modal modeling
1 benchmarks · 1 papers
Multimodal Assessment
1 benchmarks · 1 papers
Downstream Generation Impact (Qwen2.5-VL)
1 benchmarks · 1 papers
Vision-Language Deepfake Detection
Page 79 of 83
Previous
Next