Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Multimodal AI research – Page 81 · SOTA2 Research
Back to Research
Research domain
Multimodal
Search tasks
Search
1 benchmarks · 1 papers
Downstream Generation Impact (LLaVA-1.6)
1 benchmarks · 1 papers
Vision-Language In-Context Learning
1 benchmarks · 1 papers
Downstream Generation Impact (InstructBLIP)
1 benchmarks · 1 papers
Commonsense knowledge probing
1 benchmarks · 1 papers
Masked multi-modal modeling
1 benchmarks · 1 papers
Multimodal Assessment
1 benchmarks · 1 papers
Downstream Generation Impact (Qwen2.5-VL)
1 benchmarks · 1 papers
Online Editing
1 benchmarks · 1 papers
Compositional Image Reasoning
1 benchmarks · 1 papers
Vision-Language Deepfake Detection
1 benchmarks · 3 papers
Image Pointing
1 benchmarks · 1 papers
Multimodal Role-Play (T2T2I)
1 benchmarks · 1 papers
Multimodal Robustness
1 benchmarks · 1 papers
Video Scene Description
1 benchmarks · 1 papers
Active Panoramic Referring Expression Segmentation
1 benchmarks · 1 papers
Multimodal Video Inference
1 benchmarks · 1 papers
text-guided dyadic HHOI generation
1 benchmarks · 1 papers
recipe2im
1 benchmarks · 1 papers
Detour video retrieval
1 benchmarks · 1 papers
Video Embedding Evaluation
1 benchmarks · 1 papers
High-Resolution Multi-modal Search
1 benchmarks · 1 papers
Text-and-Image to Audio-Visual Generation
1 benchmarks · 1 papers
Text to Audio-Visual Generation
1 benchmarks · 1 papers
Visual Speech Translation
1 benchmarks · 1 papers
Trajectory-based Video Question Answering
1 benchmarks · 1 papers
Temporal clip retrieval (text-to-video)
1 benchmarks · 1 papers
Multimodal Complaint Scene Reasoning
1 benchmarks · 1 papers
Multi-modal Role-playing
1 benchmarks · 1 papers
Multimodal Role-playing
1 benchmarks · 1 papers
Ground-Truth-to-Generated Retrieval
1 benchmarks · 1 papers
Multi-modal Multi-image Reasoning
1 benchmarks · 1 papers
Multi-turn visual dialogue
Page 81 of 83
Previous
Next