Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Multimodal AI research – Page 3 · SOTA2 Research
Back to Research
Research domain
Multimodal
Search tasks
Search
22 benchmarks · 21 papers
Open-Ended Visual Question Answering
21 benchmarks · 2 papers
Multimodal Abstractive Summarization
21 benchmarks · 20 papers
Long Video Question Answering
20 benchmarks · 17 papers
Image Captioning Evaluation
20 benchmarks · 14 papers
Visual Storytelling
20 benchmarks · 60 papers
Object Hallucination
20 benchmarks · 12 papers
Image-Text Matching
20 benchmarks · 18 papers
Open-ended Video Question Answering
19 benchmarks · 136 papers
Text-based Visual Question Answering
19 benchmarks · 2 papers
Vision-Language Modeling
18 benchmarks · 5 papers
Multimodal Fake News Detection
18 benchmarks · 27 papers
Multi-modal Question Answering
18 benchmarks · 16 papers
Text-Guided Image Editing
17 benchmarks · 11 papers
fMRI-to-image reconstruction
17 benchmarks · 50 papers
Multimodal Benchmarking
17 benchmarks · 13 papers
Spatial Grounding
17 benchmarks · 7 papers
Gaze Following
17 benchmarks · 8 papers
Audio-visual speech separation
17 benchmarks · 6 papers
Fine-grained Visual Reasoning
17 benchmarks · 11 papers
Visual Spatial Reasoning
17 benchmarks · 7 papers
Multimodal Generation
17 benchmarks · 17 papers
Visual Instruction Following
16 benchmarks · 8 papers
Audio-to-image retrieval
16 benchmarks · 16 papers
Gesture Generation
16 benchmarks · 5 papers
Sketch-based image retrieval
16 benchmarks · 6 papers
Video-to-Music Generation
16 benchmarks · 14 papers
Multimodal Machine Translation
16 benchmarks · 22 papers
Audio-visual understanding
16 benchmarks · 9 papers
Multimodal Safety Evaluation
15 benchmarks · 6 papers
Multimodal Summarization
15 benchmarks · 33 papers
Vision Understanding
15 benchmarks · 5 papers
Token-level hallucination detection
Page 3 of 83
Previous
Next