LLaVA
Benchmarks
Task NameDataset NameSOTA ResultTrendResults
LLaVA cc3m
3.3IR@1
6
LLaVA 1.5
0.96KMR (a)
6
LLaVA-NeXT Inference
7.998Inference Time (s)
6
LLaVA Evaluation Suite Flickr30k 1.5
78.52VQAv2 Accuracy
6
LLaVA Eval
75.19Helpfulness Rating
6
LLaVA FOA-Attack in-domain
98Precision
5
LLaVA M-Attack in-domain
98.8Precision
5
LLaVA SSA-CWA in-domain
99Precision
5
LLaVA 7B 1.5
67ASR
5
LLaVA Evaluation Suite v1.5
59.7MMBench
5
LLaVA Adv N=8 OneVision (full evaluation set)
60.16Accuracy
4
LLaVA Adv N=4 OneVision (full evaluation set)
74.32Accuracy
4
LLaVA Random N=8 OneVision (full evaluation set)
94.92Accuracy
4
LLaVA Random N=4 full OneVision (evaluation)
99.04Accuracy
4
LLaVA
16.8ASR
4
LLaVA Video Sequence 1.6
2.2Compression Ratio (×)
3
LLaVA (VQAv2, GQA, VisWiz, SQA, VQAT, POPE, MMBench) 1.5 (test val)
67.6Overall Average Score
3
LLaVA Benchmark (LLV^B)
84.7LLV^B Score
3
LLaVA v1 (test)
84.4Conversation Score
3
LLaVA Pre-training
85ASR
2
LLava img-chat workload
94.6Pairwise Accuracy
2
LLava vid-chat workload
90.8Pairwise Accuracy
2
LLaVA-W In-the-Wild (subset of 500 samples)
81.2AIM-CoT Win Rate
2