Vision-Language Conversation and Reasoning on LLaVA-Bench (test)
70.3Complex AccuracyDAC
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| DACBase Model=LLaVA-1.5, Sampling strategy=nucleus sampling (p = 1), Evaluator=text-only GPT-4 API2025.02 | 70.3 | 50 | 72.7 | 64.3 | |
| VCDBase Model=LLaVA-1.5, Sampling strategy=nucleus sampling (p = 1), Evaluator=text-only GPT-4 API2025.02 | 69.6 | 51.6 | 57.3 | 61.6 | |
| SIDBase Model=LLaVA-1.5, Sampling strategy=nucleus sampling (p = 1), Evaluator=text-only GPT-4 API2025.02 | 66.7 | 51.3 | 66.3 | 60.4 | |
| OPERABase Model=LLaVA-1.5, Sampling strategy=beam search, Evaluator=text-only GPT-4 API2025.02 | 66.4 | 56.9 | 44 | 61.3 | |
| BaselineBase Model=LLaVA-1.5, Sampling strategy=nucleus sampling (p = 1), Evaluator=text-only GPT-4 API2025.02 | 66.3 | 46.7 | 68.7 | 60.6 | |
| CCABase Model=LLaVA-1.5, Sampling strategy=nucleus sampling (p = 1), Evaluator=text-only GPT-4 API2025.02 | 66.1 | 53.9 | 69.4 | 64.3 |