Descriptive Report Generation on Gut-VLM
0.89ROUGE-LLLaVa-1.6-7b
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| LLaVa-1.6-7bBackbone=LLaVa-1.6-7b, Training Protocol=Hallucination-aware finetuning2025.05 | 0.89 | 0.82 | 0.9 | 3.96 | 90.89 | |
| DeepSeek7bBackbone=DeepSeek-7b, Training Protocol=Hallucination-aware finetuning2025.05 | 0.88 | 0.81 | 0.9 | 3.63 | 88.77 | |
| Qwen7bBackbone=Qwen-7b, Training Protocol=Hallucination-aware finetuning2025.05 | 0.88 | 0.82 | 0.9 | 4.04 | 90.53 | |
| ChatGPT-4 OmniBackbone=ChatGPT-4, Training Protocol=Proprietary/Original2025.05 | 0.87 | 0.8 | 0.89 | 2.97 | 85.99 | |
| MPlugOwl2bBackbone=MPlugOwl-2b, Training Protocol=Hallucination-aware finetuning2025.05 | 0.85 | 0.77 | 0.87 | 3.72 | 88.4 | |
| DeepSeek7bBackbone=DeepSeek-7b, Training Protocol=Standard Finetuned2025.05 | 0.55 | 0.37 | 0.65 | 3.76 | 83.73 | |
| LLaVa-1.6-7bBackbone=LLaVa-1.6-7b, Training Protocol=Standard Finetuned2025.05 | 0.54 | 0.35 | 0.63 | 3.71 | 83.07 | |
| Qwen7bBackbone=Qwen-7b, Training Protocol=Standard Finetuned2025.05 | 0.54 | 0.37 | 0.64 | 3.78 | 83.27 | |
| MPlugOwl2bBackbone=MPlugOwl-2b, Training Protocol=Standard Finetuned2025.05 | 0.5 | 0.32 | 0.6 | 3.68 | 82.9 | |
| Qwen7bBackbone=Qwen-7b, Training Protocol=Pretrained2025.05 | 0.32 | 0.12 | 0.48 | 1.74 | 67.57 | |
| DeepSeek7bBackbone=DeepSeek-7b, Training Protocol=Pretrained2025.05 | 0.29 | 0.11 | 0.39 | 1.65 | 65.2 | |
| LLaVa-1.6-7bBackbone=LLaVa-1.6-7b, Training Protocol=Pretrained2025.05 | 0.26 | 0.1 | 0.47 | 1.36 | 50.89 | |
| MPlugOwl2bBackbone=MPlugOwl-2b, Training Protocol=Pretrained2025.05 | 0.26 | 0.09 | 0.44 | 1.34 | 55.29 |