General Multimodal Understanding on Vision-Language Benchmark Suite
55.07Average ScoreSmoothSMoE (k=2.5)
Evaluation Results
| Method | Links | |
|---|---|---|
| SmoothSMoE (k=2.5)k=2.5, Smoothing mechanism=SmoothSMoE, Backbone=ViT-based, Model parameters=5.6B, Fine-tuning dataset=50% of the LLaVA-665K dataset, Number of experts=42026.06 | 55.07 | |
| SmoothSMoE annealed (k=2)k=2, Smoothing mechanism=Annealed SmoothSMoE, Backbone=ViT-based, Model parameters=5.6B, Fine-tuning dataset=50% of the LLaVA-665K dataset, Number of experts=42026.06 | 54.98 | |
| SMoE (k=2)k=2, Smoothing mechanism=None, Backbone=ViT-based, Model parameters=5.6B, Fine-tuning dataset=50% of the LLaVA-665K dataset, Number of experts=42026.06 | 54.53 |