Vision-Language Understanding on MMStar
42.64AccuracySmoothSMoE annealed (k=2)
Evaluation Results
| Method | Links | |
|---|---|---|
| SmoothSMoE annealed (k=2)k=2, Smoothing mechanism=Annealed SmoothSMoE, Backbone=ViT-based, Model parameters=5.6B, Fine-tuning dataset=50% of the LLaVA-665K dataset, Number of experts=42026.06 | 42.64 | |
| SmoothSMoE (k=2.5)k=2.5, Smoothing mechanism=SmoothSMoE, Backbone=ViT-based, Model parameters=5.6B, Fine-tuning dataset=50% of the LLaVA-665K dataset, Number of experts=42026.06 | 41.98 | |
| SMoE (k=2)k=2, Smoothing mechanism=None, Backbone=ViT-based, Model parameters=5.6B, Fine-tuning dataset=50% of the LLaVA-665K dataset, Number of experts=42026.06 | 40.95 |