Image Captioning on Natural Scenes Dataset (NSD) (test)
38.91BLEU-1Neuro-Vision to Language
Evaluation Results
| Method | Links | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Neuro-Vision to Languagenumber_of_models=1, ViT3D=true2024.04 | 38.91 | 24.02 | 15.24 | 12.41 | 18.44 | 27.83 | 42.58 | 18.41 | 56.16 | — | — | |
| Neuro-Vision to Languagenumber_of_models=1, ViT3D=false2024.04 | 33.57 | 18.95 | 11.09 | 6.13 | 15.56 | 23.8 | 20.23 | 16.21 | 51.47 | — | — | |
| Mind-Omni (B→I&T)Trainable Parameters=442M, External LLM=None, # Models=1, # Tasks=72026.05 | 29.12 | 17.63 | 11.36 | — | 26.05 | 30.54 | 12.26 | 13.25 | 53.67 | 52.75 | 87.73 | |
| UMBRAETrainable Parameters=146.57M, External LLM=Vicuna (13B), # Models=1, # Tasks=22026.05 | 21.39 | 11.86 | 6.31 | — | 11.31 | 17.6 | 6.04 | 6.43 | 60.85 | 65.96 | — | |
| OneLLMTrainable Parameters=7B, External LLM=LLaMA2 (7B), # Models=4, # Tasks=72026.05 | 18.41 | 12.13 | 5.34 | — | 9.37 | 17.13 | 9.41 | 6.34 | 50.31 | 51.43 | 86.18 | |
| Mind-Omni (B→T)Trainable Parameters=442M, External LLM=None, # Models=1, # Tasks=72026.05 | 14.92 | 9.03 | 5.83 | — | 13.86 | 13.35 | 6.28 | 6.78 | 48.47 | 47 | 80.21 |