Human Understanding on Ego-in-Exo, NeXTQA, VideoMME, ADL-X
-11.04Delta SourceStandard Adaptation
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| Standard AdaptationVL-C=true, VE=true, LLM=true, Training data=VIMA-Bench + xArm-Det2025.10 | -11.04 | 64.5 | 83 | 63 | 36.04 | 62.24 | 61.76 | |
| Adaptation Strategy (Trainable VE, Trainable LLM)Training data=VIMA-Bench + xArm-Det, VE=Trainable, LLM=Trainable2025.10 | -11.04 | — | — | — | — | — | — | |
| VISCOPVL-C=true, VE=VISCOP, LLM=true, Training data=VIMA-Bench + xArm-Det2025.10 | -11 | 59.59 | 82.98 | 63.26 | 36.32 | 66.83 | 61.79 | |
| VISCOPTraining data=VIMA-Bench + xArm-Det2025.10 | -11 | — | — | — | — | — | — | |
| Standard AdaptationVL-C=true, VE=true, LLM=true, Training data=VIMA-Bench2025.10 | -8.87 | 56.92 | 83.24 | 62.74 | 52.21 | 64.5 | 63.92 | |
| Finetuning Strategy (Trainable VE, Trainable LLM)Vision Encoder=trainable, LLM=trainable, Vision Language Connector=trainable2025.10 | -8.87 | 56.92 | 83.24 | 62.74 | 52.21 | 64.5 | 63.92 | |
| Adaptation Strategy (Trainable VE, Trainable LLM)Training data=VIMA-Bench, VE=Trainable, LLM=Trainable2025.10 | -8.87 | — | — | — | — | — | — | |
| Finetuning Strategy (Trainable VE, Frozen LLM)Vision Encoder=trainable, LLM=frozen, Vision Language Connector=trainable2025.10 | -7.84 | 59.42 | 83.16 | 64.41 | 52.92 | 64.86 | 64.95 | |
| VISCOPVL-C=true, VE=VISCOP, LLM=true, Training data=VIMA-Bench2025.10 | -4.58 | 71.19 | 83.71 | 63.67 | 55.89 | 66.62 | 68.22 | |
| VISCOP2025.10 | -4.58 | 71.19 | 83.71 | 63.67 | 55.89 | 66.62 | 68.22 | |
| VISCOPTraining data=VIMA-Bench2025.10 | -4.58 | — | — | — | — | — | — | |
| Base VLM2025.10 | — | 66.27 | 84.32 | 65.37 | 77.36 | 70.65 | 72.79 | |
| Base VLM2025.10 | — | 66.27 | 84.32 | 65.37 | 77.36 | 70.65 | 72.79 |