Robotic Control on VIMA-Bench Simulation (evaluation set)
69.62L1 ScoreFinetuning Strategy (Trainable VE, Trainable LLM)
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Finetuning Strategy (Trainable VE, Trainable LLM)Vision Encoder=trainable, LLM=trainable, Vision Language Connector=trainable2025.10 | 69.62 | 60.77 | 65 | 65.13 | 65.13 | |
| VISCOP2025.10 | 67.69 | 65.77 | 70 | 67.82 | 67.82 | |
| Finetuning Strategy (Trainable VE, Frozen LLM)Vision Encoder=trainable, LLM=frozen, Vision Language Connector=trainable2025.10 | 63.46 | 63.08 | 68.75 | 65.1 | 65.1 | |
| Base VLM2025.10 | 0 | 0 | 0 | 0 | — |