Video-to-command generation on standard object sets
0.618BLEU-1Human-to-Robot (Ours)
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Human-to-Robot (Ours)Backbone=ResNet-1012026.02 | 0.618 | 0.511 | 0.403 | 0.351 | |
| Human-to-Robot (Ours)Backbone=ResNet-502026.02 | 0.607 | 0.506 | 0.397 | 0.337 | |
| Human-to-Robot (Ours w/o key frame)Backbone=ResNet-1012026.02 | 0.577 | 0.465 | 0.363 | 0.298 | |
| Human-to-Robot (Ours w/o key frame)Backbone=ResNet-502026.02 | 0.562 | 0.464 | 0.361 | 0.297 | |
| Human-to-Robot (Ours w/o object sel.)Backbone=ResNet-1012026.02 | 0.524 | 0.411 | 0.303 | 0.237 | |
| Human-to-Robot (Ours w/o object sel.)Backbone=ResNet-502026.02 | 0.521 | 0.407 | 0.3 | 0.237 | |
| Watch-and-ActBackbone=InceptionV32026.02 | 0.399 | 0.287 | 0.263 | 0.199 | |
| Watch-and-ActBackbone=ResNet-502026.02 | 0.394 | 0.271 | 0.248 | 0.187 | |
| V2CBackbone=ResNet-502026.02 | 0.357 | 0.231 | 0.201 | 0.153 | |
| V2CBackbone=InceptionV32026.02 | 0.355 | 0.227 | 0.201 | 0.164 | |
| Watch-and-ActBackbone=ResNet-1012026.02 | 0.339 | 0.171 | 0.155 | 0.134 | |
| V2CBackbone=ResNet-1012026.02 | 0.324 | 0.163 | 0.153 | 0.131 | |
| Video2CommandBackbone=InceptionV32026.02 | 0.323 | 0.196 | 0.162 | 0.154 | |
| Video2CommandBackbone=ResNet-1012026.02 | 0.319 | 0.183 | 0.175 | 0.148 | |
| Video2CommandBackbone=ResNet-502026.02 | 0.312 | 0.196 | 0.173 | 0.15 |