Video-to-command generation on Novel object sets Zero-shot (test)
48.1BLEU-2Human-to-Robot
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Human-to-RobotBackbone=ResNet-1012026.02 | 48.1 | 57.2 | 36.1 | 26.5 | |
| Human-to-RobotBackbone=ResNet-502026.02 | 47.1 | 57.7 | 35.6 | 26.1 | |
| Human-to-Robot (w/o key frame)Backbone=ResNet-1012026.02 | 43.5 | 53.9 | 29.8 | 21.8 | |
| Human-to-Robot (w/o key frame)Backbone=ResNet-502026.02 | 43.1 | 53.5 | 29.5 | 21.7 | |
| Human-to-Robot (w/o object sel.)Backbone=ResNet-1012026.02 | 36.2 | 47.8 | 25.8 | 19.3 | |
| Human-to-Robot (w/o object sel.)Backbone=ResNet-502026.02 | 36.1 | 47.4 | 25.1 | 19.1 | |
| Watch-and-ActBackbone=ResNet-502026.02 | 15.7 | 31 | 13.4 | 10.7 | |
| Video2CommandBackbone=ResNet-502026.02 | 15 | — | 9 | 8.2 | |
| Watch-and-ActBackbone=InceptionV32026.02 | 14.1 | 29.3 | 13.3 | 11.6 | |
| Video2CommandBackbone=InceptionV32026.02 | 13.2 | 24.4 | 7.5 | 6.2 | |
| V2CBackbone=ResNet-502026.02 | 11.9 | 25.9 | 9.2 | 8.7 | |
| V2CBackbone=InceptionV32026.02 | 11.7 | 27.9 | 9.3 | 8.7 | |
| Watch-and-ActBackbone=ResNet-1012026.02 | 11.3 | 23.3 | 10.9 | — | |
| Video2CommandBackbone=ResNet-1012026.02 | 10.7 | 22.6 | 9.4 | 8.1 | |
| V2CBackbone=ResNet-1012026.02 | 10.2 | 22.2 | 9.5 | 8.6 |