Semantic Generalization on RL4VLA MultiPlate (test)
64.6Success RateVLA Grounder
Evaluation Results
| Method | Links | |
|---|---|---|
| VLA GrounderVLA Backbone=OpenVLA, Command Policy=Qwen3.5-9B, Learning Method=GRPO2026.07 | 64.6 | |
| VLA GrounderVLA Backbone=OpenVLA, Command Policy=Qwen3.5-9B, Learning Method=w/o GRPO2026.07 | 56.3 | |
| OpenVLAVLA Backbone=OpenVLA, Command Policy=Original instruction (orig)2026.07 | 54.7 | |
| VLA GrounderVLA Backbone=π0, Command Policy=Qwen3.5-9B, Learning Method=GRPO2026.07 | 17.2 | |
| VLA GrounderVLA Backbone=π0, Command Policy=Qwen3.5-9B, Learning Method=w/o GRPO2026.07 | 10.4 | |
| π0VLA Backbone=π0, Command Policy=Original instruction (orig)2026.07 | 9.9 |