Semantic Generalization on RL4VLA MultiCarrot (test)
63.5Success RateVLA Grounder
Evaluation Results
| Method | Links | |
|---|---|---|
| VLA GrounderVLA Backbone=OpenVLA, Command Policy=Qwen3.5-9B, Learning Method=GRPO2026.07 | 63.5 | |
| VLA GrounderVLA Backbone=OpenVLA, Command Policy=Qwen3.5-9B, Learning Method=w/o GRPO2026.07 | 60.9 | |
| OpenVLAVLA Backbone=OpenVLA, Command Policy=Original instruction (orig)2026.07 | 59.4 | |
| VLA GrounderVLA Backbone=π0, Command Policy=Qwen3.5-9B, Learning Method=GRPO2026.07 | 29.7 | |
| VLA GrounderVLA Backbone=π0, Command Policy=Qwen3.5-9B, Learning Method=w/o GRPO2026.07 | 21.4 | |
| π0VLA Backbone=π0, Command Policy=Original instruction (orig)2026.07 | 17.2 |