ResearchBenchmarksPreference-Aligned Reinforcement Learning on Straighten Rope Pass WaterFollow4.1Label CountPrefVLM0.0441.0972.153.203Jul 2, 2026Evaluation ResultsMethodMethodLinksLabel CountImage Statistic ValuePrefVLM2026.074.10.4RL-VLM-F2026.0724ERL-VLM2026.070.90.9CoRe2026.070.21.6