Efficient Alignment of Unconditioned Action Prior for Language-Conditioned Pick and Place in Clutter (I)
We study the task of language-conditioned pick and place in clutter, where a robot should grasp a target object in open clutter and move it to a specified place. Some approaches learn end-to-end policies with features from vision foundation models, requiring large datasets. Others combine foundation…