CoVAR: Co-Generation of Video and Action for Robotic Manipulation Via Multi-Modal Diffusion
We present a method to generate video–action pairs that follow text instructions, starting from an initial image observation and the robot’s joint states. Our approach automatically provides action labels for video diffusion mod- els, overcoming the common lack of action annotations and enabling the…