S3K: Self-Supervised Semantic Keypoints for Robotic Manipulation via Multi-View Consistency
A robot’s ability to act is fundamentally constrained by what it can perceive. Many existing approaches to visual representation learning utilize general-purpose training criteria, e.g. image reconstruction, smoothness in latent space, or usefulness for control, or else make use of large datasets an