EgoRoC: Towards Egocentric Robotic Control via Task-Agnostic Visual Alignment
Recent Vision-Language-Action (VLA) models map visual-textual inputs to robotic actions via end-to-end architectures, yet this approach entangles visual understanding with task-specific actions. This leads to an exhaustive collection of full operational sequences and parameter redundancy across task