CEI: A Unified Interface for Cross-Embodiment Visuomotor Policy Learning in 3D Space
Tong Wu, Shoujie Li, Junhao Gong, Changqing Guo, Xingting Li, Shilong Mu, Wenbo Ding
Abstract
Robotic foundation models trained on large-scale manipulation datasets have shown promise in learning generalist policies, but they often overfit to specific viewpoints, robot arms, and especially parallel-jaw grippers due to dataset biases. To address this limitation, we propose Cross-Embodiment Interface (<italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">CEI</i>), a framework for cross-embodiment learning that enables the transfer of demonstrations across different robot arm and end-effector morphologies. <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">CEI</i> introduces the concept of <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">functional similarity</i>, which is quantified using Directional Chamfer Distance. Then it aligns robot trajectories through gradient-based optimization, followed by synthesizing observations and actions for unseen robot arms and end-effectors. In experiments, <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">CEI</i> transfers data and policies from a Franka Panda robot to <bold xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">16</b> different embodiments across <bold xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">3</b> tasks in simulation, and supports bidirectional transfer between a UR5+AG95 gripper robot and a UR5+Xhand robot across <bold xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">6</b> real-world tasks, achieving an average transfer ratio of 82.4%. Finally, we demonstrate that <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">CEI</i> can also be extended with spatial generalization and multimodal motion generation capabilities using our proposed techniques. Project website: <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://cross-embodiment-interface.github.io/</uri>.
BibTeX
@inproceedings{ral2026_ceiaunifiedinter,
title = {CEI: A Unified Interface for Cross-Embodiment Visuomotor Policy Learning in 3D Space},
author = {Tong Wu and Shoujie Li and Junhao Gong and Changqing Guo and Xingting Li and Shilong Mu and Wenbo Ding},
booktitle = {RA-L 2026},
year = {2026}
}