2025
CapsDT: Diffusion-Transformer for Capsule Robot Manipulation
IROS 2025
Vision-Language-Action (VLA) models have emerged as a prominent research area, showcasing significant potential across a variety of applications. However, their performance in endoscopy robotics, particularly endoscopy capsule robots that perform actions within the digestive system, remains unexplor