ICRA 2026poster0 citations

Multi-View Gating Unit with KL-Based Alignment Toward Real-World Robot Control

Kei Igarashi, Shingo Murata

Abstract

This paper proposes a framework for integrating latent representations from multi-view images, using adaptive weighting based on situational context to facilitate the generation of robot actions. Specifically, we introduce the multi-view gating unit (MGU), which assigns context-dependent weights to each dimension of the latent representations extracted from different viewpoints. By summing the corresponding dimensions across all viewpoints, we construct a fused latent representation that serves as input to a policy model. To enhance the effectiveness of the MGU and improve the accuracy of action generation, we incorporate a Kullback–Leibler (KL)-based alignment objective that encourages consistency between individual viewpoint representations and the fused representation. We evaluate the proposed framework through imitation-learning experiments in a kitchen-like real-robot environment across five tasks. The experimental results show that the MGU dynamically adapts to different contexts, thereby enabling successful task execution. Additionally, we compare our approach with a modified Action Chunking with Transformers (ACT) baseline and conduct an ablation study to assess the contribution of each component. The results show that our method achieves a task success rate of 84%, outperforming all baseline methods and validating the effectiveness of both the individual components and their integration within the proposed framework.

Cognitive Control ArchitecturesMachine Learning for Robot ControlRepresentation Learning
Multi-View Gating Unit with KL-Based Alignment Toward Real-World Robot Control · ICRA 2026