Semantic-Aware Behavior Optimization With Safety Reinforcement Feedback for Language-Conditioned Manipulation
Xiuxiu Qi, Jiannong Cao, Julie A. McCann, Chongshan Fan, Hanqian Luo, Hongpeng Wang
Abstract
Language-conditioned manipulation (LCM) couples vision, language, and control to enable natural instruction following, making it a promising direction in embodied AI research. Researchers have combined imitation learning (IL) with reinforcement learning (RL), where IL provides sample-efficient initialization and RL refines the policy through feedback to maximize long-term returns. Despite progress, safe LCM execution faces two challenges: (1) a modality gap that disrupts semantic-spatial correspondence, degrading action quality, and (2) sparse safety costs in high-dimensional states that make constraint estimation and semantic–safety alignment difficult. In this paper, we present Semantic-Aware Behavior Optimization with Safety Reinforcement Feedback, a novel framework that mitigates multimodal misalignment through context-aware initialization and enforces semantic–safety consistency during execution. We introduce spatial–semantic priors to align language goals with spatial observations and embed safety cues. Specifically, a goal prior aligns modalities for context-aware initialization during behavioral cloning, while a constraint prior combined with a feasibility metric provides learnable supervision for safety reinforcement fine-tuning. Evaluations show our framework achieves a 6.7% relative improvement in simulation and effectively reduces safety violations in real-world tasks across diverse scenarios.
BibTeX
@inproceedings{ral2026_semanticawarebeh,
title = {Semantic-Aware Behavior Optimization With Safety Reinforcement Feedback for Language-Conditioned Manipulation},
author = {Xiuxiu Qi and Jiannong Cao and Julie A. McCann and Chongshan Fan and Hanqian Luo and Hongpeng Wang},
booktitle = {RA-L 2026},
year = {2026}
}