IJCAI 20260 citations

RepSAM: Bridging Foundation Models to Robotic Vision via Representation-Guided Adaptation

Wenhui Chu

Abstract

Robotic perception in unstructured environments remains challenging despite the zero-shot capabilities of foundation models such as SAM. This work attributes performance degradation to non-uniform representation shifts across transformer layers: shallow layers exhibit substantial domain gaps (CKA < 0.5), whereas deep layers transfer effectively (CKA > 0.7). Based on this observation, we propose RepSAM, a representation-guided parameter-efficient fine-tuning (PEFT) framework for adapting foundation models to robotic vision. RepSAM employs a theoretically grounded CKA-guided rank allocation strategy combined with a multi-modal fusion module for robust handling of challenging robotic scenarios, including transparent objects and cluttered scenes. Experimental evaluation across six benchmarks and robotic manipulation tasks demonstrates that RepSAM achieves 97.9% of full fine-tuning performance (89.0% vs. 90.9% mIoU) while reducing trainable parameters by 158× (from 632M to 4.0M). RepSAM outperforms DoRA by 7.9% mIoU with just 4 hours of training on a single A100 GPU (a 96× reduction from full fine-tuning, which takes 384 GPU-hours). These improvements are statistically significant (p<0.01) and translate to a 12.0% absolute improvement in robotic manipulation success rates over the LoRA (RGB) baseline.

Generative AI, robotic foundation models, and reinforcement learning: Grounding large models in real robot interaction with cross-task, cross-environment, cross-platform, and open-domain generalization and transferLearning to understand, generalize, and explain actions: Robust generalization and transfer across tasks, objects, environments, embodiments, and long-horizon scenariosSafety, trustworthiness, generalizability, and evaluation: Verification and validation methods for AI-powered robotic systemsSafety, trustworthiness, generalizability, and evaluation: Theoretical and algorithmic guarantees on safety, robustness, and out-of-distribution generalization
BibTeX
@inproceedings{ijcai2026_repsambridgingfo,
  title = {RepSAM: Bridging Foundation Models to Robotic Vision via Representation-Guided Adaptation},
  author = {Wenhui Chu},
  booktitle = {IJCAI 2026},
  year = {2026}
}
RepSAM: Bridging Foundation Models to Robotic Vision via Representation-Guided Adaptation · IJCAI 2026