SARA: Semantic-Anchored Referential Alignment for Ego–Exo Instance Correspondence
Yueyang Ge, Ye Lin, Bojun Yang, Yilin Huang, Yi Guo
Abstract
Achieving visual coordination between egocentric (ego) and exocentric (exo) perspectives is a cornerstone of augmented reality and human-machine collaboration. However, extreme viewpoint shifts and occlusions often cause semantic drift in existing methods, making robust cross-view correspondence a formidable challenge. In this paper, we propose SARA (Semantic-Anchored Referential Alignment), a framework designed to bridge the ego-exo gap by leveraging viewpoint-invariant, high-level semantic information as region-level semantic anchors for robust feature mapping. Specifically, our approach incorporates: (1) the Semantic Anchoring Module extracts cross-view category priors and invariant features to provide stable region-level references across perspectives, and (2) the Multimodal Semantic Vector Fusion mechanism achieves language-guided manifold fusion of semantic vectors to synthesize unified embeddings, effectively establishing representational correspondence across disparate perspectives. Extensive experiments across diverse, complex scenarios demonstrate SARA’s superior performance in establishing stable mappings,validating the effectiveness of semantic-invariant features and multimodal fusion in addressing the core hurdles of cross-view understanding.
BibTeX
@inproceedings{ijcai2026_sarasemanticanch,
title = {SARA: Semantic-Anchored Referential Alignment for Ego–Exo Instance Correspondence},
author = {Yueyang Ge and Ye Lin and Bojun Yang and Yilin Huang and Yi Guo},
booktitle = {IJCAI 2026},
year = {2026}
}