← Search

Xinhang Song

19 accepted papers

2026

Joint Navigation and Manipulation Planning with 3D Interaction Chains

ICML 2026poster

Open-vocabulary mobile manipulation (OVMM) requires long-horizon navigation in unseen environments and object-centric manipulation. Most existing methods treat navigation and manipulation as separate stages, which can yield navigation endpoints that are poor for manipulation or manipulation-friendly…

Cited by 0SourceScholar
2026

Multi-Scale Gaussian-Language Map for Zero-shot Embodied Navigation and Reasoning

CVPR 2026

Understanding the geometric and semantic structure of environments is essential for embodied agents. Existing semantic mapping methods trade off between explicit geometry and multi-scale semantics,and lack a native interface for large models, thus requiring additional training of feature projection

Cited by 0SourceScholar
2026

TrajRAG: Retrieving Geometric-Semantic Experience for Zero-Shot Object Navigation

CVPR 2026

Existing zero-shot Object Goal Navigation (ObjectNav) methods often exploit commonsense knowledge from large language or vision-language models to guide navigation. However, such knowledge arises from internet-scale text rather than embodied 3D experience, and episodic observations collected during

Cited by 0SourceScholar
2025

Function-centric Bayesian Network for Zero-Shot Object Goal Navigation

ICCV 2025poster

Object goal navigation requires an agent to navigate to a specified target in unseen environments without an explicit map, which demands an understanding of object-scene context to infer the target's location based on partial observations. The function of an object plays a crucial role in its catego…

Cited by 0SourcePDFScholar
2025

Learning on the Go: A Meta-learning Object Navigation Model

ICCV 2025poster

Object navigation tasks require an agent to locate a target object using visual observations in unseen environments, where unfamiliar layouts and novel object appearances can hinder navigation. Most existing methods lack the adaptability needed to handle these uncertainties, as their navigation mode…

Cited by 0SourcePDFScholar
2024

A Category Agnostic Model for Visual Rearrangment

CVPR 2024poster

This paper presents a novel category agnostic model for visual rearrangement task which can help an embodied agent to physically recover the shuffled scene configuration without any category concepts to the goal configuration. Previous methods usually follow a similar architecture completing the rea…

Cited by 2SourcePDFScholar
2024

An Interactive Navigation Method with Effect-oriented Affordance

CVPR 2024poster

Visual navigation is to let the agent reach the target according to the continuous visual input. In most previous works visual navigation is usually assumed to be done in a static and ideal environment: the target is always reachable with no need to alter the environment. However the "messy" environ…

2024

Imagine Before Go: Self-Supervised Generative Map for Object Goal Navigation

CVPR 2024poster

The Object Goal navigation (ObjectNav) task requires the agent to navigate to a specified target in an unseen environment. Since the environment layout is unknown the agent needs to infer the unknown contextual objects from partially observations thereby deducing the likely location of the target. P…

2024

Trajectory Diffusion for ObjectGoal Navigation

NeurIPS 2024poster

Object goal navigation requires an agent to navigate to a specified object in an unseen environment based on visual observations and user-specified goals. Human decision-making in navigation is sequential, planning a most likely sequence of actions toward the goal. However, existing ObjectNav meth…

Cited by 1SourcePDFScholar
2023

CaMP: Causal Multi-policy Planning for Interactive Navigation in Multi-room Scenes

NeurIPS 2023poster

Visual navigation has been widely studied under the assumption that there may be several clear routes to reach the goal. However, in more practical scenarios such as a house with several messy rooms, there may not. Interactive Navigation (InterNav) considers agents navigating to their goals more eff…

2023

Layout-Based Causal Inference for Object Navigation

CVPR 2023poster

Previous works for ObjectNav task attempt to learn the association (e.g. relation graph) between the visual inputs and the goal during training. Such association contains the prior knowledge of navigating in training environments, which is denoted as the experience. The experience performs a positiv…

Cited by 38SourcePDFScholar
2022

Generative Meta-Adversarial Network for Unseen Object Navigation

ECCV 2022poster

"Object navigation is a task to let the agent navigate to a target object. Prevailing works attempt to expand navigation ability in new environments and achieve reasonable performance on the seen object categories that have been observed in training environments. However, this setting is somewhat li…

2021

Hierarchical Object-to-Zone Graph for Object Navigation

ICCV 2021poster

The goal of object navigation is to reach the expected objects according to visual information in the unseen environments. Previous works usually implement deep models to train an agent to predict actions in real-time. However, in the unseen environment, when the target object is not in egocentric v…

Cited by 85PDFcodeScholar
2021

See More for Scene: Pairwise Consistency Learning for Scene Classification

NeurIPS 2021poster

Scene classification is a valuable classification subtask and has its own characteristics which still needs more in-depth studies. Basically, scene characteristics are distributed over the whole image, which cause the need of “seeing” comprehensive and informative regions. Previous works mainly focu…

Cited by 6SourcePDFScholar
2015

Joint Multi-Feature Spatial Context for Scene Recognition on the Semantic Manifold

CVPR 2015poster

In the semantic multinomial framework patches and images are modeled as points in a semantic probability simplex. Patch theme models are learned resorting to weak supervision via image labels, which leads the problem of scene categories co-occurring in this semantic space. Fortunately, each category…

Cited by 38SourcePDFScholar