← Search

Shuqiang Jiang

33 accepted papers

2026

Eating for a Sustainable Planet: Personalized Sustainable Diet Recommendation via Constraint-Aware Decision-Making Modeling

ICML 2026poster

A sustainable diet represents a multi-dimensional synergy among four essential pillars: nutrition adequacy, economic affordability, cultural acceptability, and environmental respect. Despite the prevalence of population-level sustainability modeling, practical implementation relies on effective indi…

Cited by 0SourceScholar
2026

GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation

CVPR 2026

Despite significant progress in Vision-Language Navigation (VLN), existing approaches still rely on dense RGB videos that produce excessive patch tokens and lack explicit spatial structure, resulting in substantial computational overhead and limited spatial reasoning. To address these issues, we int

Cited by 0SourceScholar
2026

Joint Navigation and Manipulation Planning with 3D Interaction Chains

ICML 2026poster

Open-vocabulary mobile manipulation (OVMM) requires long-horizon navigation in unseen environments and object-centric manipulation. Most existing methods treat navigation and manipulation as separate stages, which can yield navigation endpoints that are poor for manipulation or manipulation-friendly…

Cited by 0SourceScholar
2026

Multi-Scale Gaussian-Language Map for Zero-shot Embodied Navigation and Reasoning

CVPR 2026

Understanding the geometric and semantic structure of environments is essential for embodied agents. Existing semantic mapping methods trade off between explicit geometry and multi-scale semantics,and lack a native interface for large models, thus requiring additional training of feature projection

Cited by 0SourceScholar
2026

OmniFood8K: Single-Image Nutrition Estimation via Hierarchical Frequency-Aligned Fusion

CVPR 2026

Accurate estimation of food nutrition plays a vital role in promoting healthy dietary habits and personalized diet management. Most existing food datasets primarily focus on Western cuisines and lack sufficient coverage of Chinese dishes, which restricts accurate nutritional estimation for Chinese m

Cited by 0SourcecodeScholar
2026

TrajRAG: Retrieving Geometric-Semantic Experience for Zero-Shot Object Navigation

CVPR 2026

Existing zero-shot Object Goal Navigation (ObjectNav) methods often exploit commonsense knowledge from large language or vision-language models to guide navigation. However, such knowledge arises from internet-scale text rather than embodied 3D experience, and episodic observations collected during

Cited by 0SourceScholar
2025

DiffGen: Robot Demonstration Generation via Differentiable Physics Simulation, Differentiable Rendering, and Vision-Language Model

IROS 2025

Generating robot demonstrations through simulation is widely recognized as an effective way to scale up robot data. Previous work often trained reinforcement learning agents to generate expert policies, but this approach lacks sample efficiency. Recently, a line of work has attempted to generate rob

Cited by 3SourceScholar
2025

Function-centric Bayesian Network for Zero-Shot Object Goal Navigation

ICCV 2025poster

Object goal navigation requires an agent to navigate to a specified target in unseen environments without an explicit map, which demands an understanding of object-scene context to infer the target's location based on partial observations. The function of an object plays a crucial role in its catego…

Cited by 0SourcePDFScholar
2025

Learning on the Go: A Meta-learning Object Navigation Model

ICCV 2025poster

Object navigation tasks require an agent to locate a target object using visual observations in unseen environments, where unfamiliar layouts and novel object appearances can hinder navigation. Most existing methods lack the adaptability needed to handle these uncertainties, as their navigation mode…

Cited by 0SourcePDFScholar
2024

A Category Agnostic Model for Visual Rearrangment

CVPR 2024poster

This paper presents a novel category agnostic model for visual rearrangement task which can help an embodied agent to physically recover the shuffled scene configuration without any category concepts to the goal configuration. Previous methods usually follow a similar architecture completing the rea…

Cited by 2SourcePDFScholar
2024

An Interactive Navigation Method with Effect-oriented Affordance

CVPR 2024poster

Visual navigation is to let the agent reach the target according to the continuous visual input. In most previous works visual navigation is usually assumed to be done in a static and ideal environment: the target is always reachable with no need to alter the environment. However the "messy" environ…

2024

Imagine Before Go: Self-Supervised Generative Map for Object Goal Navigation

CVPR 2024poster

The Object Goal navigation (ObjectNav) task requires the agent to navigate to a specified target in an unseen environment. Since the environment layout is unknown the agent needs to infer the unknown contextual objects from partially observations thereby deducing the likely location of the target. P…

2024

Lookahead Exploration with Neural Radiance Representation for Continuous Vision-Language Navigation

CVPR 2024highlight

Vision-and-language navigation (VLN) enables the agent to navigate to a remote location following the natural language instruction in 3D environments. At each navigation step the agent selects from possible candidate locations and then makes the move. For better navigation planning the lookahead exp…

2024

Sim-to-Real Transfer via 3D Feature Fields for Vision-and-Language Navigation

CoRL 2024poster

Vision-and-language navigation (VLN) enables the agent to navigate to a remote location in 3D environments following the natural language instruction. In this field, the agent is usually trained and evaluated in the navigation simulators, lacking effective approaches for sim-to-real transfer. The VL…

Cited by 11SourcecodeScholar
2024

Trajectory Diffusion for ObjectGoal Navigation

NeurIPS 2024poster

Object goal navigation requires an agent to navigate to a specified object in an unseen environment based on visual observations and user-specified goals. Human decision-making in navigation is sequential, planning a most likely sequence of actions toward the goal. However, existing ObjectNav meth…

Cited by 1SourcePDFScholar
2023

CaMP: Causal Multi-policy Planning for Interactive Navigation in Multi-room Scenes

NeurIPS 2023poster

Visual navigation has been widely studied under the assumption that there may be several clear routes to reach the goal. However, in more practical scenarios such as a house with several messy rooms, there may not. Interactive Navigation (InterNav) considers agents navigating to their goals more eff…

2023

Data-Free Knowledge Distillation via Feature Exchange and Activation Region Constraint

CVPR 2023poster

Despite the tremendous progress on data-free knowledge distillation (DFKD) based on synthetic data generation, there are still limitations in diverse and efficient data synthesis. It is naive to expect that a simple combination of generative network-based data synthesis and data augmentation will so…

2023

GridMM: Grid Memory Map for Vision-and-Language Navigation

ICCV 2023poster

Vision-and-language navigation (VLN) enables the agent to navigate to a remote location following the natural language instruction in 3D environments. To represent the previously visited environment, most approaches for VLN implement memory using recurrent states, topological maps, or top-down seman…

Cited by 59PDFcodeScholar
2023

KERM: Knowledge Enhanced Reasoning for Vision-and-Language Navigation

CVPR 2023poster

Vision-and-language navigation (VLN) is the task to enable an embodied agent to navigate to a remote location following the natural language instruction in real scenes. Most of the previous approaches utilize the entire features or object-centric features to represent navigable candidates. However,…

2023

Layout-Based Causal Inference for Object Navigation

CVPR 2023poster

Previous works for ObjectNav task attempt to learn the association (e.g. relation graph) between the visual inputs and the goal during training. Such association contains the prior knowledge of navigating in training environments, which is denoted as the experience. The experience performs a positiv…

Cited by 38SourcePDFScholar
2022

Generative Meta-Adversarial Network for Unseen Object Navigation

ECCV 2022poster

"Object navigation is a task to let the agent navigate to a target object. Prevailing works attempt to expand navigation ability in new environments and achieve reasonable performance on the seen object categories that have been observed in training environments. However, this setting is somewhat li…

2022

Rethinking the Optimization of Average Precision: Only Penalizing Negative Instances before Positive Ones Is Enough

AAAI 2022technical

Optimising the approximation of Average Precision (AP) has been widely studied for image retrieval. Limited by the definition of AP, such methods consider both negative and positive instances ranking before each positive instance. However, we claim that only penalizing negative instances before posi…

2021

Hierarchical Object-to-Zone Graph for Object Navigation

ICCV 2021poster

The goal of object navigation is to reach the expected objects according to visual information in the unseen environments. Previous works usually implement deep models to train an agent to predict actions in real-time. However, in the unseen environment, when the target object is not in egocentric v…

Cited by 85PDFcodeScholar
2021

See More for Scene: Pairwise Consistency Learning for Scene Classification

NeurIPS 2021poster

Scene classification is a valuable classification subtask and has its own characteristics which still needs more in-depth studies. Basically, scene characteristics are distributed over the whole image, which cause the need of “seeing” comprehensive and informative regions. Previous works mainly focu…

Cited by 6SourcePDFScholar
2021

What If We Could Not See? Counterfactual Analysis for Egocentric Action Anticipation

IJCAI 2021poster

Egocentric action anticipation aims at predicting the near future based on past observation in first-person vision. While future actions may be wrongly predicted due to the dataset bias, we present a counterfactual analysis framework for egocentric action anticipation (CA-EAA) to enhance the capacit…

Cited by 16SourcePDFScholar
2020

Multi-attention Meta Learning for Few-shot Fine-grained Image Recognition

IJCAI 2020poster

The goal of few-shot image recognition is to distinguish different categories with only one or a few training samples. Previous works of few-shot learning mainly work on general object images. And current solutions usually learn a global image representation from training tasks to adapt novel tasks.…

Cited by 0SourcePDFScholar
2015

Joint Multi-Feature Spatial Context for Scene Recognition on the Semantic Manifold

CVPR 2015poster

In the semantic multinomial framework patches and images are modeled as points in a semantic probability simplex. Patch theme models are learned resorting to weak supervision via image labels, which leads the problem of scene categories co-occurring in this semantic space. Fortunately, each category…

Cited by 38SourcePDFScholar