← Search

Qiao Gu

11 accepted papers

2026

MomaGraph: State-Aware Unified Scene Graphs with Vision-Language Models for Embodied Task Planning

ICLR 2026oral

Mobile manipulators in households must both navigate and manipulate. This requires a compact, semantically rich scene representation that captures where objects are, how they function, and which parts are actionable. Scene graphs are a natural choice, yet prior work often separates spatial and funct…

Cited by 0SourcecodeScholar
2026

SICNav-Diffusion: Safe and Interactive Crowd Navigation with Diffusion Trajectory Predictions

ICRA 2026poster

To navigate crowds without collisions, robots must interact with humans by forecasting their future motion and reacting accordingly. While learning-based prediction models have shown success in generating likely human trajectory predictions, integrating these stochastic models into a robot controlle…

2025

SAFE: Multitask Failure Detection for Vision-Language-Action Models

NeurIPS 2025poster

While vision-language-action models (VLAs) have shown promising robotic behaviors across a diverse set of manipulation tasks, they achieve limited success rates when deployed on novel tasks out of the box. To allow these policies to safely interact with their environments, we need a failure detector…

Cited by 0SourcecodeScholar
2025

SICNav-Diffusion: Safe and Interactive Crowd Navigation With Diffusion Trajectory Predictions

RA-L 2025

To navigate crowds without collisions, robots must interact with humans by forecasting their future motion and reacting accordingly. While learning-based prediction models have shown success in generating likely human trajectory predictions, integrating these stochastic models into a robot controlle

Cited by 10SourcecodeScholar
2024

ConceptGraphs: Open-Vocabulary 3D Scene Graphs for Perception and Planning

ICRA 2024poster

For robots to perform a wide variety of tasks, they require a 3D representation of the world that is semantically rich, yet compact and efficient for task-driven perception and planning. Recent approaches have attempted to leverage features from large vision-language models to encode semantics in 3D…

Cited by 202SourceScholar
2023

ConceptFusion: Open-set multimodal 3D mapping

RSS 2023poster

Building 3D maps of the environment is central to robot navigation, planning, and interaction with objects in a scene. Most existing approaches that integrate semantic concepts with 3D maps largely remain confined to the closed-set setting: they can only reason about a finite set of concepts, pre-de…

2023

Preserving Linear Separability in Continual Learning by Backward Feature Projection

CVPR 2023poster

Catastrophic forgetting has been a major challenge in continual learning, where the model needs to learn new tasks with limited or no access to data from previously seen tasks. To tackle this challenge, methods based on knowledge distillation in feature space have been proposed and shown to reduce f…

2021

Deep Video Matting via Spatio-Temporal Alignment and Aggregation

CVPR 2021poster

Despite the significant progress made by deep learning in natural image matting, there has been so far no representative work on deep learning for video matting due to the inherent technical challenges in reasoning temporal domain and lack of large-scale video matting datasets. In this paper, we pro…

Cited by 64PDFcodeScholar
2019

LADN: Local Adversarial Disentangling Network for Facial Makeup and De-Makeup

ICCV 2019poster

We propose a local adversarial disentangling network (LADN) for facial makeup and de-makeup. Central to our method are multiple and overlapping local adversarial discriminators in a content-style disentangling network for achieving local detail transfer between facial images, with the use of asymmet…

Cited by 130PDFcodeScholar