← Search

Ying Wu

42 accepted papers

2026

VLION: Vision-Language Guided Interactive Object Navigation with Mobile Manipulation

ICRA 2026poster

Object navigation for mobile robots typically assumes that targets are visible and paths are unobstructed. However, real-world scenarios often involve occluded targets like objects hidden behind doors or inside containers. Such scenarios require interactive navigation and manipulation by mobile mani…

Cited by 0Scholar
2025

A Generic Continuous Multi-Joint Spinal Robotic System for Agile and Accurate Behaviors with GNN-MPC method

RSS 2025poster

The biomimetic research of vertebrates is challenging in both mechanism design and control methods. Motivated by natural acrobatics exhibited by cats and humans, this paper presents a generic multi-joint continuous spinal system and a learning-based algorithm for agile and accurate control. The spin…

Cited by 0PDFScholar
2025

AutoScape: Geometry-Consistent Long-Horizon Scene Generation

ICCV 2025poster

This paper proposes AutoScape, a long-horizon driving scene generation framework. At its core is a novel RGB-D diffusion model that iteratively generates sparse, geometrically consistent keyframes, serving as reliable anchors for the scene's appearance and geometry. To maintain long-range geometric…

Cited by 0SourcePDFScholar
2025

GPVK-VL: Geometry-Preserving Virtual Keyframes for Visual Localization under Large Viewpoint Changes

CVPR 2025poster

Visual localization, the task of determining the position and orientation of a camera, typically involves three core components: offline construction of a keyframe database, efficient online keyframes retrieval, and robust local feature matching. However, significant challenges arise when there are…

Cited by 0SourcePDFScholar
2025

Learning Variable Whole-Body Control for Agile Aerial Manipulation in Strong Winds

RA-L 2025

Aerial manipulation provides an effective alternative to human labor in high-risk outdoor situations. Complex and variable environments demand the system to respond quickly with minimal latency to external disturbances. To address this challenge, we propose a learning-based variable whole-body model

Cited by 1SourceScholar
2024

AIDE: An Automatic Data Engine for Object Detection in Autonomous Driving

CVPR 2024poster

Autonomous vehicle (AV) systems rely on robust perception models as a cornerstone of safety assurance. However objects encountered on the road exhibit a long-tailed distribution with rare or unseen categories posing challenges to a deployed perception model. This necessitates an expensive process of…

Cited by 15SourcePDFScholar
2024

Active Open-Vocabulary Recognition: Let Intelligent Moving Mitigate CLIP Limitations

CVPR 2024poster

Active recognition which allows intelligent agents to explore observations for better recognition performance serves as a prerequisite for various embodied AI tasks such as grasping navigation and room arrangements. Given the evolving environment and the multitude of object classes it is impractical…

Cited by 4SourcePDFScholar
2024

Evidential Active Recognition: Intelligent and Prudent Open-World Embodied Perception

CVPR 2024poster

Active recognition enables robots to intelligently explore novel observations thereby acquiring more information while circumventing undesired viewing conditions. Recent approaches favor learning policies from simulated or collected data wherein appropriate actions are more frequently selected when…

Cited by 6SourcePDFScholar
2024

Learning to Ask Denotative and Connotative Questions for Knowledge-based VQA

EMNLP 2024finding

Large language models (LLMs) have attracted increasing attention due to its prominent performance on various tasks. Recent works seek to leverage LLMs on knowledge-based visual question answering (VQA) tasks which require common sense knowledge to answer the question about an image, since LLMs have…

Cited by 0SourcePDFScholar
2024

Robust and Energy-Efficient Control for Multi-task Aerial Manipulation with Automatic Arm-switching

ICRA 2024poster

Aerial manipulation has received increasing research interest with wide applications of drones. To perform specific tasks, robotic arms with various mechanical structures will be mounted on the drone. It results in sudden disturbances to the aerial manipulator when switching the robotic arm or inter…

Cited by 4SourceScholar
2023

Flexible Visual Recognition by Evidential Modeling of Confusion and Ignorance

ICCV 2023poster

In real-world scenarios, typical visual recognition systems could fail under two major causes, i.e., the misclassification between known classes and the excusable misbehavior on unknown-class images. To tackle these deficiencies, flexible visual recognition should dynamically predict multiple classe…

Cited by 5PDFScholar
2023

SynMob: Creating High-Fidelity Synthetic GPS Trajectory Dataset for Urban Mobility Analysis

NeurIPS 2023poster

Urban mobility analysis has been extensively studied in the past decade using a vast amount of GPS trajectory data, which reveals hidden patterns in movement and human activity within urban landscapes. Despite its significant value, the availability of such datasets often faces limitations due to pr…

2022

Balancing between Forgetting and Acquisition in Incremental Subpopulation Learning

ECCV 2022poster

"The subpopulation shifting challenge, known as some subpopulations of a category that are not seen during training, severely limits the classification performance of the state-of-the-art convolutional neural networks. Thus, to mitigate this practical issue, we explore incremental subpopulation lear…

2021

Biomimetic Flip-and-Flap Strategy of Flying Objects for Perching on Inclined Surfaces

RA-L 2021

Animals can use the maneuver of a flipping body and flapping wings to reduce the normal rebound force of impact during landing, decreasing the adsorption force required by the contact point. This capability aids aerial vehicles with landing not only on vertical surfaces, but also on inclined surface

Cited by 11SourceScholar
2021

Contrastive Learning for Label Efficient Semantic Segmentation

ICCV 2021poster

Collecting labeled data for the task of semantic segmentation is expensive and time-consuming, as it requires dense pixel-level annotations. While recent Convolutional Neural Network (CNN) based semantic segmentation approaches have achieved impressive results by using large amounts of labeled train…

Cited by 209PDFcodeScholar
2020

Object Detection with a Unified Label Space from Multiple Datasets

ECCV 2020poster

Given multiple datasets with different label spaces, the goal of this work is to train a single object detector predicting over the union of all the label spaces. The practical benefits of such an object detector are obvious and significant---application-relevant categories can be picked and merged…

2020

Online Joint Multi-Metric Adaptation From Frequent Sharing-Subset Mining for Person Re-Identification

CVPR 2020poster

Person Re-IDentification (P-RID), as an instance-level recognition problem, still remains challenging in computer vision community. Many P-RID works aim to learn faithful and discriminative features/metrics from offline training data and directly use them for the unseen online testing data. However,…

Cited by 61PDFScholar
2020

Uncertainty-Aware Score Distribution Learning for Action Quality Assessment

CVPR 2020oral

Assessing action quality from videos has attracted growing attention in recent years. Most existing approaches usually tackle this problem based on regression algorithms, which ignore the intrinsic ambiguity in the score labels caused by multiple judges or their subjective appraisals. To address thi…

Cited by 171PDFcodeScholar
2019

Learning Robust Facial Landmark Detection via Hierarchical Structured Ensemble

ICCV 2019poster

Heatmap regression-based models have significantly advanced the progress of facial landmark detection. However, the lack of structural constraints always generates inaccurate heatmaps resulting in poor landmark detection performance. While hierarchical structure modeling methods have been proposed t…

Cited by 80PDFScholar
2019

Recognizing Part Attributes With Insufficient Data

ICCV 2019poster

Recognizing the attributes of objects and their parts is central to many computer vision applications. Although great progress has been made to apply object-level recognition, recognizing the attributes of parts remains less applicable since the training data for part attributes recognition is usual…

Cited by 22PDFcodeScholar
2019

Visual Query Answering by Entity-Attribute Graph Matching and Reasoning

CVPR 2019poster

Visual Query Answering (VQA) is of great significance in offering people convenience: one can raise a question for details of objects, or high-level understanding about the scene, over an image. This paper proposes a novel method to address the VQA problem. In contrast to prior works, our method th…

Cited by 24PDFScholar
2018

A Modulation Module for Multi-task Learning with Applications in Image Retrieval

ECCV 2018poster

Multi-task learning has been widely adopted in many computer vision tasks to improve overall computation efficiency or boost the performance of individual tasks, under the assumption that those tasks are correlated and complementary to each other. However, the relationships between the tasks are com…

2018

Easy Identification From Better Constraints: Multi-Shot Person Re-Identification From Reference Constraints

CVPR 2018poster

Multi-shot person re-identification (MsP-RID) utilizes multiple images from the same person to facilitate identification. Considering the fact that motion information may not be discriminative nor reliable enough for MsP-RID, this paper is focused on handling the large variations in the visual appea…

Cited by 21SourcePDFScholar
2018

Exploiting Vector Fields for Geometric Rectification of Distorted Document Images

ECCV 2018poster

This paper proposes a segment-free method for geometric rectification of a distorted document image captured by a hand-held camera. The method can recover the 3D page shape by exploiting the intrinsic vector fields of the image. Based on the assumption that the curled page shape is a general cylindr…

Cited by 29SourcePDFScholar
2017

Efficient Online Local Metric Adaptation via Negative Samples for Person Re-Identification

ICCV 2017poster

Many existing person re-identification (PRID) methods typically attempt to train a faithful global metric offline to cover the enormous visual appearance variations, so as to directly use it online on various probes for identity matching. However, their need for a huge set of positive training pairs…

Cited by 98PDFScholar