← Search

Xiang Yu

34 accepted papers

2026

Learning-Based Observer for Coupled Disturbance

ICRA 2026poster

Achieving high-precision control for robotic systems is hindered by the low-fidelity dynamical model and external disturbances. Especially, the intricate coupling between internal uncertainties and external disturbances further exacerbates this challenge. This study introduces an effective and conve…

2026

Unified Meta-Representation and Feedback Calibration for General Disturbance Estimation

ICRA 2026poster

Precise control in modern robotic applications is always an open issue due to unknown time-varying disturbances. Existing meta-learning-based approaches require a shared representation of environmental structures, which lack flexibility for realistic non-structural disturbances. Besides, representat…

2025

Feedback Favors the Generalization of Neural ODEs

ICLR 2025oral

The well-known generalization problem hinders the application of artificial neural networks in continuous-time prediction tasks with varying latent dynamics. In sharp contrast, biological systems can neatly adapt to evolving environments benefiting from real-time feedback mechanisms. Inspired by the…

Cited by 1SourcePDFScholar
2025

Revisiting Source-Free Domain Adaptation: Insights into Representativeness, Generalization, and Variety

CVPR 2025poster

Domain adaptation addresses the challenge where the distribution of target inference data differs from that of the source training data. Recently, data privacy has become a significant constraint, limiting access to the source domain. To mitigate this issue, Source-Free Domain Adaptation (SFDA) meth…

Cited by 0SourcePDFScholar
2024

Flying in Narrow Spaces: Prioritizing Safety With Disturbance-Aware Control

RA-L 2024

Safe and autonomous flight of quadrotors in enclosed environments still remains formidable challenge due to the aerodynamic proximity effect and restricted free space. This letter develops an integrated planning and control architecture for disturbance-aware and safety control of quadrotors. By expl

Cited by 8SourceScholar
2023

A Safety Planning and Control Architecture Applied to a Quadrotor Autopilot

RA-L 2023

This letter presents a safety trajectory planning and tracking architecture for a quadrotor autopilot. Motor saturation constraints are explicitly considered in obstacle avoidance mission. Two challenging cases are covered: agile flight with a short task time and stable flight with actuator degradat

Cited by 12SourceScholar
2023

DeFormer: Integrating Transformers with Deformable Models for 3D Shape Abstraction from a Single Image

ICCV 2023poster

Explicit 3D shape abstraction from a single 2D image is a long-standing problem in computer vision and graphics. By leveraging a set of primitives to represent the target shape, recent methods have achieved promising results. However, these methods either use a relatively larger number of primitives…

Cited by 8PDFScholar
2023

Domain Generalization Guided by Gradient Signal to Noise Ratio of Parameters

ICCV 2023poster

Overfitting to the source domain is a common issue in gradient-based training of deep neural networks. To compensate for the over-parameterized models, numerous regularization techniques have been introduced such as those based on dropout. While these methods achieve significant improvements on clas…

Cited by 5PDFScholar
2023

Q: How To Specialize Large Vision-Language Models to Data-Scarce VQA Tasks? A: Self-Train on Unlabeled Images!

CVPR 2023poster

Finetuning a large vision language model (VLM) on a target dataset after large scale pretraining is a dominant paradigm in visual question answering (VQA). Datasets for specialized tasks such as knowledge-based VQA or VQA in non natural-image domains are orders of magnitude smaller than those for ge…

2023

Selective Structured State-Spaces for Long-Form Video Understanding

CVPR 2023poster

Effective modeling of complex spatiotemporal dependencies in long-form videos remains an open problem. The recently proposed Structured State-Space Sequence (S4) model with its linear complexity offers a promising direction in this space. However, we demonstrate that treating all image-tokens equall…

Cited by 126SourcePDFScholar
2022

Accurate High-Maneuvering Trajectory Tracking for Quadrotors: A Drag Utilization Method

RA-L 2022

The balanceness between the tracking performance and the aerodynamic drag treatment is of paramount importance especially in the presence of the quadrotor aggressive maneuvers. Different from standard approaches that achieve precise tracking by feedforward compensating the estimated drag, this work

Cited by 39SourceScholar
2022

Controllable Dynamic Multi-Task Architectures

CVPR 2022oral

Multi-task learning commonly encounters competition for resources among tasks, specifically when model capacity is limited. This challenge motivates models which allow control over the relative importance of tasks and total compute cost during inference time. In this work, we propose such a controll…

Cited by 35PDFScholar
2022

Learning Phase Mask for Privacy-Preserving Passive Depth Estimation

ECCV 2022poster

"With over a billion sold each year, cameras are not only becoming ubiquitous, but are driving progress in a wide range of domains such as mixed reality, robotics, and more. However, severe concerns regarding the privacy implications of camera-based solutions currently limit the range of environment…

Cited by 16SourcePDFScholar
2022

Learning To Learn Across Diverse Data Biases in Deep Face Recognition

CVPR 2022poster

Convolutional Neural Networks have achieved remarkable success in face recognition, in part due to the abundant availability of data. However, the data used for training CNNs is often imbalanced. Prior works largely focus on the long-tailed nature of face datasets in data volume per identity, or foc…

Cited by 26PDFScholar
2022

MMDF: Multi-Modal Deep Feature Based Place Recognition of Mobile Robots With Applications on Cross-Scene Navigation

RA-L 2022

Although the navigation of robots in urban environments has achieved great performance, there is still a problem of insufficient robustness in cross-scene (ground, water surface) navigation applications. An intuitive idea is to introduce multi-modal complementary data to improve the robustness of th

Cited by 16SourceScholar
2022

On Generalizing Beyond Domains in Cross-Domain Continual Learning

CVPR 2022poster

In the real world, humans have the ability to accumulate new knowledge in any conditions. However, deeplearning suffers from the phenomenon so-called catastrophic forgetting of the previously observed knowledge after learning a new task. Many recent methods focus on preventing catastrophic forgettin…

Cited by 42PDFScholar
2022

Single-Stream Multi-level Alignment for Vision-Language Pretraining

ECCV 2022poster

"Self-supervised vision-language pretraining from pure images and text with a contrastive loss is effective, but ignores fine-grained alignment due to a dual-stream architecture that aligns image and text representations only on a global level. Earlier, supervised, non-contrastive methods were capab…

2021

Cross-Domain Similarity Learning for Face Recognition in Unseen Domains

CVPR 2021poster

Face recognition models trained under the assumption of identical training and test distributions often suffer from poor generalization when faced with unknown variations, such as a novel ethnicity or unpredictable individual make-ups during test time. In this paper, we introduce a novel cross-domai…

Cited by 30PDFScholar
2021

Learning Cross-Modal Contrastive Features for Video Domain Adaptation

ICCV 2021poster

Learning transferable and domain adaptive feature representations from videos is important for video-relevant tasks such as action recognition. Existing video domain adaptation methods mainly rely on adversarial feature alignment, which has been derived from the RGB image space. However, video data…

Cited by 94PDFScholar
2020

Improving Face Recognition by Clustering Unlabeled Faces in the Wild

ECCV 2020poster

While deep face recognition has benefited significantly from large-scale labeled data, current research is focused on leveraging unlabeled data to further boost performance, reducing the cost of human annotation. Prior work has mostly been in controlled settings, where the labeled and unlabeled data…

Cited by 22SourcePDFScholar
2020

Towards Universal Representation Learning for Deep Face Recognition

CVPR 2020poster

Recognizing wild faces is extremely hard as they appear with all kinds of variations. Traditional methods either train with specifically annotated variation data from target domains, or by introducing unlabeled target variation data to adapt from the training data. Instead, we propose a universal re…

Cited by 197PDFScholar
2019

Feature Transfer Learning for Face Recognition With Under-Represented Data

CVPR 2019poster

Despite the large volume of face recognition datasets, there is a significant portion of subjects, of which the samples are insufficient and thus under-represented. Ignoring such significant portion results in insufficient training data. Training with under-represented data leads to biased classifie…

Cited by 396PDFScholar
2019

Gotta Adapt 'Em All: Joint Pixel and Feature-Level Domain Adaptation for Recognition in the Wild

CVPR 2019poster

Recent developments in deep domain adaptation have allowed knowledge transfer from a labeled source domain to an unlabeled target domain at the level of intermediate features or input pixels. We propose that advantages may be derived by combining them, in the form of different insights that lead to…

Cited by 54PDFScholar
2019

Unsupervised Domain Adaptation for Distance Metric Learning

ICLR 2019poster

Unsupervised domain adaptation is a promising avenue to enhance the performance of deep neural networks on a target domain, using labels only from a source domain. However, the two predominant methods, domain discrepancy reduction learning and semi-supervised learning, are not readily applicable whe…

Cited by 66SourcePDFScholar
2017

Deep Supervision With Shape Concepts for Occlusion-Aware 3D Object Parsing

CVPR 2017poster

Monocular 3D object parsing is highly desirable in various scenarios including occlusion reasoning and holistic scene interpretation. We present a deep convolutional neural network (CNN) architecture to localize semantic parts in 2D image and 3D space while inferring their visibility states, given a…

Cited by 110PDFScholar
2017

Learning Efficient Object Detection Models with Knowledge Distillation

NeurIPS 2017poster

Despite significant accuracy improvement in convolutional neural networks (CNN) based object detectors, they often require prohibitive runtimes to process an image for real-time applications. State-of-the-art models often use very deep networks with a large number of floating point operations. Effor…

Cited by 1342SourcePDFScholar
2017

Reconstruction-Based Disentanglement for Pose-Invariant Face Recognition

ICCV 2017poster

Deep neural networks (DNNs) trained on large-scale datasets have recently achieved impressive improvements in face recognition. But a persistent challenge remains to develop methods capable of handling large pose variations that are relatively under-represented in training data. This paper presents…

Cited by 189PDFScholar
2017

Unsupervised Domain Adaptation for Face Recognition in Unlabeled Videos

ICCV 2017poster

Despite rapid advances in face recognition, there remains a clear gap between the performance of still image-based face recognition and video-based face recognition, due to the vast difference in visual quality between the domains and the difficulty of curating diverse large-scale video datasets. Th…

Cited by 147PDFScholar