← Search

Xiaodong Yang

35 accepted papers

2026

ScaLoRA: Optimally Scaled Low-Rank Adaptation for Efficient High-Rank Fine-Tuning

ICML 2026poster

As large language models (LLMs) continue to scale in size, the computational overhead has become a major bottleneck for task-specific fine-tuning. While low-rank adaptation (LoRA) effectively curtails this cost by confining the weight updates to a low-dimensional subspace, such a restriction can hin…

Cited by 0SourceScholar
2025

All-directional Disparity Estimation for Real-world QPD Images

CVPR 2025highlight

Quad Photodiode (QPD) sensors represent an evolution by providing four sub-views, whereas dual-pixel (DP) sensors are limited to two sub-views. In addition to enhancing auto-focus performance, QPD sensors also enable disparity estimation in horizontal and vertical directions. However, the characteri…

Cited by 0SourcePDFScholar
2025

EfficientSleepNet: A Novel Lightweight End-to-End Model for Automated Sleep Staging on Single-Channel EEG

ICASSP 2025accepted

Sleep staging is critical for evaluating sleep quality and regulating sleep patterns. While deep learning has shown potential for automatically scoring sleep stages from raw signals, many existing models are overly complex, computationally intensive, and rely on future information, limiting their us…

Cited by 0SourceScholar
2025

MoPFormer: Motion-Primitive Transformer for Wearable-Sensor Activity Recognition

NeurIPS 2025poster

Human Activity Recognition (HAR) with wearable sensors is challenged by limited interpretability, which significantly impacts cross-dataset generalization. To address this challenge, we propose Motion-Primitive Transformer (MoPFormer), a novel self-supervised framework that enhances interpretability…

Cited by 0SourceScholar
2025

Quad-Pixel Image Defocus Deblurring: A New Benchmark and Model

CVPR 2025poster

Defocus deblurring is a challenging task due to the spatially varying blur. Recent works have shown impressive results in data-driven approaches using dual-pixel (DP) sensors. Quad-pixel (QP) sensors represent an advanced evolution of DP sensors, providing four distinct sub-aperture views in contras…

Cited by 0SourcePDFScholar
2024

Benign Oscillation of Stochastic Gradient Descent with Large Learning Rate

ICLR 2024poster

In this work, we theoretically investigate the generalization properties of neural networks (NN) trained by stochastic gradient descent (SGD) with large learning rates. Under such a training regime, our finding is that, the oscillation of the NN weights caused by SGD with large learning rates turns…

Cited by 15SourcePDFScholar
2024

Cross-Modal Self-Supervised Learning with Effective Contrastive Units for LiDAR Point Clouds

IROS 2024poster

3D perception in LiDAR point clouds is crucial for a self-driving vehicle to properly act in 3D environment. However, manually labeling point clouds is hard and costly. There has been a growing interest in self-supervised pre-training of 3D perception models. Following the success of contrastive lea…

Cited by 2SourcecodeScholar
2024

FedES: Federated Early-Stopping for Hindering Memorizing Heterogeneous Label Noise

IJCAI 2024poster

Federated learning (FL) facilitates collaborative model training across distributed clients while maintaining privacy. Federated noisy label learning (FNLL) is more of a challenge for data inaccessibility and noise heterogeneity. Existing works primarily assume clients are either noisy or clean, whi…

Cited by 1SourcePDFScholar
2024

RAM-Avatar: Real-time Photo-Realistic Avatar from Monocular Videos with Full-body Control

CVPR 2024poster

This paper focuses on advancing the applicability of human avatar learning methods by proposing RAM-Avatar which learns a Real-time photo-realistic Avatar that supports full-body control from Monocular videos. To achieve this goal RAM-Avatar leverages two statistical templates responsible for modeli…

Cited by 3SourcePDFScholar
2023

DistillBEV: Boosting Multi-Camera 3D Object Detection with Cross-Modal Knowledge Distillation

ICCV 2023poster

3D perception based on the representations learned from multi-camera bird's-eye-view (BEV) is trending as cameras are cost-effective for mass production in autonomous driving industry. However, there exists a distinct performance gap between multi-camera BEV and LiDAR based 3D object detection. One…

Cited by 36PDFcodeScholar
2023

PillarNeXt: Rethinking Network Designs for 3D Object Detection in LiDAR Point Clouds

CVPR 2023poster

In order to deal with the sparse and unstructured raw point clouds, most LiDAR based 3D object detection research focuses on designing dedicated local point aggregators for fine-grained geometrical modeling. In this paper, we revisit the local point aggregators from the perspective of allocating com…

2023

ProphNet: Efficient Agent-Centric Motion Forecasting With Anchor-Informed Proposals

CVPR 2023highlight

Motion forecasting is a key module in an autonomous driving system. Due to the heterogeneous nature of multi-sourced input, multimodality in agent behavior, and low latency required by onboard deployment, this task is notoriously challenging. To cope with these difficulties, this paper proposes a no…

Cited by 70SourcePDFScholar
2023

Transcendental Idealism of Planner: Evaluating Perception from Planning Perspective for Autonomous Driving

ICML 2023poster

Evaluating the performance of perception modules in autonomous driving is one of the most critical tasks in developing the complex intelligent system. While module-level unit test metrics adopted from traditional computer vision tasks are feasible to some extent, it remains far less explored to meas…

2020

Bridging Cross-Tasks Gap for Cognitive Assessment via Fine-Grained Domain Adaptation

IJCAI 2020poster

Discriminating pathologic cognitive decline from the expected decline of normal aging is an important research topic for elderly care and health monitoring. However, most cognitive assessment methods only work when data distributions of the training set and testing set are consistent. Enabling exist…

Cited by 0SourcePDFScholar
2020

Contrastive Learning for Weakly Supervised Phrase Grounding

ECCV 2020poster

Phrase grounding, the problem of associating image regions to caption words, is a crucial component of vision-language tasks. We show that phrase grounding can be learned by optimizing word-region attention to maximize a lower bound on mutual information between images and caption words. Given pairs…

2020

Instance-Aware, Context-Focused, and Memory-Efficient Weakly Supervised Object Detection

CVPR 2020poster

Weakly supervised learning has emerged as a compelling tool for object detection by reducing the need for strong supervision during training. However, major challenges remain: (1) differentiation of object instances can be ambiguous; (2) detectors tend to focus on discriminative parts rather than en…

Cited by 261PDFcodeScholar
2020

Joint Disentangling and Adaptation for Cross-Domain Person Re-Identification

ECCV 2020poster

Although a significant progress has been witnessed in supervised person re-identification (re-id), it remains challenging to generalize re-id models to new domains due to the huge domain gaps. Recently, there has been a growing interest in using unsupervised domain adaptation to address this scalabi…

2020

Simulating Content Consistent Vehicle Datasets with Attribute Descent

ECCV 2020poster

This paper uses a graphic engine to simulate a large amount of training data with free annotations. Between synthetic and real data, there is a two-level domain gap, i.e., content level and appearance level. While the latter has been widely studied, we focus on reducing the content gap in attributes…

2020

UFO²: A Unified Framework towards Omni-supervised Object Detection

ECCV 2020poster

Existing work on object detection often relies on a single form of annotation: the model is trained using either accurate yet costly bounding boxes or cheaper but less expressive image-level tags. However, real-world annotations are often diverse in form, which challenges these existing works. In th…

2019

A Delay Metric for Video Object Detection: What Average Precision Fails to Tell

ICCV 2019poster

Average precision (AP) is a widely used metric to evaluate detection accuracy of image and video object detectors. In this paper, we analyze the object detection from video and point out that mAP alone is not sufficient to capture the temporal nature of video object detection. To tackle this problem…

Cited by 60PDFcodeScholar
2019

CityFlow: A City-Scale Benchmark for Multi-Target Multi-Camera Vehicle Tracking and Re-Identification

CVPR 2019oral

Urban traffic optimization using traffic cameras as sensors is driving the need to advance state-of-the-art multi-target multi-camera (MTMC) tracking. This work introduces CityFlow, a city-scale traffic camera dataset consisting of more than 3 hours of synchronized HD videos from 40 cameras across 1…

Cited by 511PDFcodeScholar
2019

Dancing to Music

NeurIPS 2019poster

Dancing to music is an instinctive move by humans. Learning to model the music-to-dance generation process is, however, a challenging problem. It requires significant efforts to measure the correlation between music and dance as one needs to simultaneously consider multiple aspects, such as style an…

2019

Joint Discriminative and Generative Learning for Person Re-Identification

CVPR 2019oral

Person re-identification (re-id) remains challenging due to significant intra-class variations across different cameras. Recently, there has been a growing interest in using generative models to augment training data and enhance the invariance to input changes. The generative pipelines in existing m…

Cited by 1005PDFScholar
2019

PAMTRI: Pose-Aware Multi-Task Learning for Vehicle Re-Identification Using Highly Randomized Synthetic Data

ICCV 2019poster

In comparison with person re-identification (ReID), which has been widely studied in the research community, vehicle ReID has received less attention. Vehicle ReID is challenging due to 1) high intra-class variability (caused by the dependency of shape and appearance on viewpoint), and 2) small inte…

Cited by 147PDFcodeScholar
2019

STEP: Spatio-Temporal Progressive Learning for Video Action Detection

CVPR 2019oral

In this paper, we propose Spatio-TEmporal Progressive (STEP) action detector--a progressive learning framework for spatio-temporal action detection in videos. Starting from a handful of coarse-scale proposal cuboids, our approach progressively refines the proposals towards actions over a few steps.…

Cited by 200PDFScholar
2018

MoCoGAN: Decomposing Motion and Content for Video Generation

CVPR 2018poster

Visual signals in a video can be divided into content and motion. While content specifies which objects are in the video, motion describes their dynamics. Based on this prior, we propose the Motion and Content decomposed Generative Adversarial Network (MoCoGAN) framework for video generation. The pr…

2018

PWC-Net: CNNs for Optical Flow Using Pyramid, Warping, and Cost Volume

CVPR 2018poster

We present a compact but effective CNN model for optical flow, called PWC-Net. PWC-Net has been designed according to simple and well-established principles: pyramidal processing, warping, and the use of a cost volume. Cast in a learnable feature pyramid, PWC-Net uses the current optical flow estima…

2017

Dynamic Facial Analysis: From Bayesian Filtering to Recurrent Neural Network

CVPR 2017poster

Facial analysis in videos, including head pose estimation and facial landmark localization, is key for many applications such as facial animation capture, human activity recognition, and human-computer interaction. In this paper, we propose to use a recurrent neural network (RNN) for joint estimatio…

Cited by 163PDFScholar
2016

Online Detection and Classification of Dynamic Hand Gestures With Recurrent 3D Convolutional Neural Network

CVPR 2016poster

Automatic detection and classification of dynamic hand gestures in real-world systems intended for human computer interaction is challenging as: 1) there is a large diversity in how people perform gestures, making detection and classification difficult; 2) the system must work online in order to avo…

Cited by 822PDFScholar