← Search

Yipeng Qin

28 accepted papers

2026

3DGS-HPC: Distractor-free 3D Gaussian Splatting with Hybrid Patch-wise Classification

ICML 2026poster

3D Gaussian Splatting (3DGS) has demonstrated remarkable performance in novel view synthesis and 3D scene reconstruction, but its quality often degrades in real-world environments due to transient distractors, such as moving objects and varying shadows. Existing methods commonly rely on semantic cue…

Cited by 0SourceScholar
2026

Distinguishing Imitation Error from Intrinsic Motion Learning Difficulty

ICML 2026poster

Physics-based motion imitation is central to humanoid control, yet current evaluation metrics(e.g., MPJPE) only quantify imitation outcomes, not their underlying causes. This conflation obscures a critical diagnostic question: when imitation error occurs, does it stem from policy limitations or the …

Cited by 0SourceScholar
2026

Improving Sparse IMU-based Motion Capture with Motion Label Smoothing

AAAI 2026technical

Sparse Inertial Measurement Units (IMUs) based human motion capture has gained significant momentum, driven by the adaptation of fundamental AI tools such as recurrent neural networks (RNNs) and transformers that are tailored for temporal and spatial modeling. Despite these achievements, current res

Cited by 0SourcePDFScholar
2026

LoFA: Learning to Predict Personalized Prior for Fast Adaptation of Visual Generative Models

CVPR 2026

Personalizing visual generative models to meet specific user needs has gained increasing attention, yet current methods like Low-Rank Adaptation (LoRA) remain impractical due to their demand for task-specific data and lengthy optimization. While a few hypernetwork-based approaches attempt to predict

Cited by 0SourcecodeScholar
2026

LookasideVLN: Direction-Aware Aerial Vision-and-Language Navigation

CVPR 2026

Aerial Vision-and-Language Navigation (Aerial VLN) enables unmanned aerial vehicles (UAVs) to follow natural language instructions and navigate complex urban environments.While recent advances have achieved progress through large-scale memory graphs and lookahead path planning, they remain limited b

Cited by 0SourceScholar
2025

Hierarchically Controlled Deformable 3D Gaussians for Talking Head Synthesis

AAAI 2025technical

Audio-driven talking head synthesis is a critical task in digital human modeling. While recent advances using diffusion models and Neural Radiance Fields (NeRF) have improved visual quality, they often require substantial computational resources, limiting practical deployment. We present a novel fra…

Cited by 1SourcePDFScholar
2025

LLM-driven Multimodal and Multi-Identity Listening Head Generation

CVPR 2025poster

Generating natural listener responses in conversational scenarios is crucial for creating engaging digital humans and avatars. Recent work has shown that large language models (LLMs) can be effectively leveraged for this task, demonstrating remarkable capabilities in generating contextually appropri…

Cited by 0SourcePDFScholar
2025

ToF-IP: Time-of-Flight Enhanced Sparse Inertial Poser for Real-time Human Motion Capture

NeurIPS 2025poster

Sparse inertial measurement units (IMUs) provide a portable, low-cost solution for human motion tracking but struggle with error accumulation from drift and sensor noise when estimating joint position through time-based linear acceleration integration (i.e., indirect measurement). To address this,…

Cited by 0SourceScholar
2025

VTON 360: High-Fidelity Virtual Try-On from Any Viewing Direction

CVPR 2025poster

Virtual Try-On (VTON) is a transformative technology in e-commerce and fashion design, enabling realistic digital visualization of clothing on individuals. In this work, we propose VTON 360, a novel 3D VTON method that addresses the open challenge of achieving high-fidelity VTON that supports any-vi…

Cited by 1SourcePDFScholar
2024

Accurate and Steady Inertial Pose Estimation through Sequence Structure Learning and Modulation

NeurIPS 2024poster

Transformer models excel at capturing long-range dependencies in sequential data, but lack explicit mechanisms to leverage structural patterns inherent in fixed-length input sequences. In this paper, we propose a novel sequence structure learning and modulation approach that endows Transformers wit…

Cited by 1SourcePDFScholar
2024

Deep Generative Model based Rate-Distortion for Image Downscaling Assessment

CVPR 2024poster

In this paper we propose Image Downscaling Assessment by Rate-Distortion (IDA-RD) a novel measure to quantitatively evaluate image downscaling algorithms. In contrast to image-based methods that measure the quality of downscaled images ours is process-based that draws ideas from rate-distortion theo…

2024

Efficient Precision and Recall Metrics for Assessing Generative Models using Hubness-aware Sampling

ICML 2024spotlight

Despite impressive results, deep generative models require massive datasets for training, and as dataset size increases, effective evaluation metrics like precision and recall (P&R) become computationally infeasible on commodity hardware. In this paper, we address this challenge by proposing efficie…

2024

Loose Inertial Poser: Motion Capture with IMU-attached Loose-Wear Jacket

CVPR 2024poster

Existing wearable motion capture methods typically demand tight on-body fixation (often using straps) for reliable sensing limiting their application in everyday life. In this paper we introduce Loose Inertial Poser a novel motion capture solution with high wearing comfortableness by integrating fou…

2024

NeRF-HuGS: Improved Neural Radiance Fields in Non-static Scenes Using Heuristics-Guided Segmentation

CVPR 2024poster

Neural Radiance Field (NeRF) has been widely recognized for its excellence in novel view synthesis and 3D scene reconstruction. However their effectiveness is inherently tied to the assumption of static scenes rendering them susceptible to undesirable artifacts when confronted with transient distrac…

2024

PICTURE: PhotorealistIC virtual Try-on from UnconstRained dEsigns

CVPR 2024poster

In this paper we propose a novel virtual try-on from unconstrained designs (ucVTON) task to enable photorealistic synthesis of personalized composite clothing on input human images. Unlike prior arts constrained by specific input types our method allows flexible specification of style (text or image…

Cited by 8SourcePDFScholar
2024

SuDA: Support-based Domain Adaptation for Sim2Real Hinge Joint Tracking with Flexible Sensors

ICML 2024poster

Flexible sensors hold promise for human motion capture (MoCap), offering advantages such as wearability, privacy preservation, and minimal constraints on natural movement. However, existing flexible sensor-based MoCap methods rely on deep learning and necessitate large and diverse labeled datasets f…

Cited by 1SourcePDFScholar
2023

Feature Proliferation -- the "Cancer" in StyleGAN and its Treatments

ICCV 2023poster

Despite the success of StyleGAN in image synthesis, the images it synthesizes are not always perfect and the well-known truncation trick has become a standard post-processing technique for StyleGAN to synthesize high-quality images. Although effective, it has long been noted that the truncation tric…

Cited by 0PDFcodeScholar
2023

Parametric Implicit Face Representation for Audio-Driven Facial Reenactment

CVPR 2023poster

Audio-driven facial reenactment is a crucial technique that has a range of applications in film-making, virtual avatars and video conferences. Existing works either employ explicit intermediate face representations (e.g., 2D facial landmarks or 3D face models) or implicit ones (e.g., Neural Radiance…

Cited by 20SourcePDFScholar
2023

Self-Adaptive Motion Tracking against On-body Displacement of Flexible Sensors

NeurIPS 2023poster

Flexible sensors are promising for ubiquitous sensing of human status due to their flexibility and easy integration as wearable systems. However, on-body displacement of sensors is inevitable since the device cannot be firmly worn at a fixed position across different sessions. This displacement issu…

Cited by 6SourcePDFScholar
2022

Centrality and Consistency: Two-Stage Clean Samples Identification for Learning with Instance-Dependent Noisy Labels

ECCV 2022poster

"Deep models trained with noisy labels are prone to over-fitting and struggle in generalization. Most existing solutions are based on an ideal assumption that the label noise is class-conditional, i.e., instances of the same class share the same noise model, and are independent of features. While in…

2022

Exploring and Exploiting Hubness Priors for High-Quality GAN Latent Sampling

ICML 2022spotlight

Despite the extensive studies on Generative Adversarial Networks (GANs), how to reliably sample high-quality images from their latent spaces remains an under-explored topic. In this paper, we propose a novel GAN latent sampling method by exploring and exploiting the hubness priors of GAN latent dist…

2022

Multi-level Consistency Learning for Semi-supervised Domain Adaptation

IJCAI 2022poster

Semi-supervised domain adaptation (SSDA) aims to apply knowledge learned from a fully labeled source domain to a scarcely labeled target domain. In this paper, we propose a Multi-level Consistency Learning (MCL) framework for SSDA. Specifically, our MCL regularizes the consistency of different views…

2020

SEAN: Image Synthesis With Semantic Region-Adaptive Normalization

CVPR 2020oral

We propose semantic region-adaptive normalization (SEAN), a simple but effective building block for Generative Adversarial Networks conditioned on segmentation masks that describe the semantic regions in the desired output image. Using SEAN normalization, we can build a network architecture that can…

Cited by 731PDFcodeScholar