← Search

Xuecheng Nie

19 accepted papers

2026

All-in-One Slider for Attribute Manipulation in Diffusion Models

CVPR 2026

Text-to-image (T2I) diffusion models have made significant strides in generating high-quality images. However, progressively manipulating certain attributes of generated images to meet the desired user expectations remains challenging, particularly for content with rich details, such as human faces.

Cited by 0SourcecodeScholar
2025

CamPoint: Boosting Point Cloud Segmentation with Virtual Camera

CVPR 2025poster

Local features aggregation and global information perception are the fundamental to point cloud segmentation. However, existing works often fall short in effectively identifying semantic relevant neighbors and face challenges in endowing each point with high-level information. Here, we propose CamPo…

Cited by 0SourcePDFScholar
2025

StyO: Stylize Your Face in Only One-Shot

AAAI 2025technical

This paper focuses on face stylization with a single artistic target. Existing works for this task often fail to retain the source content while achieving geometry variation. Here, we present a novel StyO model, i.e., Stylize the face in only One-shot, to solve the above problem. In particular, StyO…

Cited by 9SourcePDFScholar
2024

BlazeBVD: Make Scale-Time Equalization Great Again for Blind Video Deflickering

ECCV 2024poster

"Developing blind video deflickering (BVD) algorithms to enhance video temporal consistency, is gaining importance amid the flourish of image processing and video generation. However, the intricate nature of video data complicates the training of deep learning methods, leading to high resource consu…

Cited by 0SourcePDFScholar
2024

Customize your NeRF: Adaptive Source Driven 3D Scene Editing via Local-Global Iterative Training

CVPR 2024poster

In this paper we target the adaptive source driven 3D scene editing task by proposing a CustomNeRF model that unifies a text description or a reference image as the editing prompt. However obtaining desired editing results conformed with the editing prompt is nontrivial since there exist two signifi…

Cited by 14SourcePDFScholar
2023

Towards Consistent Video Editing with Text-to-Image Diffusion Models

NeurIPS 2023poster

Existing works have advanced Text-to-Image (TTI) diffusion models for video editing in a one-shot learning manner. Despite their low requirements of data and computation, these methods might produce results of unsatisfied consistency with text prompt as well as temporal sequence, limiting their appl…

Cited by 32SourcePDFScholar
2022

Distribution-Aware Single-Stage Models for Multi-Person 3D Pose Estimation

CVPR 2022poster

In this paper, we present a novel Distribution-Aware Single-stage (DAS) model for tackling the challenging multi-person 3D pose estimation problem. Different from existing top-down and bottom-up methods, the proposed DAS model simultaneously localizes person positions and their corresponding body jo…

Cited by 50PDFScholar
2022

Single-Stage Is Enough: Multi-Person Absolute 3D Pose Estimation

CVPR 2022poster

The existing multi-person absolute 3D pose estimation methods are mainly based on two-stage paradigm, i.e., top-down or bottom-up, leading to redundant pipelines with high computation cost. We argue that it is more desirable to simplify such two-stage paradigm to a single-stage one to promote both e…

Cited by 56PDFScholar
2020

Adversarial Self-Supervised Learning for Semi-Supervised 3D Action Recognition

ECCV 2020poster

We consider the problem of semi-supervised 3D action recognition which has been rarely explored before. Its major challenge lies in how to effectively learn motion representations from unlabeled data. Self-supervised learning (SSL) has been proved very effective at learning representations from unla…

Cited by 83SourcePDFScholar
2020

Inference Stage Optimization for Cross-scenario 3D Human Pose Estimation

NeurIPS 2020poster

Existing 3D human pose estimation models suffer performance drop when applying to new scenarios with unseen poses due to their limited generalizability. In this work, we propose a novel framework, Inference Stage Optimization (ISO), for improving the generalizability of 3D pose models when source an…

Cited by 55SourcePDFScholar
2019

Dynamic Kernel Distillation for Efficient Pose Estimation in Videos

ICCV 2019poster

Existing video-based human pose estimation methods extensively apply large networks onto every frame in the video to localize body joints, which suffer high computational cost and hardly meet the low-latency requirement in realistic applications. To address this issue, we propose a novel Dynamic Ker…

Cited by 92PDFcodeScholar
2018

Pose Partition Networks for Multi-Person Pose Estimation

ECCV 2018poster

This paper proposes a novel Pose Partition Network (PPN) to address the challenging multi-person pose estimation problem. The proposed PPN is favorably featured by low complexity and high accuracy of joint detection and partition. In particular, PPN performs dense regressions from global joint candi…

Cited by 102SourcePDFScholar
2017

Recurrent 3D-2D Dual Learning for Large-Pose Facial Landmark Detection

ICCV 2017poster

Despite remarkable progress of face analysis techniques, detecting landmarks on large-pose faces is still difficult due to self-occlusion, subtle landmark difference and incomplete information. To address these challenging issues, we introduce a novel recurrent 3D-2D dual learning model that alterna…

Cited by 64PDFScholar