← Search

Haoye Dong

19 accepted papers

2025

CLIP-GS: Unifying Vision-Language Representation with 3D Gaussian Splatting

ICCV 2025poster

Recent works in 3D representation learning and multimodal pre-training have made remarkable progress. However, typically multimodal 3D models are only capable of handling point clouds. Compared to the emerging 3D representation technique, 3D Gaussian Splatting (3DGS), the spatially sparse point clou…

Cited by 0SourcePDFScholar
2025

DuCos: Duality Constrained Depth Super-Resolution via Foundation Model

ICCV 2025poster

We introduce DuCos, a novel depth super-resolution framework grounded in Lagrangian duality theory, offering a flexible integration of multiple constraints and reconstruction objectives to enhance accuracy and robustness. Our DuCos is the first to significantly improve generalization across diverse…

2025

Learnable Infinite Taylor Gaussian for Dynamic View Rendering

CVPR 2025poster

Capturing the temporal evolution of Gaussian properties such as position, rotation, and scale is a challenging task due to the vast number of time-varying parameters and the limited photometric data available, which generally results in convergence issues, making it difficult to find an optimal solu…

Cited by 0SourcePDFScholar
2025

MV-SSM: Multi-View State Space Modeling for 3D Human Pose Estimation

CVPR 2025poster

While significant progress has been made in single-view 3D human pose estimation, multi-view 3D human pose estimation remains challenging, particularly in terms of generalizing to new camera configurations. Existing attention-based transformers often struggle to accurately model the spatial arrangem…

2024

Generalizable Human Gaussians for Sparse View Synthesis

ECCV 2024poster

"Recent progress in neural rendering has brought forth pioneering methods, such as NeRF and Gaussian Splatting, which revolutionize view rendering across various domains like AR/VR, gaming, and content creation. While these methods excel at interpolating within the training data, the challenge of ge…

2024

Hamba: Single-view 3D Hand Reconstruction with Graph-guided Bi-Scanning Mamba

NeurIPS 2024poster

3D Hand reconstruction from a single RGB image is challenging due to the articulated motion, self-occlusion, and interaction with objects. Existing SOTA methods employ attention-based transformers to learn the 3D hand pose and shape, yet they do not fully achieve robust and accurate performance, pri…

2023

Coordinate Transformer: Achieving Single-stage Multi-person Mesh Recovery from Videos

ICCV 2023poster

Multi-person 3D mesh recovery from videos is a critical first step towards automatic perception of group behavior in virtual reality, physical therapy and beyond. However, existing approaches rely on multi-stage paradigms, where the person detection and tracking stages are performed in a multi-perso…

Cited by 5PDFcodeScholar
2023

GP-VTON: Towards General Purpose Virtual Try-On via Collaborative Local-Flow Global-Parsing Learning

CVPR 2023poster

Image-based Virtual Try-ON aims to transfer an in-shop garment onto a specific person. Existing methods employ a global warping module to model the anisotropic deformation for different garment parts, which fails to preserve the semantic information of different parts when receiving challenging inpu…

2023

Human MotionFormer: Transferring Human Motions with Vision Transformers

ICLR 2023poster

Human motion transfer aims to transfer motions from a target dynamic person to a source static one for motion synthesis. An accurate matching between the source person and the target motion in both large and subtle motion changes is vital for improving the transferred motion quality. In this paper,…

2023

XFormer: Fast and Accurate Monocular 3D Body Capture

IJCAI 2023poster

We present XFormer, a novel human mesh and motion capture method that achieves real-time performance on consumer CPUs given only monocular images as input. The proposed network architecture contains two branches: a keypoint branch that estimates 3D human mesh vertices given 2D keypoints, and an imag…

Cited by 2SourcePDFScholar
2021

M3D-VTON: A Monocular-to-3D Virtual Try-On Network

ICCV 2021poster

Virtual 3D try-on can provide an intuitive and realistic view for online shopping and has a huge potential commercial value. However, existing 3D virtual try-on methods mainly rely on annotated 3D human shapes and garment templates, which hinders their applications in practical scenarios. 2D virtual…

Cited by 77PDFcodeScholar
2021

Towards Scalable Unpaired Virtual Try-On via Patch-Routed Spatially-Adaptive GAN

NeurIPS 2021poster

Image-based virtual try-on is one of the most promising applications of human-centric image generation due to its tremendous real-world potential. Yet, as most try-on approaches fit in-shop garments onto a target person, they require the laborious and restrictive construction of a paired training da…

2020

Fashion Editing With Adversarial Parsing Learning

CVPR 2020poster

Interactive fashion image manipulation, which enables users to edit images with sketches and color strokes, is an interesting research problem with great application value. Existing works often treat it as a general inpainting task and do not fully leverage the semantic structural information in fas…

Cited by 92PDFScholar
2019

FW-GAN: Flow-Navigated Warping GAN for Video Virtual Try-On

ICCV 2019poster

Beyond current image-based virtual try-on systems that have attracted increasing attention, we move a step forward to developing a video virtual try-on system that precisely transfers clothes onto the person and generates visually realistic videos conditioned on arbitrary poses. Besides the challeng…

Cited by 125PDFScholar
2019

Towards Multi-Pose Guided Virtual Try-On Network

ICCV 2019poster

Virtual try-on systems under arbitrary human poses have significant application potential, yet also raise extensive challenges, such as self-occlusions, heavy misalignment among different poses, and complex clothes textures. Existing virtual try-on methods can only transfer clothes given a fixed hum…

Cited by 252PDFScholar
2018

Deep Generative Models with Learnable Knowledge Constraints

NeurIPS 2018poster

The broad set of deep generative models (DGMs) has achieved remarkable advances. However, it is often difficult to incorporate rich structured domain knowledge with the end-to-end DGMs. Posterior regularization (PR) offers a principled framework to impose structured constraints on probabilistic mode…

Cited by 99SourcePDFScholar
2018

Soft-Gated Warping-GAN for Pose-Guided Person Image Synthesis

NeurIPS 2018poster

Despite remarkable advances in image synthesis research, existing works often fail in manipulating images under the context of large geometric transformations. Synthesizing person images conditioned on arbitrary poses is one of the most representative examples where the generation quality largely re…

Cited by 205SourcePDFScholar