← Search

Yuhao Cheng

14 accepted papers

2026

BAMI: Training-Free Bias Mitigation in GUI Grounding

CVPR 2026

GUI grounding is a critical capability for enabling GUI agents to execute tasks such as clicking and dragging. However, in complex scenarios like the ScreenSpot-Pro benchmark, existing models often suffer from suboptimal performance. Utilizing the proposed Masked Prediction Distribution (MPD) attrib

Cited by 0SourcecodeScholar
2026

MoRE: 3D Visual Geometry Reconstruction Meets Mixture-of-Experts

CVPR 2026

Recent advances in language and vision have demonstrated that scaling up model capacity consistently improves performance across diverse tasks.In 3D visual geometry reconstruction, large-scale training has likewise proven effective for learning versatile representations.However, further scaling of 3

Cited by 0SourcecodeScholar
2026

ProPhy: Progressive Physical Alignment for Dynamic World Simulation

CVPR 2026

Recent advances in video generation have shown remarkable potential for constructing world simulators. However, current models still struggle to produce physically consistent results, particularly when handling large-scale or complex dynamics. This limitation arises primarily because existing approa

Cited by 0SourceScholar
2025

GDrag:Towards General-Purpose Interactive Editing with Anti-ambiguity Point Diffusion

ICLR 2025poster

Recent interactive point-based image manipulation methods have gained considerable attention for being user-friendly. However, these methods still face two types of ambiguity issues that can lead to unsatisfactory outcomes, namely, intention ambiguity which misinterprets the purposes of users, and c…

Cited by 1SourcePDFScholar
2025

Multimodal Latent Diffusion Model for Complex Sewing Pattern Generation

ICCV 2025poster

Generating sewing patterns in garment design is receiving increasing attention due to its CG-friendly and flexible-editing nature. Previous sewing pattern generation methods have been able to produce exquisite clothing, but struggle to design complex garments with detailed control. To address these…

Cited by 0SourcePDFScholar
2025

S^3-Face: SSS-Compliant Facial Reflectance Estimation via Diffusion Priors

CVPR 2025poster

Recent 3D face reconstruction methods have made remarkable advancements, yet achieving high-quality facial reflectance from monocular input remains challenging. Existing methods rely on the light-stage captured data to learn facial reflectance models. However, limited subject diversity in these data…

Cited by 0SourcePDFScholar
2025

Towards High-fidelity 3D Talking Avatar with Personalized Dynamic Texture

CVPR 2025poster

Significant progress has been made for speech-driven 3D face animation, but most works focus on learning the motion of mesh/geometry, ignoring the impact of dynamic texture. In this work, we reveal that dynamic texture plays a key role in rendering high-fidelity talking avatars, and introduce a high…

2024

3D-Aware Face Editing via Warping-Guided Latent Direction Learning

CVPR 2024poster

3D facial editing a longstanding task in computer vision with broad applications is expected to fast and intuitively manipulate any face from arbitrary viewpoints following the user's will. Existing works have limitations in terms of intuitiveness generalization and efficiency. To overcome these cha…

2024

GarmentAligner: Text-to-Garment Generation via Retrieval-augmented Multi-level Corrections

ECCV 2024poster

"General text-to-image models bring revolutionary innovation to the fields of arts, design, and media. However, when applied to garment generation, even the state-of-the-art text-to-image models suffer from fine-grained semantic misalignment, particularly concerning the quantity, position, and inter…

Cited by 5SourcePDFScholar
2024

Learning Interaction-aware 3D Gaussian Splatting for One-shot Hand Avatars

NeurIPS 2024poster

In this paper, we propose to create animatable avatars for interacting hands with 3D Gaussian Splatting (GS) and single-image inputs. Existing GS-based methods designed for single subjects often yield unsatisfactory results due to limited input views, various hand poses, and occlusions. To address t…

2024

Monocular Identity-Conditioned Facial Reflectance Reconstruction

CVPR 2024poster

Recent 3D face reconstruction methods have made remarkable advancements yet there remain huge challenges in monocular high-quality facial reflectance reconstruction. Existing methods rely on a large amount of light-stage captured data to learn facial reflectance models. However the lack of subject d…

Cited by 3SourcePDFScholar
2024

Topo4D: Topology-Preserving Gaussian Splatting for High-Fidelity 4D Head Capture

ECCV 2024poster

"Recent significant advances in high-quality face reconstruction have been made, but challenges remain in 4D face asset reconstruction. 4D head capture aims to generate dynamic topological meshes and corresponding texture maps from videos, which is widely utilized in movies and games for its ability…

2023

GANHead: Towards Generative Animatable Neural Head Avatars

CVPR 2023poster

To bring digital avatars into people's lives, it is highly demanded to efficiently generate complete, realistic, and animatable head avatars. This task is challenging, and it is difficult for existing methods to satisfy all the requirements at once. To achieve these goals, we propose GANHead (Genera…

Cited by 20SourcePDFScholar
2022

CageNeRF: Cage-based Neural Radiance Field for Generalized 3D Deformation and Animation

NeurIPS 2022accept

While implicit representations have achieved high-fidelity results in 3D rendering, it remains challenging to deforming and animating the implicit field. Existing works typically leverage data-dependent models as deformation priors, such as SMPL for human body animation. However, this dependency on…

Cited by 59SourcePDFScholar