← Search

Zilong Dong

26 accepted papers

2026

Head-wise Adaptive Rotary Positional Encoding for Fine-Grained Image Generation

CVPR 2026

Transformers rely on explicit positional encoding to model structure in data. WhileRotary Position Embedding (RoPE) excels in 1D domains, its application to image generation reveals significant limitations such as fine-grained spatial relationmodeling, color cues, and object counting. This paper ide

Cited by 0SourcecodeScholar
2026

Large Depth Completion Model from Sparse Observations

ICLR 2026poster

This work presents the Large Depth Completion Model (LDCM), a simple, effective, and robust framework for single-view metric depth estimation with sparse observations. Without relying on complex architectural designs, LDCM generates metric-accurate dense depth maps in one large transformer. It outpe…

Cited by 0SourceScholar
2026

Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers

ICML 2026poster

Multimodal Diffusion Transformers (MMDiTs) for text-to-image generation maintain separate text and image branches, with bidirectional information flow between text tokens and visual latents throughout denoising. In this setting, we observe a prompt forgetting phenomenon: the semantics of the prompt …

Cited by 0SourceScholar
2025

AniGS: Animatable Gaussian Avatar from a Single Image with Inconsistent Gaussian Reconstruction

CVPR 2025poster

Generating animatable human avatars from a single image is essential for various digital human modeling applications. Existing 3D reconstruction methods often struggle to capture fine details in animatable models, while generative approaches for controllable animation, though avoiding explicit 3D mo…

2025

Dirichlet-Constrained Variational Codebook Learning for Temporally Coherent Video Face Restoration

ICCV 2025poster

Video face restoration faces a critical challenge in maintaining temporal consistency while recovering fine facial details from degraded inputs. This paper presents a novel approach that extends Vector-Quantized Variational Autoencoders (VQ-VAEs), pretrained on static high-quality portraits, into a…

2025

HyPlaneHead: Rethinking Tri-plane-like Representations in Full-Head Image Synthesis

NeurIPS 2025poster

Tri-plane-like representations have been widely adopted in 3D-aware GANs for head image synthesis and other 3D object/scene modeling tasks due to their efficiency. However, querying features via Cartesian coordinate projection often leads to feature entanglement, which results in mirroring artifacts…

Cited by 0SourceScholar
2025

LHM: Large Animatable Human Reconstruction Model for Single Image to 3D in Seconds

ICCV 2025poster

Animatable 3D human reconstruction from a single image is a challenging problem due to the ambiguity in decoupling geometry, appearance, and deformation. Recent advances in 3D human reconstruction mainly focus on static human modeling, and the reliance of using synthetic 3D scans for training limits…

2025

LaMP: Language-Motion Pretraining for Motion Generation, Retrieval, and Captioning

ICLR 2025poster

Language plays a vital role in the realm of human motion. Existing methods have largely depended on CLIP text embeddings for motion generation, yet they fall short in effectively aligning language and motion due to CLIP’s pretraining on static image-text pairs. This work introduces LaMP, a novel Lan…

2025

Motions as Queries: One-Stage Multi-Person Holistic Human Motion Capture

CVPR 2025poster

Existing methods for capturing multi-person holistic human motions from a monocular video usually involve integrating the detector, the tracker, and the human pose & shape estimator into a cascaded system. Differently, we develop a one-stage multi-person holistic human motion capture system, which 1…

2024

Champ: Controllable and Consistent Human Image Animation with 3D Parametric Guidance

ECCV 2024poster

"In this study, we introduce a methodology for human image animation by leveraging a 3D human parametric model within a latent diffusion framework to enhance shape alignment and motion guidance in current human generative techniques. The methodology utilizes the SMPL(Skinned Multi-Person Linear) mod…

2024

GIC: Gaussian-Informed Continuum for Physical Property Identification and Simulation

NeurIPS 2024oral

This paper studies the problem of estimating physical properties (system identification) through visual observations. To facilitate geometry-aware guidance in physical property estimation, we introduce a novel hybrid framework that leverages 3D Gaussian representation to not only capture explicit sh…

2024

GPLD3D: Latent Diffusion of 3D Shape Generative Models by Enforcing Geometric and Physical Priors

CVPR 2024poster

State-of-the-art man-made shape generative models usually adopt established generative models under a suitable implicit shape representation. A common theme is to perform distribution alignment which does not explicitly model important shape priors. As a result many synthetic shapes are not connecte…

Cited by 8SourcePDFScholar
2024

High-Fidelity 3D Textured Shapes Generation by Sparse Encoding and Adversarial Decoding

ECCV 2024poster

"3D vision is inherently characterized by sparse spatial structures, which propels the necessity for an efficient paradigm tailored to 3D generation. Another discrepancy is the amount of training data, which undeniably affects generalization if we only use limited 3D data. To solve these, we design…

2024

High-Fidelity and Transferable NeRF Editing by Frequency Decomposition

ECCV 2024poster

"This paper enables high-fidelity, transferable NeRF editing by frequency decomposition. Recent NeRF editing pipelines lift 2D stylization results to 3D scenes while suffering from blurry results, and fail to capture detailed structures caused by the inconsistency between 2D editings. Our critical i…

2024

IPoD: Implicit Field Learning with Point Diffusion for Generalizable 3D Object Reconstruction from Single RGB-D Images

CVPR 2024highlight

Generalizable 3D object reconstruction from single-view RGB-D images remains a challenging task particularly with real-world data. Current state-of-the-art methods develop Transformer-based implicit field learning necessitating an intensive learning paradigm that requires dense query-supervision uni…

2024

MoGenTS: Motion Generation based on Spatial-Temporal Joint Modeling

NeurIPS 2024poster

Motion generation from discrete quantization offers many advantages over continuous regression, but at the cost of inevitable approximation errors. Previous methods usually quantize the entire body pose into one code, which not only faces the difficulty in encoding all joints within one vector but a…

Cited by 2SourcePDFScholar
2024

Open-Vocabulary Category-Level Object Pose and Size Estimation

RA-L 2024

This letter studies a new open-set problem, the open-vocabulary category-level object pose and size estimation. Given human text descriptions of arbitrary novel object categories, the robot agent seeks to predict the position, orientation, and size of the target object in the observed scene image. T

Cited by 11SourceScholar
2024

PanoContext-Former: Panoramic Total Scene Understanding with a Transformer

CVPR 2024poster

Panoramic images enable deeper understanding and more holistic perception of 360 surrounding environment which can naturally encode enriched scene context information compared to standard perspective image. Previous work has made lots of effort to solve the scene understanding task in a hybrid solut…

Cited by 11SourcePDFScholar
2024

RichDreamer: A Generalizable Normal-Depth Diffusion Model for Detail Richness in Text-to-3D

CVPR 2024highlight

Lifting 2D diffusion for 3D generation is a challenging problem due to the lack of geometric prior and the complex entanglement of materials and lighting in natural images. Existing methods have shown promise by first creating the geometry through score-distillation sampling (SDS) applied to rendere…

2023

$\mathcal {S}{2}$Net: Accurate Panorama Depth Estimation on Spherical Surface

RA-L 2023

Monocular depth estimation is an ambiguous problem, thus global structural cues play an important role in current data-driven single-view depth estimation methods. Panorama images capture the complete spatial information of their surroundings utilizing the equirectangular projection which introduces

Cited by 10SourceScholar
2023

DENSE RGB SLAM WITH NEURAL IMPLICIT MAPS

ICLR 2023poster

There is an emerging trend of using neural implicit functions for map representation in Simultaneous Localization and Mapping (SLAM). Some pioneer works have achieved encouraging results on RGB-D SLAM. In this paper, we present a dense RGB SLAM method with neural implicit map representation. To reac…

2023

DRO: Deep Recurrent Optimizer for Video to Depth

RA-L 2023

There are increasing interests of studying the video-to-depth (V2D) problem with machine learning techniques. While earlier methods directly learn a mapping from images to depth maps and camera poses, more recent works enforce multi-view geometry constraints through optimization embedded in the lear

Cited by 21SourcecodeScholar
2021

Single-Shot is Enough: Panoramic Infrastructure Based Calibration of Multiple Cameras and 3D LiDARs

IROS 2021poster

The integration of multiple cameras and 3D Li-DARs has become basic configuration of augmented reality devices, robotics, and autonomous vehicles. The calibration of multi-modal sensors is crucial for a system to properly function, but it remains tedious and impractical for mass production. Moreover…

Cited by 27SourceScholar