← Search

Zhiyang Dou

29 accepted papers

2026

GeoPT: Scaling Physics Simulation via Lifted Geometric Pre-Training

ICML 2026poster

Neural simulators promise efficient surrogates for physics simulation, but scaling them is bottlenecked by the prohibitive cost of generating high-fidelity training data. Pre-training on abundant off-the-shelf geometries offers a natural alternative, yet faces a fundamental gap: supervision on stati…

Cited by 0SourceScholar
2026

MeshMosaic: Scaling Artist Mesh Generation via Local-to-Global Assembly

CVPR 2026

Scaling artist-designed meshes to high triangle numbers remains challenging for autoregressive generative models. Existing transformer-based methods suffer from long-sequence bottlenecks and limited quantization resolution, primarily due to the large number of tokens required and constrained quantiz

Cited by 0SourcecodeScholar
2026

NeuralActuator: Neural Actuation Modeling for Robot Dynamics and External Force Perception

RSS 2026poster

Differentiable simulators have advanced policy learning and model-based control across diverse robotic tasks. To date, actuator dynamics remain underexplored and are a major source of sim-to-real error, especially on low-cost platforms where the linear current–torque model τ = K_tI breaks down under…

Cited by 0SourceScholar
2026

PartSAM: A Scalable Promptable Part Segmentation Model Trained on Native 3D Data

ICLR 2026poster

Segmenting 3D objects into parts is a long-standing challenge in computer vision. To overcome taxonomy constraints and generalize to unseen 3D objects, recent works turn to open-world part segmentation. These approaches typically transfer supervision from 2D foundation models, such as SAM, by liftin…

Cited by 0SourcecodeScholar
2026

PointNSP: Autoregressive 3D Point Cloud Generation with Next-Scale Level-of-Detail Prediction

CVPR 2026

Autoregressive point cloud generation has long lagged behind diffusion-based approaches in quality. The performance gap stems from the fact that autoregressive models impose an artificial ordering on inherently unordered point sets, forcing shape generation to proceed as a sequence of local predicti

Cited by 4SourcecodeScholar
2026

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation

ICLR 2026poster

Generating high-quality videos from complex temporal descriptions, which refer to prompts containing multiple sequential actions, remains a significant challenge. Existing methods are constrained by an inherent trade-off: using multiple short prompts fed sequentially into the model improves action f…

Cited by 0SourcecodeScholar
2026

Text2Interact: High-Fidelity and Diverse Text-to-Two-Person Interaction Generation

ICLR 2026poster

Generating realistic and diverse human-human interactions from text is a crucial yet challenging task in computer vision, graphics, and robotics. Despite recent advances, existing methods have two key limitations. First, two-person interaction synthesis is highly complex, simultaneously requiring in…

Cited by 0SourcecodeScholar
2025

Align3R: Aligned Monocular Depth Estimation for Dynamic Videos

CVPR 2025highlight

Recent developments in monocular depth estimation methods enable high-quality depth estimation of single-view images but fail to estimate consistent video depth across different frames. Recent works address this problem by applying a video diffusion model to generate video depth conditioned on the i…

Cited by 14SourcePDFScholar
2025

Boosting Segment Anything Model Towards Open-Vocabulary Learning

AAAI 2025technical

The recent Segment Anything Model (SAM) has emerged as a new paradigmatic vision foundation model, showcasing potent zero-shot generalization and flexible prompting. Despite SAM finding applications and adaptations in various domains, its primary limitation lies in the inability to grasp object sema…

2025

CityAnchor: City-scale 3D Visual Grounding with Multi-modality LLMs

ICLR 2025poster

In this paper, we present a 3D visual grounding method called CityAnchor for localizing an urban object in a city-scale point cloud. Recent developments in multiview reconstruction enable us to reconstruct city-scale point clouds but how to conduct visual grounding on such a large-scale urban point…

Cited by 0SourcePDFScholar
2025

CoDA: Coordinated Diffusion Noise Optimization for Whole-Body Manipulation of Articulated Objects

NeurIPS 2025poster

Synthesizing whole-body manipulation of articulated objects, including body motion, hand motion, and object motion, is a critical yet challenging task with broad applications in virtual humans and robotics. The core challenges are twofold. First, achieving realistic whole-body motion requires tight…

Cited by 0SourceScholar
2025

DICE: End-to-end Deformation Capture of Hand-Face Interactions from a Single Image

ICLR 2025poster

Reconstructing 3D hand-face interactions with deformations from a single image is a challenging yet crucial task with broad applications in AR, VR, and gaming. The challenges stem from self-occlusions during single-view hand-face interactions, diverse spatial relationships between hands and face, co…

2025

Go to Zero: Towards Zero-shot Motion Generation with Million-scale Data

ICCV 2025poster

Generating diverse and natural human motion sequences based on textual descriptions constitutes a fundamental and challenging research area within the domains of computer vision, graphics, and robotics. Despite significant advancements in this field, current methodologies often face challenges regar…

2025

MVTokenFlow: High-quality 4D Content Generation using Multiview Token Flow

ICLR 2025poster

In this paper, we present MVTokenFlow for high-quality 4D content creation from monocular videos. Recent advancements in generative models such as video diffusion models and multiview diffusion models enable us to create videos or 3D models. However, extending these generative models for dynamic 4D…

2025

PhysCtrl: Generative Physics for Controllable and Physics-Grounded Video Generation

NeurIPS 2025poster

Existing video generation models excel at producing photo-realistic videos from text or images, but often lack physical plausibility and 3D controllability. To overcome these limitations, we introduce PhysCtrl, a novel framework for physics-grounded image-to-video generation with physical parameters…

Cited by 0SourceScholar
2025

SIMS: Simulating Stylized Human-Scene Interactions with Retrieval-Augmented Script Generation

ICCV 2025poster

Simulating stylized human-scene interactions (HSI) in physical environments is a challenging yet fascinating task. Prior works emphasize long-term execution but fall short in achieving both diverse style and physical plausibility. To tackle this challenge, we introduce a novel hierarchical framework…

Cited by 0SourcePDFScholar
2025

ScaMo: Exploring the Scaling Law in Autoregressive Motion Generation Model

CVPR 2025poster

The scaling law has been validated in various domains, such as natural language processing (NLP) and massive computer vision tasks; however, its application to motion generation remains largely unexplored. In this paper, we introduce a scalable motion generation framework that includes the motion to…

Cited by 6SourcePDFScholar
2025

SyncHuman: Synchronizing 2D and 3D Generative Models for Single-view Human Reconstruction

NeurIPS 2025poster

Photorealistic 3D full-body human reconstruction from a single image is a critical yet challenging task for applications in films and video games due to inherent ambiguities and severe self-occlusions. While recent approaches leverage SMPL estimation and SMPL-conditioned image generative models to h…

Cited by 0SourceScholar
2025

TokenHSI: Unified Synthesis of Physical Human-Scene Interactions through Task Tokenization

CVPR 2025poster

Synthesizing diverse and physically plausible Human-Scene Interactions (HSI) is pivotal for both computer animation and embodied AI. Despite encouraging progress, current methods mainly focus on developing separate controllers, each specialized for a specific interaction task. This significantly hin…

Cited by 3SourcePDFScholar
2025

TrackingWorld: World-centric Monocular 3D Tracking of Almost All Pixels

NeurIPS 2025poster

Monocular 3D tracking aims to capture the long-term motion of pixels in 3D space from a single monocular video and has witnessed rapid progress in recent years. However, we argue that the existing monocular 3D tracking methods still fall short in separating the camera motion from foreground dynamic…

Cited by 0SourceScholar
2025

Vid2Sim: Generalizable, Video-based Reconstruction of Appearance, Geometry and Physics for Mesh-free Simulation

CVPR 2025poster

Faithfully reconstructing textured shapes and physical properties from videos presents an intriguing yet challenging problem. Significant efforts have been dedicated to advancing such a system identification problem in this area. Previous methods often rely on heavy optimization pipelines with a dif…

Cited by 0SourcePDFScholar
2025

🎧MOSPA: Human Motion Generation Driven by Spatial Audio

NeurIPS 2025spotlight

Enabling virtual humans to dynamically and realistically respond to diverse auditory stimuli remains a key challenge in character animation, demanding the integration of perceptual modeling and motion synthesis. Despite its significance, this task remains largely unexplored. Most previous works have…

Cited by 0SourcecodeScholar
2024

AutoPSV: Automated Process-Supervised Verifier

NeurIPS 2024poster

In this work, we propose a novel method named \textbf{Auto}mated \textbf{P}rocess-\textbf{S}upervised \textbf{V}erifier (\textbf{\textsc{AutoPSV}}) to enhance the reasoning capabilities of large language models (LLMs) by automatically annotating the reasoning steps. \textsc{AutoPSV} begins by traini…

2024

Disentangled Clothed Avatar Generation from Text Descriptions

ECCV 2024poster

"In this paper, we introduce a novel text-to-avatar generation method that separately generates the human body and the clothes and allows high-quality animation on the generated avatar. While recent advancements in text-to-avatar generation have yielded diverse human avatars from text prompts, these…

Cited by 24SourcePDFScholar
2024

Surf-D: Generating High-Quality Surfaces of Arbitrary Topologies Using Diffusion Models

ECCV 2024poster

"We present Surf-D, a novel method for generating high-quality 3D shapes as Surfaces with arbitrary topologies using Diffusion models. Previous methods explored shape generation with different representations and they suffer from limited topologies and poor geometry details. To generate high-quality…

Cited by 1SourcePDFScholar
2024

TLControl: Trajectory and Language Control for Human Motion Synthesis

ECCV 2024poster

"Controllable human motion synthesis is essential for applications in AR/VR, gaming and embodied AI. Existing methods often focus solely on either language or full trajectory control, lacking precision in synthesizing motions aligned with user-specified trajectories, especially for multi-joint contr…

Cited by 49SourcePDFScholar
2024

Wonder3D: Single Image to 3D using Cross-Domain Diffusion

CVPR 2024highlight

In this work we introduce Wonder3D a novel method for generating high-fidelity textured meshes from single-view images with remarkable efficiency. Recent methods based on the Score Distillation Sampling (SDS) loss methods have shown the potential to recover 3D geometry from 2D diffusion priors but t…

Cited by 414SourcePDFScholar
2023

TORE: Token Reduction for Efficient Human Mesh Recovery with Transformer

ICCV 2023poster

In this paper, we introduce a set of simple yet effective TOken REduction (TORE) strategies for Transformer-based Human Mesh Recovery from monocular images. Current SOTA performance is achieved by Transformer-based structures. However, they suffer from high model complexity and computation cost caus…

Cited by 51PDFcodeScholar