← Search

Zehuan Huang

10 accepted papers

2026

InterMoE: Individual-Specific 3D Human Interaction Generation via Dynamic Temporal-Selective MoE

AAAI 2026technical

Generating high-quality human interactions holds significant value for applications like virtual reality and robotics. However, existing methods often fail to preserve unique individual characteristics or fully adhere to textual descriptions. To address these challenges, we introduce InterMoE, a nov

Cited by 0SourcePDFScholar
2026

PhysForge: Generating Physics-Grounded 3D Assets for Interactive Virtual World

ICML 2026poster

Synthesizing physics-grounded 3D assets is a critical bottleneck for interactive virtual worlds and embodied AI. Existing methods predominantly focus on static geometry, overlooking the functional properties essential for interaction. We propose that interactive asset generation must be rooted in fu…

Cited by 0SourceScholar
2026

Repurposing 3D Generative Model for Autoregressive Layout Generation

CVPR 2026

We introduce LaviGen, a framework that repurposes 3D generative models for 3D layout generation. Unlike previous methods that infer object layouts from textual descriptions, LaviGen operates directly in the native 3D space, formulating layout generation as an autoregressive process that explicitly m

Cited by 0SourcecodeScholar
2026

Stereo World Model: Camera-Guided Stereo Video Generation

CVPR 2026

We present StereoWorld, a camera-conditioned stereo world model that jointly learns appearance and binocular geometry for end-to-end stereo video generation.Unlike monocular RGB or RGBD approaches, StereoWorld operates exclusively within the RGB modality, while simultaneously grounding geometry dire

Cited by 0SourcecodeScholar
2025

MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation

CVPR 2025poster

This paper introduces MIDI, a novel paradigm for compositional 3D scene generation from a single image. Unlike existing methods that rely on reconstruction or retrieval techniques or recent approaches that employ multi-stage object-by-object generation, MIDI extends pre-trained image-to-3D object ge…

Cited by 1SourcePDFScholar
2025

MV-Adapter: Multi-View Consistent Image Generation Made Easy

ICCV 2025poster

Existing multi-view image generation methods often make invasive modifications to pre-trained text-to-image (T2I) models and require full fine-tuning, leading to high computational costs and degradation in image quality due to scarce high-quality 3D data. This paper introduces MV-Adapter, an efficie…

Cited by 0SourcePDFScholar
2025

Ouroboros3D: Image-to-3D Generation via 3D-aware Recursive Diffusion

CVPR 2025poster

Existing image-to-3D creation methods typically split the task into two individual stage: multi-view image generation and 3D reconstruction, leading to two main limitations: (1) In multi-view generation stage, the multi-view generated images present a challenge to preserving 3D consistency;; (2) In…

2024

EpiDiff: Enhancing Multi-View Synthesis via Localized Epipolar-Constrained Diffusion

CVPR 2024poster

Generating multiview images from a single view facilitates the rapid generation of a 3D mesh conditioned on a single image. Recent methods that introduce 3D global representation into diffusion models have shown the potential to generate consistent multiviews but they have reduced generation speed a…

2024

Text to Layer-wise 3D Clothed Human Generation

ECCV 2024poster

"This paper addresses the task of 3D clothed human generation from textural descriptions. Previous works usually encode the human body and clothes as a holistic model and generate the whole model in a single-stage optimization, which makes them struggle for clothing editing and meanwhile lose fine-g…

Cited by 12SourcePDFScholar