← Search

Chenguo Lin

10 accepted papers

2026

Diff4Splat: Repurposing Video Diffusion Models for Dynamic Scene Generation

CVPR 2026

We introduce Diff4Splat, a feed-forward framework for dynamic scene generation from a single image. Our method synergizes the powerful generative priors of video diffusion models with geometric and motion constraints learned from a large-scale 4D dataset. Given a single image, a camera trajectory, a

Cited by 0SourcecodeScholar
2026

ID-Crafter: VLM-Grounded Online RL for Compositional Multi-Subject Video Generation

CVPR 2026

Significant progress has been achieved in high-fidelity video synthesis, yet current paradigms often fall short in effectively integrating identity information from multiple subjects. This leads to semantic conflicts and suboptimal performance in preserving identities and interactions, limiting cont

Cited by 0SourceScholar
2026

MoVieS: Motion-Aware 4D Dynamic View Synthesis in One Second

CVPR 2026

We present MoVieS, a Motion-aware View Synthesis model that reconstructs 4D dynamic scenes from monocular videos in one second. It represents dynamic 3D scenes with pixel-aligned Gaussian primitives and explicitly supervises their time-varying motions. This allows, for the first time, the unified mo

Cited by 0SourcecodeScholar
2025

DiffSplat: Repurposing Image Diffusion Models for Scalable Gaussian Splat Generation

ICLR 2025poster

Recent advancements in 3D content generation from text or a single image struggle with limited high-quality 3D datasets and inconsistency from 2D multi-view generation. We introduce DiffSplat, a novel 3D generative framework that natively generates 3D Gaussian splats by taming large-scale text-to-im…

Cited by 5SourcePDFScholar
2025

DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling

NeurIPS 2025poster

Understanding the dynamic physical world, characterized by its evolving 3D structure, real-world motion, and semantic content with textual descriptions, is crucial for human-agent interaction and enables embodied agents to perceive and act within real environments with human‑like capabilities. Howev…

Cited by 0SourceScholar
2025

OmniPhysGS: 3D Constitutive Gaussians for General Physics-Based Dynamics Generation

ICLR 2025poster

Recently, significant advancements have been made in the reconstruction and generation of 3D assets, including static cases and those with physical interactions. To recover the physical properties of 3D assets, existing methods typically assume that all materials belong to a specific predefined cate…

Cited by 2SourcePDFScholar
2025

PartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion Transformers

NeurIPS 2025poster

We introduce PartCrafter, the first structured 3D generative model that jointly synthesizes multiple semantically meaningful and geometrically distinct 3D meshes from a single RGB image. Unlike existing methods that either produce monolithic 3D shapes or follow two-stage pipelines, i.e. first segmen…

Cited by 0SourceScholar
2024

HumanSplat: Generalizable Single-Image Human Gaussian Splatting with Structure Priors

NeurIPS 2024poster

Despite recent advancements in high-fidelity human reconstruction techniques, the requirements for densely captured images or time-consuming per-instance optimization significantly hinder their applications in broader scenarios. To tackle these issues, we present **HumanSplat**, which predicts the 3…

2024

InstructScene: Instruction-Driven 3D Indoor Scene Synthesis with Semantic Graph Prior

ICLR 2024spotlight

Comprehending natural language instructions is a charming property for 3D indoor scene synthesis systems. Existing methods directly model object joint distributions and express object relations implicitly within a scene, thereby hindering the controllability of generation. We introduce InstructScene…

2021

A Survey on Universal Adversarial Attack

IJCAI 2021poster

The intriguing phenomenon of adversarial examples has attracted significant attention in machine learning and what might be more surprising to the community is the existence of universal adversarial perturbations (UAPs), i.e. a single perturbation to fool the target DNN for most images. With the foc…

Cited by 111SourcePDFScholar