← Search

Sherwin Bahmani

7 accepted papers

2026

Lyra: Generative 3D Scene Reconstruction via Video Diffusion Model Self-Distillation

ICLR 2026poster

The ability to generate virtual environments is crucial for applications ranging from gaming to physical AI domains such as robotics, autonomous driving, and industrial AI. Current learning-based 3D reconstruction methods rely on the availability of captured real-world multi-view data, which is not…

Cited by 0SourcecodeScholar
2025

AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers

CVPR 2025poster

Numerous works have recently integrated 3D camera control into foundational text-to-video models, but the resulting camera control is often imprecise, and video generation quality suffers. In this work, we analyze camera motion from a first principles perspective, uncovering insights that enable pre…

Cited by 10SourcePDFScholar
2025

SG-I2V: Self-Guided Trajectory Control in Image-to-Video Generation

ICLR 2025poster

Methods for image-to-video generation have achieved impressive, photo-realistic quality. However, adjusting specific elements in generated videos, such as object motion or camera movement, is often a tedious process of trial and error, e.g., involving re-generating videos with different random seed…

2025

VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

ICLR 2025poster

Modern text-to-video synthesis models demonstrate coherent, photorealistic generation of complex videos from a text description. However, most existing models lack fine-grained control over camera movement, which is critical for downstream applications related to content creation, visual effects, an…

Cited by 38SourcePDFScholar
2024

4D-fy: Text-to-4D Generation Using Hybrid Score Distillation Sampling

CVPR 2024poster

Recent breakthroughs in text-to-4D generation rely on pre-trained text-to-image and text-to-video models to generate dynamic 3D scenes. However current text-to-4D methods face a three-way tradeoff between the quality of scene appearance 3D structure and motion. For example text-to-image models and t…

2023

CC3D: Layout-Conditioned Generation of Compositional 3D Scenes

ICCV 2023poster

In this work, we introduce CC3D, a conditional generative model that synthesizes complex 3D scenes conditioned on 2D semantic scene layouts, trained using single-view images. Different from most existing 3D GANs that limit their applicability to aligned single objects, we focus on generating complex…

Cited by 44PDFScholar
2022

Towards Robust and Adaptive Motion Forecasting: A Causal Representation Perspective

CVPR 2022poster

Learning behavioral patterns from observational data has been a de-facto approach to motion forecasting. Yet, the current paradigm suffers from two shortcomings: brittle under distribution shifts and inefficient for knowledge transfer. In this work, we propose to address these challenges from a caus…

Cited by 68PDFScholar