← Search

huan ling

25 accepted papers

2026

ChronoEdit: Towards Temporal Reasoning for In-Context Image Editing and World Simulation

ICLR 2026poster

Recent advances in large generative models have significantly advanced image editing and in-context image generation, yet a critical gap remains in ensuring physical consistency, where edited objects must remain coherent. This capability is especially vital for world simulation related tasks. In thi…

Cited by 0SourcecodeScholar
2026

Lyra: Generative 3D Scene Reconstruction via Video Diffusion Model Self-Distillation

ICLR 2026poster

The ability to generate virtual environments is crucial for applications ranging from gaming to physical AI domains such as robotics, autonomous driving, and industrial AI. Current learning-based 3D reconstruction methods rely on the availability of captured real-world multi-view data, which is not…

Cited by 0SourcecodeScholar
2026

MHLA: Restoring Expressivity of Linear Attention via Token-Level Multi-Head

ICLR 2026poster

While the Transformer architecture dominates many fields, its quadratic self-attention complexity hinders its use in large-scale applications. **Linear attention** offers an efficient alternative, but its direct application often degrades performance, with existing fixes typically re-introducing com…

Cited by 0SourceScholar
2026

SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer

ICLR 2026oral

We introduce SANA-Video, a small diffusion model that can efficiently generate videos up to 720×1280 resolution and minute-length duration. SANA-Video synthesizes high-resolution, high-quality and long videos with strong text-video alignment at a remarkably fast speed, deployable on RTX 5090 GPU. Tw…

Cited by 0SourcecodeScholar
2025

DIFIX3D+: Improving 3D Reconstructions with Single-Step Diffusion Models

CVPR 2025award

Neural Radiance Fields and 3D Gaussian Splatting have revolutionized 3D reconstruction and novel-view synthesis task. However, achieving photorealistic rendering from extreme novel viewpoints remains challenging, as artifacts persist across representations. In this work, we introduce Difix3D+, a nov…

2025

Diffusion Renderer: Neural Inverse and Forward Rendering with Video Diffusion Models

CVPR 2025poster

Understanding and modeling lighting effects are fundamental tasks in computer vision and graphics. Classic physically-based rendering (PBR) accurately simulates the light transport, but relies on precise scene representations--explicit 3D geometry, high-quality material properties, and lighting cond…

Cited by 3SourcePDFScholar
2025

Feed-Forward Bullet-Time Reconstruction of Dynamic Scenes from Monocular Videos

NeurIPS 2025poster

Recent advancements in static feed-forward scene reconstruction have demonstrated significant progress in high-quality novel view synthesis. However, these models often struggle with generalizability across diverse environments and fail to effectively handle dynamic content. We present BTimer (short…

Cited by 0SourceScholar
2025

GEN3C: 3D-Informed World-Consistent Video Generation with Precise Camera Control

CVPR 2025highlight

We present GEN3C, a generative video model with precise Camera Control and temporal 3D Consistency. Prior video models already generate realistic videos, but they tend to leverage little 3D information, leading to inconsistencies, such as objects popping in and out of existence. Camera control, if i…

2024

3DiffTection: 3D Object Detection with Geometry-Aware Diffusion Features

CVPR 2024poster

3DiffTection introduces a novel method for 3D object detection from single images utilizing a 3D-aware diffusion model for feature extraction. Addressing the resource-intensive nature of annotating large-scale 3D image data our approach leverages pretrained diffusion models traditionally used for 2D…

Cited by 13SourcePDFScholar
2024

Align Your Gaussians: Text-to-4D with Dynamic 3D Gaussians and Composed Diffusion Models

CVPR 2024highlight

Text-guided diffusion models have revolutionized image and video generation and have also been successfully used for optimization-based 3D object synthesis. Here we instead focus on the underexplored text-to-4D setting and synthesize dynamic animated 3D objects using score distillation methods with…

Cited by 110SourcePDFScholar
2024

L4GM: Large 4D Gaussian Reconstruction Model

NeurIPS 2024poster

We present L4GM, the first 4D Large Reconstruction Model that produces animated objects from a single-view video input -- in a single feed-forward pass that takes only a second. Key to our success is a novel dataset of multiview videos containing curated, rendered animated objects from Objaverse. Th…

Cited by 38SourcePDFScholar
2024

SCube: Instant Large-Scale Scene Reconstruction using VoxSplats

NeurIPS 2024poster

We present SCube, a novel method for reconstructing large-scale 3D scenes (geometry, appearance, and semantics) from a sparse set of posed images. Our method encodes reconstructed scenes using a novel representation VoxSplat, which is a set of 3D Gaussians supported on a high-resolution sparse-voxel…

Cited by 10SourcePDFScholar
2023

Align Your Latents: High-Resolution Video Synthesis With Latent Diffusion Models

CVPR 2023poster

Latent Diffusion Models (LDMs) enable high-quality image synthesis while avoiding excessive compute demands by training a diffusion model in a compressed lower-dimensional latent space. Here, we apply the LDM paradigm to high-resolution video generation, a particularly resource-intensive task. We fi…

2023

DreamTeacher: Pretraining Image Backbones with Deep Generative Models

ICCV 2023poster

In this work, we introduce a self-supervised feature representation learning framework DreamTeacher that utilizes generative networks for pre-training downstream image backbones. We propose to distill knowledge from a trained generative model into standard image backbones that have been well enginee…

Cited by 23PDFScholar
2022

BigDatasetGAN: Synthesizing ImageNet With Pixel-Wise Annotations

CVPR 2022poster

Annotating images with pixel-wise labels is a time-consuming and costly process. Recently, DatasetGAN showcased a promising alternative - to synthesize a large labeled dataset via a generative adversarial network (GAN) by exploiting a small set of manually labeled, GAN-generated images. Here, we sca…

Cited by 121PDFScholar
2021

DatasetGAN: Efficient Labeled Data Factory With Minimal Human Effort

CVPR 2021poster

We introduce DatasetGAN: an automatic procedure to generate massive datasets of high-quality semantically segmented images requiring minimal human effort. Current deep networks are extremely data-hungry, benefiting from training on large-scale datasets, which are time-consuming to annotate. Our meth…

Cited by 393PDFcodeScholar
2021

EditGAN: High-Precision Semantic Image Editing

NeurIPS 2021poster

Generative adversarial networks (GANs) have recently found applications in image editing. However, most GAN-based image editing methods often require large-scale datasets with semantic segmentation annotations for training, only provide high-level control, or merely interpolate between different ima…

Cited by 280SourcePDFScholar
2021

Image GANs meet Differentiable Rendering for Inverse Graphics and Interpretable 3D Neural Rendering

ICLR 2021oral

Differentiable rendering has paved the way to training neural networks to perform “inverse graphics” tasks such as predicting 3D geometry from monocular photographs. To train high performing models, most of the current approaches rely on multi-view imagery which are not readily available in practice…

Cited by 145SourcePDFScholar
2020

ScribbleBox: Interactive Annotation Framework for Video Object Segmentation

ECCV 2020poster

Manually labeling video datasets for segmentation tasks is extremely time consuming. We introduce ScribbleBox, an interactive framework for annotating object instances with masks in videos with a significant boost in efficiency. In particular, we split annotation into two steps: annotating objects w…

Cited by 23SourcePDFScholar
2019

Learning to Predict 3D Objects with an Interpolation-based Differentiable Renderer

NeurIPS 2019poster

Many machine learning models operate on images, but ignore the fact that images are 2D projections formed by 3D geometry interacting with light, in a process called rendering. Enabling ML models to understand image formation might be key for generalization. However, due to an essential rasterization…

Cited by 453SourcePDFScholar
2019

Object Instance Annotation With Deep Extreme Level Set Evolution

CVPR 2019poster

In this paper, we tackle the task of interactive object segmentation. We revive the old ideas on level set segmentation which framed object annotation as curve evolution. Carefully designed energy functions ensured that the curve was well aligned with image boundaries, and generally "well behaved".…

Cited by 98PDFcodeScholar
2018

Efficient Interactive Annotation of Segmentation Datasets With Polygon-RNN++

CVPR 2018poster

Manually labeling datasets with object masks is extremely time consuming. In this work, we follow the idea of Polygon-RNN to produce polygonal annotations of objects interactively using humans-in-the-loop. We introduce several important improvements to the model: 1) we design a new CNN encoder archi…

Cited by 537SourcePDFScholar