← Search

Zian Wang

18 accepted papers

2026

ChronoEdit: Towards Temporal Reasoning for In-Context Image Editing and World Simulation

ICLR 2026poster

Recent advances in large generative models have significantly advanced image editing and in-context image generation, yet a critical gap remains in ensuring physical consistency, where edited objects must remain coherent. This capability is especially vital for world simulation related tasks. In thi…

Cited by 0SourcecodeScholar
2026

DiffusionHarmonizer: Bridging Neural Reconstruction and Photorealistic Simulation with Online Diffusion Enhancer

CVPR 2026

Simulation is essential to the development and evaluation of autonomous robots such as self-driving vehicles. Neural reconstruction is emerging as a promising solution as it enables simulating a wide variety of scenarios from real-world data alone in an automated and scalable way. However, while met

Cited by 0SourcecodeScholar
2026

Training-Free Hierarchical Working Memory for Small Language Model Agents

ICML 2026poster

Small language models (SLMs) are attractive for agent deployment, but they struggle to reliably retain and reuse decision-relevant state information over long interactions. This issue is exacerbated when working memory is maintained via unstructured natural-language summarization. Some recent work a…

Cited by 0SourceScholar
2025

Controllable Weather Synthesis and Removal with Video Diffusion Models

ICCV 2025poster

Generating realistic and controllable weather effects in videos is valuable for many applications. Physics-based weather simulation requires precise reconstructions that are hard to scale to in-the-wild videos, while current video editing often lacks realism and control.In this work, we introduce We…

Cited by 0SourcePDFScholar
2025

Diffusion Renderer: Neural Inverse and Forward Rendering with Video Diffusion Models

CVPR 2025poster

Understanding and modeling lighting effects are fundamental tasks in computer vision and graphics. Classic physically-based rendering (PBR) accurately simulates the light transport, but relies on precise scene representations--explicit 3D geometry, high-quality material properties, and lighting cond…

Cited by 3SourcePDFScholar
2025

LuxDiT: Lighting Estimation with Video Diffusion Transformer

NeurIPS 2025poster

Estimating scene lighting from a single image or video remains a longstanding challenge in computer vision and graphics. Learning-based approaches are constrained by the scarcity of ground-truth HDR environment maps, which are expensive to capture and limited in diversity. While recent generative mo…

Cited by 0SourceScholar
2025

RobustKV: Defending Large Language Models against Jailbreak Attacks via KV Eviction

ICLR 2025poster

Jailbreak attacks circumvent LLMs' built-in safeguards by concealing harmful queries within adversarial prompts. While most existing defenses attempt to mitigate the effects of adversarial prompts, they often prove inadequate as adversarial prompts can take arbitrary, adaptive forms. This paper intr…

Cited by 4SourcePDFScholar
2025

UniRelight: Learning Joint Decomposition and Synthesis for Video Relighting

NeurIPS 2025spotlight

We address the challenge of relighting a single image or video, a task that demands precise scene intrinsic understanding and high-quality light transport synthesis. Existing end-to-end relighting models are often limited by the scarcity of paired multi-illumination data, restricting their ability t…

Cited by 0SourceScholar
2025

Watermark under Fire: A Robustness Evaluation of LLM Watermarking

EMNLP 2025

Various watermarking methods (“watermarkers”) have been proposed to identify LLM-generated texts; yet, due to the lack of unified evaluation platforms, many critical questions remain under-explored: i) What are the strengths/limitations of various watermarkers, especially their attack robustness? ii

2024

Misalignment-Robust Frequency Distribution Loss for Image Transformation

CVPR 2024poster

This paper aims to address a common challenge in deep learning-based image transformation methods such as image enhancement and super-resolution which heavily rely on precisely aligned paired datasets with pixel-level alignments. However creating precisely aligned paired images presents significant…

2023

Neural Fields Meet Explicit Geometric Representations for Inverse Rendering of Urban Scenes

CVPR 2023poster

Reconstruction and intrinsic decomposition of scenes from captured imagery would enable many applications such as relighting and virtual object insertion. Recent NeRF based methods achieve impressive fidelity of 3D reconstruction, but bake the lighting and shadows into the radiance field, while mesh…

Cited by 89SourcePDFScholar
2023

Neural LiDAR Fields for Novel View Synthesis

ICCV 2023poster

We present Neural Fields for LiDAR (NFL), a method to optimise a neural field scene representation from LiDAR measurements, with the goal of synthesizing realistic LiDAR scans from novel viewpoints. NFL combines the rendering power of neural fields with a detailed, physically motivated model of the…

Cited by 61PDFScholar
2022

GET3D: A Generative Model of High Quality 3D Textured Shapes Learned from Images

NeurIPS 2022accept

As several industries are moving towards modeling massive 3D virtual worlds, the need for content creation tools that can scale in terms of the quantity, quality, and diversity of 3D content is becoming evident. In our work, we aim to train performant 3D generative models that synthesize textured me…

2022

Neural Light Field Estimation for Street Scenes with Differentiable Virtual Object Insertion

ECCV 2022poster

"We consider the challenging problem of outdoor lighting estimation for the goal of photorealistic virtual object insertion into photographs. Existing works on outdoor lighting estimation typically simplify the scene lighting into an environment map which cannot capture the spatially-varying lightin…

Cited by 41SourcePDFScholar
2021

DIB-R++: Learning to Predict Lighting and Material with a Hybrid Differentiable Renderer

NeurIPS 2021poster

We consider the challenging problem of predicting intrinsic object properties from a single image by exploiting differentiable renderers. Many previous learning-based approaches for inverse graphics adopt rasterization-based renderers and assume naive lighting and material models, which often fail t…

Cited by 68SourcePDFScholar
2020

Beyond Fixed Grid: Learning Geometric Image Representation with a Deformable Grid

ECCV 2020poster

In modern computer vision, images are typically represented as a fixed uniform grid with some stride and processed via a deep convolutional neural network. We argue that deforming the grid to better align with the high-frequency image content is a more effective strategy. We introduce mph{Deformable…

2019

Object Instance Annotation With Deep Extreme Level Set Evolution

CVPR 2019poster

In this paper, we tackle the task of interactive object segmentation. We revive the old ideas on level set segmentation which framed object annotation as curve evolution. Carefully designed energy functions ensured that the curve was well aligned with image boundaries, and generally "well behaved".…

Cited by 98PDFcodeScholar