← Search

Shuchen Weng

22 accepted papers

2026

Audio-sync Video Instance Editing with Granularity-Aware Mask Refiner

CVPR 2026

Recent advancements in video generation highlight that realistic audio-visual synchronization is crucial for engaging content creation. However, existing video editing methods largely overlook audio-visual synchronization and lack the fine-grained spatial and temporal controllability required for pr

Cited by 0SourcecodeScholar
2026

Lighting-grounded Video Generation with Renderer-based Agent Reasoning

CVPR 2026

Diffusion models have achieved remarkable progress in video generation, but their controllability remains a major limitation. Key scene factors such as layout, lighting, and camera trajectory are often entangled or only weakly modeled, restricting their applicability in domains like filmmaking and v

Cited by 0SourceScholar
2026

STAGE: Storyboard-Anchored Generation for Cinematic Multi-shot Narrative

CVPR 2026

While recent advancements in generative models have achieved remarkable visual fidelity in video synthesis, creating coherent multi-shot narratives remains a significant challenge. To address this, keyframe-based approaches have emerged as a promising alternative to computationally intensive end-to-

Cited by 0SourceScholar
2025

Audio-Sync Video Generation with Multi-Stream Temporal Control

NeurIPS 2025poster

Audio is inherently temporal and closely synchronized with the visual world, making it a naturally aligned and expressive control signal for controllable video generation (e.g., movies). Beyond control, directly translating audio into video is essential for understanding and visualizing rich audio n…

Cited by 0SourceScholar
2025

PanoWan: Lifting Diffusion Video Generation Models to 360$^\circ$ with Latitude/Longitude-aware Mechanisms

NeurIPS 2025poster

Panoramic video generation enables immersive 360$^\circ$ content creation, valuable in applications that demand scene-consistent world exploration. However, existing panoramic video generation models struggle to leverage pre-trained generative priors from conventional text-to-video models for high-q…

Cited by 0SourceScholar
2025

PhyS-EdiT: Physics-aware Semantic Image Editing with Text Description

CVPR 2025poster

Achieving joint control over material properties, lighting, and high-level semantics in images is essential for applications in digital media, advertising, and interactive design. Existing methods often isolate these properties, lacking a cohesive approach to manipulating materials, lighting, and se…

Cited by 0SourcePDFScholar
2025

VIRES: Video Instance Repainting via Sketch and Text Guided Generation

CVPR 2025poster

We introduce VIRES, a video instance repainting method with sketch and text guidance, enabling video instance repainting, replacement, generation, and removal. Existing approaches struggle with temporal consistency and accurate alignment with the provided sketch sequence. VIRES leverages the generat…

Cited by 0SourcePDFScholar
2024

Colorizing Monochromatic Radiance Fields

AAAI 2024technical

Though Neural Radiance Fields (NeRF) can produce colorful 3D representations of the world by using a set of 2D images, such ability becomes non-existent when only monochromatic images are provided. Since color is necessary in representing the world, reproducing color from monochromatic radiance fiel…

2024

L-DiffER: Single Image Reflection Removal with Language-based Diffusion Model

ECCV 2024poster

"In this paper, we introduce L-DiffER, a language-based diffusion model designed for the ill-posed single image reflection removal task. Although having shown impressive performance for image generation, existing language-based diffusion models struggle with precise control and faithfulness in image…

Cited by 4SourcePDFScholar
2023

Affective Image Filter: Reflecting Emotions from Text to Images

ICCV 2023poster

Understanding the emotions in text and presenting them visually is a very challenging problem that requires a deep understanding of natural language and high-quality image synthesis simultaneously. In this work, we propose Affective Image Filter (AIF), a novel model that is able to understand the vi…

Cited by 14PDFScholar
2023

L-CAD: Language-based Colorization with Any-level Descriptions using Diffusion Priors

NeurIPS 2023spotlight

Language-based colorization produces plausible and visually pleasing colors under the guidance of user-friendly natural language descriptions. Previous methods implicitly assume that users provide comprehensive color descriptions for most of the objects in the image, which leads to suboptimal perfor…

2023

L-CoIns: Language-Based Colorization With Instance Awareness

CVPR 2023poster

Language-based colorization produces plausible colors consistent with the language description provided by the user. Recent studies introduce additional annotation to prevent color-object coupling and mismatch issues, but they still have difficulty in distinguishing instances corresponding to the sa…

Cited by 28SourcePDFScholar
2023

LuminAIRe: Illumination-Aware Conditional Image Repainting for Lighting-Realistic Generation

NeurIPS 2023poster

We present the ilLumination-Aware conditional Image Repainting (LuminAIRe) task to address the unrealistic lighting effects in recent conditional image repainting (CIR) methods. The environment lighting and 3D geometry conditions are explicitly estimated from given background images and parsing mask…

Cited by 5SourcePDFScholar
2022

L-CoDe:Language-Based Colorization Using Color-Object Decoupled Conditions

AAAI 2022technical

Colorizing a grayscale image is inherently an ill-posed problem with multi-modal uncertainty. Language-based colorization offers a natural way of interaction to reduce such uncertainty via a user-provided caption. However, the color-object coupling and mismatch issues make the mapping from word to c…

Cited by 41SourcePDFScholar
2022

L-CoDer: Language-Based Colorization with Color-Object Decoupling Transformer

ECCV 2022poster

"Language-based colorization requires the colorized image to be consistent with the the user-provided language caption. A most recent work proposes to decouple the language into color and object conditions in solving the problem. Though decent progress has been made, its performance is limited by th…

2020

Conditional Image Repainting via Semantic Bridge and Piecewise Value Function

ECCV 2020poster

We study conditional image repainting where a model is trained to generate visual content conditioned on user inputs, and composite the generated content seamlessly onto a user provided image while preserving the semantics of users' inputs. The content generation community have been pursuing to lowe…

Cited by 6SourcePDFScholar
2020

MISC: Multi-Condition Injection and Spatially-Adaptive Compositing for Conditional Person Image Synthesis

CVPR 2020poster

In this paper, we explore synthesizing person images with multiple conditions for various backgrounds. To this end, we propose a framework named "MISC" for conditional image generation and image compositing. For conditional image generation, we improve the existing condition injection mechanisms by…

Cited by 37PDFScholar