← Search

Yuanze Lin

7 accepted papers

2025

IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation

NeurIPS 2025poster

Although diffusion-based models can generate high-quality and high-resolution video sequences from textual or image inputs, they lack explicit integration of geometric cues when controlling scene lighting and visual appearance across frames. To address this limitation, we propose IllumiCraft, an end…

Cited by 0SourceScholar
2025

Olympus: A Universal Task Router for Computer Vision Tasks

CVPR 2025highlight

We introduce Olympus, a new approach that transforms Multimodal Large Language Models (MLLMs) into a unified framework capable of handling a wide array of computer vision tasks. Utilizing a controller MLLM, Olympus delegates over 20 specialized tasks across images, videos, and 3D objects to dedicate…

2024

Text-Driven Image Editing via Learnable Regions

CVPR 2024poster

Language has emerged as a natural interface for image editing. In this paper we introduce a method for region-based image editing driven by textual prompts without the need for user-provided masks or sketches. Specifically our approach leverages an existing pre-trained text-to-image model and introd…

2023

SMAUG: Sparse Masked Autoencoder for Efficient Video-Language Pre-Training

ICCV 2023poster

Video-language pre-training is crucial for learning powerful multi-modal representation. However, it typically requires a massive amount of computation. In this paper, we develop SMAUG, an efficient pre-training framework for video-language models. The foundation component in SMAUG is masked autoenc…

Cited by 16PDFScholar
2022

AdaFocus V2: End-to-End Training of Spatial Dynamic Networks for Video Recognition

CVPR 2022poster

Recent works have shown that the computational efficiency of video recognition can be significantly improved by reducing the spatial redundancy. As a representative work, the adaptive focus method (AdaFocus) has achieved a favorable trade-off between accuracy and inference speed by dynamically ident…

Cited by 63PDFcodeScholar
2022

Pseudo-Q: Generating Pseudo Language Queries for Visual Grounding

CVPR 2022poster

Visual grounding, i.e., localizing objects in images according to natural language queries, is an important topic in visual language understanding. The most effective approaches for this task are based on deep learning, which generally require expensive manually labeled image-query or patch-query pa…

Cited by 73PDFcodeScholar