← Search

Si Li

28 accepted papers

2026

Lighting-grounded Video Generation with Renderer-based Agent Reasoning

CVPR 2026

Diffusion models have achieved remarkable progress in video generation, but their controllability remains a major limitation. Key scene factors such as layout, lighting, and camera trajectory are often entangled or only weakly modeled, restricting their applicability in domains like filmmaking and v

Cited by 0SourceScholar
2026

STAGE: Storyboard-Anchored Generation for Cinematic Multi-shot Narrative

CVPR 2026

While recent advancements in generative models have achieved remarkable visual fidelity in video synthesis, creating coherent multi-shot narratives remains a significant challenge. To address this, keyframe-based approaches have emerged as a promising alternative to computationally intensive end-to-

Cited by 0SourceScholar
2025

Audio-Sync Video Generation with Multi-Stream Temporal Control

NeurIPS 2025poster

Audio is inherently temporal and closely synchronized with the visual world, making it a naturally aligned and expressive control signal for controllable video generation (e.g., movies). Beyond control, directly translating audio into video is essential for understanding and visualizing rich audio n…

Cited by 0SourceScholar
2025

GUI Exploration Lab: Enhancing Screen Navigation in Agents via Multi-Turn Reinforcement Learning

NeurIPS 2025poster

With the rapid development of Large Vision Language Models, the focus of Graphical User Interface (GUI) agent tasks shifts from single-screen tasks to complex screen navigation challenges. However, real-world GUI environments, such as PC software and mobile Apps, are often complex and proprietary,…

Cited by 0SourceScholar
2025

PolGS: Polarimetric Gaussian Splatting for Fast Reflective Surface Reconstruction

ICCV 2025poster

Efficient shape reconstruction for surfaces with complex reflectance properties is crucial for real-time virtual reality. While 3D Gaussian Splatting (3DGS)-based methods offer fast novel view rendering by leveraging their explicit surface representation, their reconstruction quality lags behind tha…

Cited by 0SourcePDFScholar
2025

PolarAnything: Diffusion-based Polarimetric Image Synthesis

ICCV 2025poster

Polarization images facilitate image enhancement and 3D reconstruction tasks, but the limited accessibility of polarization cameras hinders their broader application. This gap drives the need for synthesizing photorealistic polarization images. The existing polarization simulator Mitsuba relies on a…

Cited by 0SourcePDFScholar
2025

VIRES: Video Instance Repainting via Sketch and Text Guided Generation

CVPR 2025poster

We introduce VIRES, a video instance repainting method with sketch and text guidance, enabling video instance repainting, replacement, generation, and removal. Existing approaches struggle with temporal consistency and accurate alignment with the provided sketch sequence. VIRES leverages the generat…

Cited by 0SourcePDFScholar
2024

Leveraging Generative Large Language Models with Visual Instruction and Demonstration Retrieval for Multimodal Sarcasm Detection

NAACL 2024long

Multimodal sarcasm detection aims to identify sarcasm in the given image-text pairs and has wide applications in the multimodal domains. Previous works primarily design complex network structures to fuse the image-text modality features for classification. However, such complicated structures may ri…

Cited by 14SourcePDFScholar
2024

SfPUEL: Shape from Polarization under Unknown Environment Light

NeurIPS 2024poster

Shape from polarization (SfP) benefits from advancements like polarization cameras for single-shot normal estimation, but its performance heavily relies on light conditions. This paper proposes SfPUEL, an end-to-end SfP method to jointly estimate surface normal and material under unknown environment…

2024

Type-Aware Decoding Via Explicitly Aggregating Event Information for Document-Level Event Extraction

ICASSP 2024accepted

Document-level event extraction (DEE) faces two main challenges: arguments-scattering and multi-event. Although previous methods attempt to address these challenges, they overlook the interference of event-unrelated sentences during event detection and neglect the mutual interference of different ev…

Cited by 0SourceScholar
2024

Visual Enhanced Entity-Level Interaction Network for Multimodal Summarization

NAACL 2024findings

MultiModal Summarization (MMS) aims to generate a concise summary based on multimodal data like texts and images and has wide application in multimodal fields.Previous works mainly focus on the coarse-level textual and visual features in which the overall features of the image interact with the whol…

2023

Affective Image Filter: Reflecting Emotions from Text to Images

ICCV 2023poster

Understanding the emotions in text and presenting them visually is a very challenging problem that requires a deep understanding of natural language and high-quality image synthesis simultaneously. In this work, we propose Affective Image Filter (AIF), a novel model that is able to understand the vi…

Cited by 14PDFScholar
2023

Complementary Intrinsics From Neural Radiance Fields and CNNs for Outdoor Scene Relighting

CVPR 2023poster

Relighting an outdoor scene is challenging due to the diverse illuminations and salient cast shadows. Intrinsic image decomposition on outdoor photo collections could partly solve this problem by weakly supervised labels with albedo and normal consistency from multi-view stereo. With neural radiance…

Cited by 9SourcePDFScholar
2023

DemoSG: Demonstration-enhanced Schema-guided Generation for Low-resource Event Extraction

EMNLP 2023long findings

Most current Event Extraction (EE) methods focus on the high-resource scenario, which requires a large amount of annotated data and can hardly be applied to low-resource domains. To address EE more effectively with limited resources, we propose the Demonstration-enhanced Schema-guided Generation (De…

Cited by 0SourceScholar
2023

L-CAD: Language-based Colorization with Any-level Descriptions using Diffusion Priors

NeurIPS 2023spotlight

Language-based colorization produces plausible and visually pleasing colors under the guidance of user-friendly natural language descriptions. Previous methods implicitly assume that users provide comprehensive color descriptions for most of the objects in the image, which leads to suboptimal perfor…

2023

L-CoIns: Language-Based Colorization With Instance Awareness

CVPR 2023poster

Language-based colorization produces plausible colors consistent with the language description provided by the user. Recent studies introduce additional annotation to prevent color-object coupling and mismatch issues, but they still have difficulty in distinguishing instances corresponding to the sa…

Cited by 28SourcePDFScholar
2022

Dependency Parsing via Sequence Generation

EMNLP 2022finding

Dependency parsing aims to extract syntactic dependency structure or semantic dependency structure for sentences.Existing methods for dependency parsing include transition-based method, graph-based method and sequence-to-sequence method.These methods obtain excellent performance and we notice them b…

2022

Entity-level Interaction via Heterogeneous Graph for Multimodal Named Entity Recognition

EMNLP 2022finding

Multimodal Named Entity Recognition (MNER) faces two specific challenges: 1) How to capture useful entity-related visual information. 2) How to alleviate the interference of visual noise. Previous works have gained progress by improving interacting mechanisms or seeking for better visual features. H…

2022

Estimating Spatially-Varying Lighting in Urban Scenes with Disentangled Representation

ECCV 2022poster

"We present an end-to-end network for spatially-varying outdoor lighting estimation in urban scenes given a single limited field-of-view LDR image and any assigned 2D pixel position. We use three disentangled latent spaces learned by our network to represent sky light, sun light, and lighting-indepe…

Cited by 15SourcePDFScholar
2022

L-CoDe:Language-Based Colorization Using Color-Object Decoupled Conditions

AAAI 2022technical

Colorizing a grayscale image is inherently an ill-posed problem with multi-modal uncertainty. Language-based colorization offers a natural way of interaction to reduce such uncertainty via a user-provided caption. However, the color-object coupling and mismatch issues make the mapping from word to c…

Cited by 41SourcePDFScholar
2022

L-CoDer: Language-Based Colorization with Color-Object Decoupling Transformer

ECCV 2022poster

"Language-based colorization requires the colorized image to be consistent with the the user-provided language caption. A most recent work proposes to decouple the language into color and object conditions in solving the problem. Though decent progress has been made, its performance is limited by th…

2021

A Joint Model for Dropped Pronoun Recovery and Conversational Discourse Parsing in Chinese Conversational Speech

ACL 2021long

In this paper, we present a neural model for joint dropped pronoun recovery (DPR) and conversational discourse parsing (CDP) in Chinese conversational speech. We show that DPR and CDP are closely related, and a joint model benefits both tasks. We refer to our model as DiscProReco, and it first encod…

2021

Unsupervised Domain Adaptation Method with Semantic-Structural Alignment for Dependency Parsing

EMNLP 2021finding

Unsupervised cross-domain dependency parsing is to accomplish domain adaptation for dependency parsing without using labeled data in target domain. Existing methods are often of the pseudo-annotation type, which generates data through self-annotation of the base model and performing iterative traini…

Cited by 2SourcePDFScholar
2019

Reflection Separation using a Pair of Unpolarized and Polarized Images

NeurIPS 2019spotlight

When we take photos through glass windows or doors, the transmitted background scene is often blended with undesirable reflection. Separating two layers apart to enhance the image quality is of vital importance for both human and machine perception. In this paper, we propose to exploit physical cons…