← Search

You Li

20 accepted papers

2026

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts

CVPR 2026

Recent Video-to-Audio (V2A) methods have achieved remarkable progress, enabling the synthesis of realistic, high-quality audio. However, they struggle with fine-grained temporal control in multi-event scenarios or when visual cues are insufficient, such as small regions, off-screen sounds, or occlud

Cited by 0SourcecodeScholar
2026

Imagination Helps Visual Reasoning, But Not Yet in Latent Space

ICML 2026poster

Latent visual reasoning aims to mimic human's *imagination* process by meditating through hidden states of Multimodal Large Language Models. While recognized as a promising paradigm for visual reasoning, the underlying mechanisms driving its effectiveness remain unclear. Motivated to demystify the t…

Cited by 0SourceScholar
2026

LLA: Enhancing Security and Privacy for Generative Models with Logic-Locked Accelerators

AAAI 2026technical

We introduce LLA, an effective intellectual property (IP) protection scheme for generative AI models. LLA leverages the synergy between hardware and software to defend against various supply chain threats, including model theft, model corruption, and information leakage. On the software side, it emb

Cited by 0SourcePDFScholar
2026

TopoMA: Topology-Guided Multi-Agent Dense RGB 3D Reconstruction via Distributed Inference

CVPR 2026

Multi-agent 3D reconstruction, as a key technology for large-scale VR/AR, robot swarms, and digital twins, has attracted growing attention. Recent end-to-end 3D reconstruction methods achieve strong performance in single-agent scenarios, but they are difficult to directly extend to multi-agent colla

Cited by 0SourceScholar
2025

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models

ACL 2025finding

The recent advancement of Multimodal Large Language Models (MLLMs) has significantly improved their fine-grained perception of single images and general comprehension across multiple images. However, existing MLLMs still face challenges in achieving precise grounding in complex multi-image scenarios…

2025

Multi-Stage LLM Fine-Tuning with a Continual Learning Setting

NAACL 2025findings

In recent years, large language models (LLMs) have made significant progress in knowledge-intensive applications. However, when adapting them to specific domains, we may encounter a multi-stage continuous learning scenario, especially in cases where domain knowledge evolves rapidly.This issue severe…

Cited by 1SourcePDFScholar
2025

Noise Fusion-based Distillation Learning for Anomaly Detection in Complex Industrial Environments

IROS 2025

Anomaly detection and localization in automated industrial manufacturing can significantly enhance production efficiency and product quality. Existing methods are capable of detecting surface defects in pre-defined or controlled imaging environments. However, accurately detecting workpiece defects i

Cited by 0SourcecodeScholar
2025

Think in Safety: Unveiling and Mitigating Safety Alignment Collapse in Multimodal Large Reasoning Model

EMNLP 2025

The rapid development of Multimodal Large Reasoning Models (MLRMs) has demonstrated broad application potential, yet their safety and reliability remain critical concerns that require systematic exploration. To address this gap, we conduct a comprehensive and systematic safety evaluation of 13 MLRMs

2024

AS-LIO: Spatial Overlap Guided Adaptive Sliding Window LiDAR-Inertial Odometry for Aggressive FOV Variation

IROS 2024poster

LiDAR-Inertial Odometry (LIO) demonstrates outstanding accuracy and stability in general low-speed and smooth motion scenarios. However, in high-speed and intense motion scenarios, such as sharp turns, two primary challenges arise: firstly, due to the limitations of IMU frequency, the error in estim…

Cited by 3SourceScholar
2024

Knowledge-Guided Dynamic Modality Attention Fusion Framework for Multimodal Sentiment Analysis

EMNLP 2024finding

Multimodal Sentiment Analysis (MSA) utilizes multimodal data to infer the users’ sentiment. Previous methods focus on equally treating the contribution of each modality or statically using text as the dominant modality to conduct interaction, which neglects the situation where each modality may beco…

2024

MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis

CVPR 2024highlight

We present a Multi-Instance Generation (MIG) task simultaneously generating multiple instances with diverse controls in one image. Given a set of predefined coordinates and their corresponding descriptions the task is to ensure that generated instances are accurately at the designated locations and…

2024

Neural Poisson Solver: A Universal and Continuous Framework for Natural Signal Blending

ECCV 2024poster

"Implicit Neural Representation (INR) has become a popular method for representing visual signals (, 2D images and 3D scenes), demonstrating promising results in various downstream applications. Given its potential as a medium for visual signals, exploring the development of a neural blending method…

Cited by 0SourcePDFScholar
2023

Learning Decomposed Spatial Relations for Multi-Variate Time-Series Modeling

AAAI 2023technical

Modeling multi-variate time-series (MVTS) data is a long-standing research subject and has found wide applications. Recently, there is a surge of interest in modeling spatial relations between variables as graphs, i.e., first learning one static graph for each dataset and then exploiting the graph s…

Cited by 20SourcePDFScholar
2023

Learning to Learn: How to Continuously Teach Humans and Machines

ICCV 2023poster

Curriculum design is a fundamental component of education. For example, when we learn mathematics at school, we build upon our knowledge of addition to learn multiplication. These and other concepts must be mastered before our first algebra lesson, which also reinforces our addition and multiplicati…

Cited by 5PDFScholar
2023

RePaint-NeRF: NeRF Editting via Semantic Masks and Diffusion Models

IJCAI 2023poster

The emergence of Neural Radiance Fields (NeRF) has promoted the development of synthesized high-fidelity views of the intricate real world. However, it is still a very demanding task to repaint the content in NeRF. In this paper, we propose a novel framework that can take RGB images as input and alt…

2023

Rethinking the Construction of Effective Metrics for Understanding the Mechanisms of Pretrained Language Models

EMNLP 2023long findings

Pretrained language models are expected to effectively map input text to a set of vectors while preserving the inherent relationships within the text. Consequently, designing a white-box model to compute metrics that reflect the presence of specific internal relations in these vectors has become a c…

Cited by 0SourcecodeScholar
2022

MDOE: A Spatiotemporal Event Representation Considering the Magnitude and Density of Events

RA-L 2022

Event-based sensors (e.g., DVS cameras) are capable of higher dynamic range, higher temporal resolution, lower time latency, and better power efficiency compared to conventional devices (e.g., RGB cameras). However, learning from these sensors remains challenging; event-based sensors output a stream

Cited by 4SourceScholar
2022

Novel Design of a Cable-Driven Continuum Robot With Multiple Motion Patterns

RA-L 2022

Cable-driven continuum robots exhibit accessing and manipulation capabilities in constrained and cluttered environments that are unachievable for traditional robots composing of discrete links and joints. However, the existing continuum robots structured by non-scalable backbones are monofunctional,

Cited by 41SourceScholar
2020

LaNoising: A Data-driven Approach for 903nm ToF LiDAR Performance Modeling under Fog

IROS 2020poster

As a critical sensor for high-level autonomous vehicles, LiDAR's limitations in adverse weather (e.g. rain, fog, snow, etc.) impede the deployment of self-driving cars in all weather conditions. In this paper, we model the performance of a popular 903nm ToF LiDAR under various fog conditions based o…

Cited by 31SourceScholar