← Search

Xiaopeng LIN

9 accepted papers

2026

AD-MIR: Bridging the Gap from Perception to Persuasion in Advertising Video Understanding via Structured Reasoning

ICML 2026poster

Multimodal understanding of advertising videos is essential for interpreting the intricate relationship between visual storytelling and abstract persuasion strategies. However, despite excelling at general search, existing agents often struggle to bridge the cognitive gap between pixel-level percept…

Cited by 0SourceScholar
2026

DIMOS: Disentangling Instance-level Moving Object Segmentation

CVPR 2026

Moving instance segmentation (MIS) attracts increasing attention due to its broad applications in traffic surveillance, autonomous driving, and animal tracking. Event cameras record asynchronous brightness changes, providing high temporal resolution and dynamic range, which makes them highly sensiti

Cited by 0SourceScholar
2026

LangForce: Bayesian Decomposition of Vision Language Action Models via Latent Action Queries

ICML 2026poster

Vision-Language-Action (VLA) models have shown promise in robot manipulation but often struggle to generalize to new instructions or complex multi-task scenarios. We identify a critical pathology in current training paradigms where goal-driven data collection creates a dataset bias. In such datasets…

Cited by 0SourceScholar
2026

MDN: Parallelizing Stepwise Momentum for Delta Linear Attention

ICML 2026poster

Linear Attention (LA) offers a promising paradigm for scaling large language models (LLMs) to long sequences by avoiding the quadratic complexity of self-attention. Recent LA models such as Mamba2 and GDN interpret linear recurrences as closed-form online stochastic gradient descent (SGD), but naive…

Cited by 0SourceScholar
2026

Scalable Event Cloud Network for Event-based Classification

ICML 2026oral

Event cameras are biologically inspired sensors garnering significant attention from both industry and academia. Mainstream methods favor frame and voxel representations, which reach a satisfactory performance while introducing time-consuming transformations, bulky models, and sacrificing fine-grain…

Cited by 0SourceScholar
2025

ClearSight: Human Vision-Inspired Solutions for Event-Based Motion Deblurring

ICCV 2025poster

Motion deblurring addresses the challenge of image blur caused by camera or scene movement. Event cameras provide motion information that is encoded in the asynchronous event streams. To efficiently leverage the temporal information of event streams, we employ Spiking Neural Networks (SNNs) for moti…

Cited by 0SourcePDFScholar
2024

CLIF: Complementary Leaky Integrate-and-Fire Neuron for Spiking Neural Networks

ICML 2024spotlight

Spiking neural networks (SNNs) are promising brain-inspired energy-efficient models. Compared to conventional deep Artificial Neural Networks (ANNs), SNNs exhibit superior efficiency and capability to process temporal information. However, it remains a challenge to train SNNs due to their undifferen…

2024

SpikePoint: An Efficient Point-based Spiking Neural Network for Event Cameras Action Recognition

ICLR 2024spotlight

Event cameras are bio-inspired sensors that respond to local changes in light intensity and feature low latency, high energy efficiency, and high dynamic range. Meanwhile, Spiking Neural Networks (SNNs) have gained significant attention due to their remarkable efficiency and fault tolerance. By syne…

Cited by 25SourcePDFScholar
2022

Underwater Image Enhancement Via Learning Water Type Desensitized Representations

ICASSP 2022accepted

We present a novel underwater image enhancement method termed SCNet to improve the image quality meanwhile cope with the degradation diversity caused by the water. SCNet is based on normalization schemes across both spatial and channel dimensions with the key idea of learning water type desensitized…

Cited by 0SourceScholar