← Search

Xiulian Peng

16 accepted papers

2026

From Content to Knowledge: Lightning Fast Long-Video Understanding with Neural Knowledge Representations

ICML 2026poster

We propose a new paradigm for long video understanding by treating a long video as a Neural Knowledge Representation (NKR). NKR represent video contents neither as a stream of tokens or pre-organized databases, but as an individual small portion of network weights attached to the VLM backbone. The N…

Cited by 0SourceScholar
2026

Hierarchical Long Video Understanding with Audiovisual Entity Cohesion and Agentic Search

CVPR 2026

Long video understanding presents significant challenges for vision-language models due to extremely long context windows. Existing solutions relying on naive chunking strategies with retrieval-augmented generation, typically suffer from information fragmentation and a loss of global coherence. We p

Cited by 0SourceScholar
2025

Bitrate-Controlled Diffusion for Disentangling Motion and Content in Video

ICCV 2025poster

We propose a novel and general framework to disentangle video data into its dynamic motion and static content components. Our proposed method is a self-supervised pipeline with less assumptions and inductive biases than previous works: it utilizes a transformer-based architecture to jointly generate…

Cited by 0SourcePDFScholar
2024

QDFormer: Towards Robust Audiovisual Segmentation in Complex Environments with Quantization-based Semantic Decomposition

CVPR 2024poster

Audiovisual segmentation (AVS) is a challenging task that aims to segment visual objects in videos according to their associated acoustic cues. With multiple sound sources and background disturbances involved establishing robust correspondences between audio and visual contents poses unique challeng…

2023

Dasformer: Deep Alternating Spectrogram Transformer For Multi/Single-Channel Speech Separation

ICASSP 2023accepted

For the task of speech separation, previous study usually treats multi-channel and single-channel scenarios as two research tracks with specialized solutions developed respectively. Instead, we propose a simple and unified architecture - DasFormer (Deep alternating spectrogram transFormer) to handle…

Cited by 0SourceScholar
2022

End-to-End Neural Speech Coding for Real-Time Communications

ICASSP 2022accepted

Deep-learning based methods have shown their advantages in audio coding over traditional ones but limited attention has been paid on real-time communications (RTC). This paper proposes the TFNet, an end-to-end neural speech codec with low latency for RTC. It takes an encoder-temporal filtering-decod…

Cited by 0SourceScholar
2022

Reference-Based Speech Enhancement via Feature Alignment and Fusion Network

AAAI 2022technical

Speech enhancement aims at recovering a clean speech from a noisy input, which can be classified into single speech enhancement and personalized speech enhancement. Personalized speech enhancement usually utilizes the speaker identity extracted from the noisy speech itself (or a clean reference spee…

2021

Interactive Speech and Noise Modeling for Speech Enhancement

AAAI 2021technical

Speech enhancement is challenging because of the diversity of background noise types. Most of the existing methods are focused on modelling the speech rather than the noise. In this paper, we propose a novel idea to model speech and noise simultaneously in a two-branch convolutional neural network,…

2018

Frequency-Domain Dynamic Pruning for Convolutional Neural Networks

NeurIPS 2018poster

Deep convolutional neural networks have demonstrated their powerfulness in a variety of applications. However, the storage and computational requirements have largely restricted their further extensions on mobile devices. Recently, pruning of unimportant parameters has been used for both network com…

Cited by 200SourcePDFScholar