← Search

Li Su

35 accepted papers

2026

Joint Learning of General and Diverse Patterns with Mixture of Memory Experts for Weakly-Supervised Video Anomaly Detection

CVPR 2026

Weakly-supervised Video Anomaly Detection (wVAD) aims to detect abnormal events using only binary labels, making it challenging to capture both the diversity of anomalies and their shared semantic cues. Existing methods either focus on a generic anomaly pattern, achieving strong generalization but w

Cited by 0SourceScholar
2026

STaR: Sensitive Trajectory Regulation for Unlearning in Large Reasoning Models

AAAI 2026technical

Large Reasoning Models (LRMs) have advanced automated multi-step reasoning, but their ability to generate complex Chain-of-Thought (CoT) trajectories introduces severe privacy risks, as sensitive information may be deeply embedded throughout the reasoning process. Existing Large Language Models (LLM

Cited by 0SourcePDFScholar
2026

SYNTHCLONER: SYNTHESIZER-STYLE AUDIO TRANSFER VIA FACTORIZED CODEC WITH ADSR ENVELOPE CONTROL

ICASSP 2026poster

Electronic synthesizer sounds are controlled by parameter settings that yield complex timbral characteristics and ADSR envelopes, making synthesizer-style audio transfer particularly challenging. Recent approaches to timbre transfer often rely on spectral objectives or implicit style matching, offer…

Cited by 0SourcePDFScholar
2026

VioPTT: Violin Technique-Aware Transcription from Synthetic Data Augmentation

ICASSP 2026poster

While automatic music transcription is well-established in music information retrieval, most models are limited to transcribing pitch and timing information from audio, and thus omit crucial expressive and instrument-specific nuances. One example is playing technique on the violin, which affords its…

Cited by 0SourcePDFScholar
2025

Change Entity-guided Heterogeneous Representation Disentangling for Change Captioning

ACL 2025finding

Change captioning aims to describe differences between a pair of images using natural language. However, learning effective difference representations is highly challenging due to distractors such as illumination and viewpoint changes. To address this, we propose a change-entity-guided disentangleme…

2025

Generalizing Single-Frame Supervision to Event-Level Understanding for Video Anomaly Detection

NeurIPS 2025poster

Video Anomaly Detection (VAD) aims to identify abnormal frames from discrete events within video sequences. Existing VAD methods suffer from heavy annotation burdens in fully-supervised paradigm, insensitivity to subtle anomalies in semi-supervised paradigm, and vulnerability to noise in weakly-supe…

Cited by 0SourceScholar
2025

Query-centric Audio-Visual Cognition Network for Moment Retrieval, Segmentation and Step-Captioning

AAAI 2025technical

Video has emerged as a favored multimedia format on the internet. To better gain video contents, a new topic HIREST is presented, including video retrieval, moment retrieval, moment segmentation, and step-captioning. The pioneering work chooses the pre-trained CLIP-based model for video retrieval,…

2024

Context-aware Difference Distilling for Multi-change Captioning

ACL 2024long

Multi-change captioning aims to describe complex and coupled changes within an image pair in natural language. Compared with single-change captioning, this task requires the model to have higher-level cognition ability to reason an arbitrary number of changes. In this paper, we propose a novel conte…

2024

Distractors-Immune Representation Learning with Cross-modal Contrastive Regularization for Change Captioning

ECCV 2024poster

"Change captioning aims to succinctly describe the semantic change between a pair of similar images, while being immune to distractors (illumination and viewpoint changes). Under these distractors, unchanged objects often appear pseudo changes about location and scale, and certain objects might over…

2024

Leveraging Catastrophic Forgetting to Develop Safe Diffusion Models against Malicious Finetuning

NeurIPS 2024spotlight

Diffusion models (DMs) have demonstrated remarkable proficiency in producing images based on textual prompts. Numerous methods have been proposed to ensure these models generate safe images. Early methods attempt to incorporate safety filters into models to mitigate the risk of generating harmful im…

Cited by 1SourcePDFScholar
2024

Prompt-Enhanced Multiple Instance Learning for Weakly Supervised Video Anomaly Detection

CVPR 2024poster

Weakly-supervised Video Anomaly Detection (wVAD) aims to detect frame-level anomalies using only video-level labels in training. Due to the limitation of coarse-grained labels Multi-Instance Learning (MIL) is prevailing in wVAD. However MIL suffers from insufficiency of binary supervision to model d…

2023

Audio-Driven Facial Landmark Generation in Violin Performance using 3DCNN Network with Self Attention Model

ICASSP 2023accepted

In a music scenario, both auditory and visual elements are essential to achieve an outstanding performance. Recent research has focused on the generation of body movements or fingering from audio in music performance. The audio-driven face generation technique in music performance is still deficient…

Cited by 0SourceScholar
2023

Decoding Musical Pitch from Human Brain Activity with Automatic Voxel-Wise Whole-Brain FMRI Feature Selection

ICASSP 2023accepted

Decoding models seek to infer stimulus or task information from neural activity and play a central role in brain-computer interfaces. However, the high spatial resolution of fMRI means that the number of available features far exceeds the number of trials in a typical experiment. Although a common a…

Cited by 0SourceScholar
2023

Note and Playing Technique Transcription of Electric Guitar Solos in Real-World Music Performance

ICASSP 2023accepted

Transcribing electric guitar solo in real-world performance is challenging because of the interference of background accompaniments, the strong coupling between music pitch and playing technique, and the limited resource of data annotation. To address these issues, we first propose a new guitar solo…

Cited by 0SourceScholar
2023

Self-supervised Cross-view Representation Reconstruction for Change Captioning

ICCV 2023poster

Change captioning aims to describe the difference between a pair of similar images. Its key challenge is how to learn a stable difference representation under pseudo changes caused by viewpoint change. In this paper, we address this by proposing a self-supervised cross-view representation reconstruc…

Cited by 36PDFcodeScholar
2022

A Sparse-Motif Ensemble Graph Convolutional Network against Over-smoothing

IJCAI 2022poster

The over-smoothing issue is a well-known challenge for Graph Convolutional Networks (GCN). Specifically, it is often observed that increasing the depth of GCN ends up in a trivial embedding subspace where the difference among node embeddings belonging to the same cluster tends to vanish. This paper…

2021

Rethinking Graph Neural Architecture Search From Message-Passing

CVPR 2021poster

Graph neural networks (GNNs) emerged recently as a standard toolkit for learning from data on graphs. Current GNN designing works depend on immense human expertise to explore different message-passing mechanisms, and require manual enumeration to determine the proper message-passing depth. Inspired…

Cited by 66PDFcodeScholar
2020

Body Movement Generation for Expressive Violin Performance Applying Neural Networks

ICASSP 2020accepted

Generating body movements based on given music audio recordings is an emerging research topic. This problem remains challenging particularly for string instruments, considering the fact that the relationship between the musical note sequences and the body movement sequences in string instruments doe…

Cited by 0SourceScholar
2020

Reverse Perspective Network for Perspective-Aware Object Counting

CVPR 2020poster

One of the critical challenges of object counting is the dramatic scale variations, which is introduced by arbitrary perspectives. We propose a reverse perspective network to solve the scale variations of input images, instead of generating perspective maps to smooth final outputs. The reverse persp…

Cited by 168PDFScholar
2020

Weakly-Supervised Crowd Counting Learns from Sorting rather than Locations

ECCV 2020poster

In crowd counting datasets, the location labels are costly, yet, they are not taken into the evaluation metrics. Besides, existing multi-task approaches employ high-level tasks to improve counting accuracy. This research tendency increases the demand for more annotations. In this paper, we propose a…

Cited by 108SourcePDFScholar
2018

Automatic Music Transcription Leveraging Generalized Cepstral Features and Deep Learning

ICASSP 2018accepted

Spectral features are limited in modeling musical signals with multiple concurrent pitches due to the challenge to suppress the interference of the harmonic peaks from one pitch to another. In this paper, we show that using multiple features represented in both the frequency and time domains with de…

Cited by 0SourceScholar
2018

Vocal Melody Extraction Using Patch-Based CNN

ICASSP 2018accepted

A patch-based convolutional neural network (CNN) model presented in this paper for vocal melody extraction in polyphonic music is inspired from object detection in image processing. The input of the model is a novel time-frequency representation which enhances the pitch contours and suppresses the h…

Cited by 0SourceScholar
2017

Automatic conversion of Pop music into chiptunes for 8-bit pixel art

ICASSP 2017accepted

In this paper, we propose an audio mosaicing method that converts Pop songs into a specific music style called “chiptune,” or “8-bit music.” The goal is to reproduce Pop songs by using the sound of the chips on the old game consoles in 1980s/1990s. The proposed method goes through a procedure that f…

Cited by 0SourceScholar
2017

Polyphonic piano note transcription with non-negative matrix factorization of differential spectrogram

ICASSP 2017accepted

Automatic music transcription is usually approached by using a time-frequency (TF) representation such as the short-time Fourier transform (STFT) spectrogram or the constant-Q transform. In this paper, we propose a novel yet simple TF representation that capitalizes the effectiveness of spectral flu…

Cited by 0SourceScholar
2015

Vocal activity informed singing voice separation with the iKala dataset

ICASSP 2015accepted

A new algorithm is proposed for robust principal component analysis with predefined sparsity patterns. The algorithm is then applied to separate the singing voice from the instrumental accompaniment using vocal activity information. To evaluate its performance, we construct a new publicly available…

Cited by 0SourceScholar