← Search

Wenwu Wang

71 accepted papers

2026

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips

CVPR 2026

Foley art plays a pivotal role in enhancing immersive auditory experiences in film, yet manual creation of spatio-temporal aligned audio remains labor-intensive. We propose FoleyDesigner, a novel framework inspired by professional Foley workflows, integrating film clip analysis, spatio-temporal cont

Cited by 0SourceScholar
2026

PHYSICS-AWARE NOVEL-VIEW ACOUSTIC SYNTHESIS WITH VISION-LANGUAGE PRIORS AND 3D ACOUSTIC ENVIRONMENT MODELING

ICASSP 2026poster

Spatial audio is essential for immersive experiences, yet novel-view acoustic synthesis (NVAS) remains challenging due to complex physical phenomena such as reflection, diffraction, and material absorption. Existing methods based on single-view or panoramic inputs improve spatial fidelity but fail t…

Cited by 0SourcePDFScholar
2026

RFM-EDITING: RECTIFIED FLOW MATCHING FOR TEXT-GUIDED AUDIO EDITING

ICASSP 2026poster

Diffusion models have shown remarkable progress in text-to-audio generation. However, text-guided audio editing remains in its early stages. This task focuses on modifying the target content within an audio signal while preserving the rest, thus demanding precise localization and faithful editing ac…

Cited by 7SourcePDFScholar
2026

Teacher-Guided Pseudo Supervision and Cross-Modal Alignment for Audio-Visual Video Parsing

ICASSP 2026poster

Weakly-supervised audio-visual video parsing (AVVP) seeks to detect audible, visible, and audio-visual events without temporal annotations. Previous work has emphasized refining global predictions through contrastive or collaborative learning, but neglected stable segment-level supervision and class…

Cited by 0SourcePDFScholar
2026

The Power of Prior: Training-Free Open-Vocabulary Semantic Segmentation with LLaVA

CVPR 2026

Multimodal Large Language Models (MLLMs) like LLaVA have demonstrated remarkable capabilities in multi-modal understanding and generation. This success motivates us to investigate whether the inherent prior knowledge embedded within such MLLMs contains sufficient spatial awareness for dense predicti

Cited by 0SourcecodeScholar
2025

Bayesian Nonparametric Clustering for Source Counting with a Small Aperture Microphone Array

ICASSP 2025accepted

Source counting (SC) in an indoor environment is an important problem in computational auditory scene analysis. However, the problem is challenging, especially when reverberation and ambient noise are present in the environment. To address this problem, we propose an augmented Bayesian non-parametri…

Cited by 0SourceScholar
2025

FlowSep: Language-Queried Sound Separation with Rectified Flow Matching

ICASSP 2025accepted

Language-queried audio source separation (LASS) focuses on separating sounds using textual descriptions of the desired sources. Current methods mainly use discriminative approaches, such as time-frequency masking, to separate target sounds and minimize interference from other sources. However, these…

Cited by 0SourceScholar
2025

Graph-Enhanced Dual-Stream Feature Fusion with Pre-Trained Model for Acoustic Traffic Monitoring

ICASSP 2025accepted

Microphone array techniques are widely used in sound source localization and smart city acoustic-based traffic monitoring, but these applications face significant challenges due to the scarcity of labeled real-world traffic audio data and the complexity and diversity of application scenarios. The DC…

Cited by 0SourceScholar
2025

Sound-Based Recognition of Touch Gestures and Emotions for Enhanced Human-Robot Interaction

ICASSP 2025accepted

Emotion recognition and touch gesture decoding are crucial for advancing human-robot interaction (HRI), especially in social environments where emotional cues and tactile perception play important roles. However, many humanoid robots, such as Pepper, Nao, and Furhat, lack full-body tactile skin, lim…

Cited by 0SourceScholar
2025

Sound-VECaps: Improving Audio Generation with Visually Enhanced Captions

ICASSP 2025accepted

Generative models have shown significant achievements in audio generation tasks. However, existing models struggle with complex and detailed prompts, leading to potential performance degradation. We hypothesize that this problem stems from the simplicity and scarcity of the training data. This work…

Cited by 0SourceScholar
2025

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey

EMNLP 2025

Text-to-speech (TTS) has advanced from generating natural-sounding speech to enabling fine-grained control over attributes like emotion, timbre, and style. Driven by rising industrial demand and breakthroughs in deep learning, e.g., diffusion and large language models (LLMs), controllable TTS has be

2024

CM-PIE: Cross-Modal Perception for Interactive-Enhanced Audio-Visual Video Parsing

ICASSP 2024accepted

Audio-visual video parsing is the task of categorizing a video with weak labels at the segment level, and predicting them as audible or visible events. Recent methods have leveraged the attention mechanism to capture the semantic correlations among the whole video across the audio-visual modalities.…

Cited by 0SourceScholar
2024

First-Shot Unsupervised Anomalous Sound Detection with Unknown Anomalies Estimated by Metadata-Assisted Audio Generation

ICASSP 2024accepted

First-shot (FS) unsupervised anomalous sound detection (ASD) is a brand-new task introduced in DCASE 2023 Challenge Task 2, where the anomalous sounds for the target machine types are unseen in training. Existing methods often rely on the availability of normal and abnormal sound data from the targe…

Cited by 0SourceScholar
2024

Fusion of Audio and Visual Embeddings for Sound Event Localization and Detection

ICASSP 2024accepted

Sound event localization and detection (SELD) combines two subtasks: sound event detection (SED) and direction of arrival (DOA) estimation. SELD is usually tackled as an audio-only problem, but visual information has been recently included. Few audio-visual (AV)-SELD works have been published and mo…

Cited by 0SourceScholar
2024

Hierarchical Metadata Information Constrained Self-Supervised Learning for Anomalous Sound Detection under Domain Shift

ICASSP 2024accepted

Self-supervised learning methods have achieved promising performance for anomalous sound detection (ASD) under domain shift by incorporating the metadata of domain shift types and machine sound attributes in feature learning. However, the relation between domain shifts and machine sound attributes h…

Cited by 7SourceScholar
2024

Learning Temporal Resolution in Spectrogram for Audio Classification

AAAI 2024technical

The audio spectrogram is a time-frequency representation that has been widely used for audio classification. One of the key attributes of the audio spectrogram is the temporal resolution, which depends on the hop size used in the Short-Time Fourier Transform (STFT). Previous works generally assume t…

2024

Multi-Level Graph Learning For Audio Event Classification And Human-Perceived Annoyance Rating Prediction

ICASSP 2024accepted

WHO’s report on environmental noise estimates that 22 M people suffer from chronic annoyance related to noise caused by audio events (AEs) from various sources. Annoyance may lead to health issues and adverse effects on metabolic and cognitive systems. In cities, monitoring noise levels does not pro…

Cited by 0SourceScholar
2024

Multi-Speaker Localization in the Circular Harmonic Domain on Small Aperture Microphone Arrays Using Deep Convolutional Networks

ICASSP 2024accepted

Acoustic signal processing in the circular harmonic domain (CHD) is an appealing method for speaker localization, since it inherently supports wideband acoustic sources and provides frequency invariant beampatterns. However, the performance of existing circular harmonic direction-of-arrival (DOA) es…

Cited by 0SourceScholar
2024

Retrieval-Augmented Text-to-Audio Generation

ICASSP 2024accepted

Despite recent progress in text-to-audio (TTA) generation, we show that the state-of-the-art models, such as AudioLDM, trained on datasets with an imbalanced class distribution, such as AudioCaps, are biased in their generation performance. Specifically, they excel in generating common audio classes…

Cited by 0SourceScholar
2024

Selective Prompting Tuning for Personalized Conversations with LLMs

ACL 2024findings

In conversational AI, personalizing dialogues with persona profiles and contextual understanding is essential. Despite large language models’ (LLMs) improved response coherence, effective persona integration remains a challenge. In this work, we first study two common approaches for personalizing LL…

2023

An Improved Optimal Transport Kernel Embedding Method with Gating Mechanism for Singing Voice Separation and Speaker Identification

ICASSP 2023accepted

Singing voice separation (SVS) and speaker identification (SI) are two classic problems in speech signal processing. Deep neural networks (DNNs) solve these two problems by extracting effective representations of the target signal from the input mixture. Since essential features of a signal can be w…

Cited by 0SourceScholar
2023

Anomalous Sound Detection Using Audio Representation with Machine ID Based Contrastive Learning Pretraining

ICASSP 2023accepted

Existing contrastive learning methods for anomalous sound detection refine the audio representation of each audio sample by using the contrast between the samples’ augmentations (e.g., with time or frequency masking). However, they might be biased by the augmented data, due to the lack of physical p…

Cited by 0SourceScholar
2023

AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

ICML 2023poster

Text-to-audio (TTA) systems have recently gained attention for their ability to synthesize general audio based on text descriptions. However, previous studies in TTA have limited generation quality with high computational costs. In this study, we propose AudioLDM, a TTA system that is built on a lat…

2023

Learning Retrieval Augmentation for Personalized Dialogue Generation

EMNLP 2023long main

Personalized dialogue generation, focusing on generating highly tailored responses by leveraging persona profiles and dialogue context, has gained significant attention in conversational AI applications. However, persona profiles, a prevalent setting in current personalized dialogue datasets, typica…

Cited by 0SourcecodeScholar
2023

Personalized Dialogue Generation with Persona-Adaptive Attention

AAAI 2023technical

Persona-based dialogue systems aim to generate consistent responses based on historical context and predefined persona. Unlike conventional dialogue generation, the persona-based dialogue needs to consider both dialogue context and persona, posing a challenge for coherent training. Specifically, thi…

2023

Simple Pooling Front-Ends for Efficient Audio Classification

ICASSP 2023accepted

Recently, there has been increasing interest in building efficient audio neural networks for on-device scenarios. Most existing approaches are designed to reduce the size of audio neural networks using methods such as model pruning. In this work, we show that instead of reducing model size using com…

Cited by 0SourceScholar
2023

Time-Weighted Frequency Domain Audio Representation with GMM Estimator for Anomalous Sound Detection

ICASSP 2023accepted

Although deep learning is the mainstream method in unsupervised anomalous sound detection, Gaussian Mixture Model (GMM) with statistical audio frequency representation as input can achieve comparable results with much lower model complexity and fewer parameters. Existing statistical frequency repres…

Cited by 0SourceScholar
2022

A Mutual Learning Framework for Few-Shot Sound Event Detection

ICASSP 2022accepted

Although prototypical network (ProtoNet) has proved to be an effective method for few-shot sound event detection, two problems still exist. Firstly, the small-scaled support set is insufficient so that the class prototypes may not represent the class center accurately. Secondly, the feature extracto…

Cited by 0SourceScholar
2022

Audio-Visual Tracking of Multiple Speakers Via a PMBM Filter

ICASSP 2022accepted

Audio-visual tracking of multiple speakers requires to estimate the state (e.g. velocity and location) of each speaker by leveraging the information of both audio and visual modalities. Estimating the number of speakers and their states jointly remains a challenging problem. We propose an Audio-Visu…

Cited by 0SourceScholar
2022

DSTAGNN: Dynamic Spatial-Temporal Aware Graph Neural Network for Traffic Flow Forecasting

ICML 2022spotlight

As a typical problem in time series analysis, traffic flow prediction is one of the most important application fields of machine learning. However, achieving highly accurate traffic flow prediction is a challenging task, due to the presence of complex dynamic spatial-temporal dependencies within a r…

2022

Partial Arithmetic Consensus based Distributed Intensity Particle Flow SMC-PHD Filter for Multi-Target Tracking

ICASSP 2022accepted

Intensity Particle Flow (IPF) SMC-PHD has been proposed recently for multi-target tracking. In this paper, we extend IPF-SMC-PHD filter to distributed setting, and develop a novel consensus method for fusing the estimates from individual sensors, based on Arithmetic Average (AA) fusion. Different fr…

Cited by 0SourceScholar
2021

An Improved Event-Independent Network for Polyphonic Sound Event Localization and Detection

ICASSP 2021accepted

Polyphonic sound event localization and detection (SELD), which jointly performs sound event detection (SED) and direction-of-arrival (DoA) estimation, detects the type and occurrence time of sound events as well as their corresponding DoA angles simultaneously. We study the SELD task from a multi-t…

Cited by 0SourceScholar
2021

Enhancing Audio Augmentation Methods with Consistency Learning

ICASSP 2021accepted

Data augmentation is an inexpensive way to increase training data diversity, and is commonly achieved via transformations of existing data. For tasks such as classification, there is a good case for learning representations of the data that are invariant to such transformations, yet this is not expl…

Cited by 0SourceScholar
2021

Low-Dimensional Denoising Embedding Transformer for ECG Classification

ICASSP 2021accepted

The transformer based model (e.g., FusingTF) has been employed recently for Electrocardiogram (ECG) signal classification. However, the high-dimensional embedding obtained via 1-D convolution and positional encoding can lead to the loss of the signal’s own temporal information and a large amount of…

Cited by 0SourceScholar
2020

Learning With Out-of-Distribution Data for Audio Classification

ICASSP 2020accepted

In supervised machine learning, the assumption that training data is labelled correctly is not always satisfied. In this paper, we investigate an instance of labelling error for classification tasks in which the dataset is corrupted with out-of-distribution (OOD) instances: data that does not belong…

Cited by 0SourceScholar
2020

Meta Metric Learning for Highly Imbalanced Aerial Scene Classification

ICASSP 2020accepted

Class imbalance is an important factor that affects the performance of deep learning models used for remote sensing scene classification. In this paper, we propose a random finetuning meta metric learning model (RF-MML) to address this problem. Derived from episodic training in meta metric learning,…

Cited by 0SourceScholar
2020

Source Separation with Weakly Labelled Data: an Approach to Computational Auditory Scene Analysis

ICASSP 2020accepted

Source separation is the task of separating an audio recording into individual sound sources. Source separation is fundamental for computational auditory scene analysis. Previous work on source separation has focused on separating particular sound classes such as speech and music. Much previous work…

Cited by 0SourceScholar
2020

Weakly Labelled Audio Tagging Via Convolutional Networks with Spatial and Channel-Wise Attention

ICASSP 2020accepted

Multiple instance learning (MIL) with convolutional neural networks (CNNs) has been proposed recently for weakly labelled audio tagging. However, features from the various CNN filtering channels and spatial regions are often treated equally, which may limit its performance in event prediction. In th…

Cited by 0SourceScholar
2019

Acoustic Scene Generation with Conditional Samplernn

ICASSP 2019accepted

Acoustic scene generation (ASG) is a task to generate waveforms for acoustic scenes. ASG can be used to generate audio scenes for movies and computer games. Recently, neural networks such as SampleRNN have been used for speech and music generation. However, ASG is more challenging due to its wide va…

Cited by 0SourceScholar
2019

Background Adaptation for Improved Listening Experience in Broadcasting

ICASSP 2019accepted

The intelligibility of speech in noise can be improved by modifying the speech. But with object-based audio, there is the possibility of altering the background sound while leaving the speech unaltered. This may prove a less intrusive approach, affording good speech intelligibility without overly co…

Cited by 0SourceScholar
2019

Enhanced Streaming Based Subspace Clustering Applied to Acoustic Scene Data Clustering

ICASSP 2019accepted

Labelled data are often required to train an acoustic scene classification system. However, it is time-consuming and expensive to label the data manually. An unsupervised clustering algorithm can be used to facilitate the labelling process by dividing the acoustic data into different categories. Nev…

Cited by 0SourceScholar
2019

Generalisation in Environmental Sound Classification: The 'Making Sense of Sounds' Data Set and Challenge

ICASSP 2019accepted

Humans are able to identify a large number of environmental sounds and categorise them according to high-level semantic categories, e.g. urban sounds or music. They are also capable of generalising from past experience to new sounds when applying these categories. In this paper we report on the crea…

Cited by 14SourceScholar
2019

Proximal Deep Recurrent Neural Network for Monaural Singing Voice Separation

ICASSP 2019accepted

The recent deep learning methods can offer state-of-the-art performance for Monaural Singing Voice Separation (MSVS). In these deep methods, the recurrent neural network (RNN) is widely employed. This work proposes a novel type of Deep RNN (DRNN), namely Proximal DRNN (P-DRNN) for MSVS, which improv…

Cited by 0SourceScholar
2019

Semantic Super-resolution for Extremely Low-resolution Vehicle License Plate

ICASSP 2019accepted

Vehicle license plate (VLP) super-resolution (SR) is of great demand in intelligent traffic systems. Super-Resolution for extremely low-resolution VLP remains challenging and the state-of-the-art SR methods hardly provide satisfying results for low-resolution (LR) VLPs. In this study, from a new per…

Cited by 0SourceScholar
2018

A Joint Separation-Classification Model for Sound Event Detection of Weakly Labelled Data

ICASSP 2018accepted

Source separation (SS) aims to separate individual sources from an audio recording. Sound event detection (SED) aims to detect sound events from an audio recording. We propose a joint separation-classification (JSC) model trained only on weakly labelled audio data, that is, only the tags of an audio…

Cited by 0SourceScholar
2018

Audio Set Classification with Attention Model: A Probabilistic Perspective

ICASSP 2018accepted

This paper investigates the Audio Set classification. Audio Set is a large scale weakly labelled dataset (WLD) of audio clips. In WLD only the presence of a label is known, without knowing the happening time of the labels. We propose an attention model to solve this WLD problem and explain the atten…

Cited by 0SourceScholar
2018

Bayesian Inference for Multi-Line Spectra in Linear Sensor Array

ICASSP 2018accepted

For a linear sensor array, using line spectra is a common technique for estimating directions of arrival (DOA) of single-tone sources. Yet, very few papers consider multitone sources. For the first time, we provide the optimal Bayesian inference for multi-line spectra, i.e. a superposition of line s…

Cited by 1SourceScholar
2018

Intelligent Signal Processing Mechanisms for Nuanced Anomaly Detection in Action Audio-Visual Data Streams

ICASSP 2018accepted

We consider the problem of anomaly detection in an audiovisual analysis system designed to interpret sequences of actions from visual and audio cues. The scene activity recognition is based on a generative framework, with a high-level inference model for contextual recognition of sequences of action…

Cited by 0SourceScholar
2018

Iterative Deep Neural Networks for Speaker-Independent Binaural Blind Speech Separation

ICASSP 2018accepted

In this paper, we propose an iterative deep neural network (DNN)-based binaural source separation scheme, for recovering two concurrent speech signals in a room environment. Besides the commonly-used spectral features, the DNN also takes non-linearly wrapped binaural spatial features as input, which…

Cited by 11SourceScholar
2018

Large-Scale Weakly Supervised Audio Classification Using Gated Convolutional Neural Network

ICASSP 2018accepted

In this paper, we present a gated convolutional neural network and a temporal attention-based localization method for audio classification, which won the 1st place in the large-scale weakly supervised sound event detection task of Detection and Classification of Acoustic Scenes and Events (DCASE) 20…

Cited by 0SourceScholar
2018

Non-Zero Diffusion Particle Flow SMC-PHD Filter for Audio-Visual Multi-Speaker Tracking

ICASSP 2018accepted

The sequential Monte Carlo probability hypothesis density (SMC-PHD) filter has been shown to be promising for audio-visual multi-speaker tracking. Recently, the zero diffusion particle flow (ZPF) has been used to mitigate the weight degeneracy problem in the SMC-PHD filter. However, this leads to a…

Cited by 0SourceScholar
2018

Synthesis of Images by Two-Stage Generative Adversarial Networks

ICASSP 2018accepted

In this paper, we propose a divide-and-conquer approach using two generative adversarial networks (GANs) to explore how a machine can draw colorful pictures (bird) using a small amount of training data. In our work, we simulate the procedure of an artist drawing a picture, where one begins with draw…

Cited by 0SourceScholar
2017

A greedy algorithm with learned statistics for sparse signal reconstruction

ICASSP 2017accepted

We address the problem of sparse signal reconstruction from a few noisy samples. Recently, a Covariance-Assisted Matching Pursuit (CAMP) algorithm has been proposed, improving the sparse coefficient update step of the classic Orthogonal Matching Pursuit (OMP) algorithm. CAMP allows the a-priori mean…

Cited by 0SourceScholar
2017

A joint detection-classification model for audio tagging of weakly labelled data

ICASSP 2017accepted

Audio tagging aims to assign one or several tags to an audio clip. Most of the datasets are weakly labelled, which means only the tags of the clip are known, without knowing the occurrence time of the tags. The labeling of an audio clip is often based on the audio events in the clip and no event lev…

Cited by 0SourceScholar
2017

Assessment of musical noise using localization of isolated peaks in time-frequency domain

ICASSP 2017accepted

Musical noise is a recurrent issue that appears in spectral techniques for denoising or blind source separation. Due to localised errors of estimation, isolated peaks may appear in the processed spectrograms, resulting in annoying tonal sounds after synthesis known as “musical noise”. In this paper,…

Cited by 0SourceScholar
2017

Fast tagging of natural sounds using marginal co-regularization

ICASSP 2017accepted

Automatic and fast tagging of natural sounds in audio collections is a very challenging task due to wide acoustic variations, the large number of possible tags, the incomplete and ambiguous tags provided by different labellers. To handle these problems, we use a co-regularization approach to learn a…

Cited by 0SourceScholar
2017

Particle flow for sequential Monte Carlo implementation of probability hypothesis density

ICASSP 2017accepted

Target tracking is a challenging task and generally no analytical solution is available, especially for the multi-target tracking systems. To address this problem, probability hypothesis density (PHD) filter is used by propagating the PHD instead of the full multi-target posterior. Recently, the par…

Cited by 0SourceScholar
2016

Identity association using PHD filters in multiple head tracking with depth sensors

ICASSP 2016accepted

The work on 3D human pose estimation has been through a significant amount of progress in recent years, particularly due to the widespread availability of commodity depth sensors. However, most pose estimation methods follow a tracking-as-detection approach which does not explicitly handle occlusion…

Cited by 0SourceScholar
2016

Social force model aided robust particle PHD filter for multiple human tracking

ICASSP 2016accepted

In this paper, we propose a novel robust multiple human tracking approach based upon processing a video signal by utilizing a social force model to enhance the particle probability hypothesis density (PHD) filter. In traditional dynamic models, the states of targets are only predicted by their own h…

Cited by 0SourceScholar