← Search

HUY PHAN

25 accepted papers

2025

Effective Techniques for Scaling Audio Encoder Pretraining

ICASSP 2025accepted

This work presents advancements in audio pretraining objectives designed to generate semantically rich embeddings, capable of addressing a wide range of audio-related tasks. Despite significant progress in the field, current methods often emphasize full fine-tuning in downstream applications, which…

Cited by 0SourceScholar
2025

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction

ICML 2025poster

Autoregressive next-token prediction with the Transformer decoder has become a de facto standard in large language models (LLMs), achieving remarkable success in Natural Language Processing (NLP) at scale. Extending this paradigm to audio poses unique challenges due to its inherently continuous natu…

Cited by 0SourcePDFScholar
2025

IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion Modeling

ICML 2025poster

Text-to-audio generation synthesizes realistic sounds or music given a natural language prompt. Diffusion-based frameworks, including the Tango and the AudioLDM series, represent the state-of-the-art in text-to-audio generation. Despite achieving high audio fidelity, they incur significant inference…

2025

LHGNN: Local-Higher Order Graph Neural Networks For Audio Classification and Tagging

ICASSP 2025accepted

Transformers have set new benchmarks in audio processing tasks, leveraging self-attention mechanisms to capture complex patterns and dependencies within audio data. However, their focus on pairwise interactions limits their ability to process the higher-order relations essential for identifying dist…

Cited by 0SourceScholar
2024

Cross-Triggering Issue in Audio Event Detection and Mitigation

ICASSP 2024accepted

Cross-triggering is a critical problem for applications of audio event detection (AED), particularly in low-resource settings. However, not much attention (if not none) has been paid to this problem in the AED research community. In this work, we tackle this problem via a regularization approach. We…

Cited by 0SourceScholar
2024

Learning from Taxonomy: Multi-Label Few-Shot Classification for Everyday Sound Recognition

ICASSP 2024accepted

Humans categorise and structure perceived acoustic signals into hierarchies of auditory objects. The semantics of these objects are thus informative in sound classification, especially in few-shot scenarios. However, existing works have only represented audio semantics as binary labels (e.g., whethe…

Cited by 0SourceScholar
2023

CSTAR: Towards Compact and Structured Deep Neural Networks with Adversarial Robustness

AAAI 2023technical

Model compression and model defense for deep neural networks (DNNs) have been extensively and individually studied. Considering the co-importance of model compactness and robustness in practical applications, several prior works have explored to improve the adversarial robustness of the sparse neura…

Cited by 13SourcePDFScholar
2023

Cross-Modal Fusion Techniques for Utterance-Level Emotion Recognition from Text and Speech

ICASSP 2023accepted

Multimodal emotion recognition (MER) is a fundamental complex research problem due to the uncertainty of human emotional expression and the heterogeneity gap between different modalities. Audio and text modalities are particularly important for a human participant in understanding emotions. Although…

Cited by 0SourceScholar
2023

Improving Automatic Sleep Staging Via Temporal Smoothness Regularization

ICASSP 2023accepted

We propose a regularization method, so-called temporal smoothness regularization, for training deep neural networks for automatic sleep staging in small data settings. In intuition, we constrain the cross-entropy losses of any two adjacent epochs in the sequential input to be as close to each other…

Cited by 0SourceScholar
2023

Modelling Black-Box Audio Effects with Time-Varying Feature Modulation

ICASSP 2023accepted

Deep learning approaches for black-box modelling of audio effects have shown promise, however, the majority of existing work focuses on nonlinear effects with behaviour on relatively short time-scales, such as guitar amplifiers and distortion. While recurrent and convolutional architectures can theo…

Cited by 24SourceScholar
2022

BATUDE: Budget-Aware Neural Network Compression Based on Tucker Decomposition

AAAI 2022technical

Model compression is very important for the efficient deployment of deep neural network (DNN) models on resource-constrained devices. Among various model compression approaches, high-order tensor decomposition is particularly attractive and useful because the decomposed model is very small and fully…

Cited by 30SourcePDFScholar
2022

Invisible and Efficient Backdoor Attacks for Compressed Deep Neural Networks

ICASSP 2022accepted

Compressed deep neural network (DNN) models have been widely deployed in many resource-constrained platforms and devices. However, the security issue of the compressed models, especially their vulnerability against backdoor attacks, is not well explored yet. In this paper, we study the feasibility o…

Cited by 0SourceScholar
2022

Polyphonic Audio Event Detection: Multi-Label or Multi-Class Multi-Task Classification Problem?

ICASSP 2022accepted

Polyphonic events are the main error source of audio event detection (AED) systems. In deep-learning context, the most common approach to deal with event overlaps is to treat the AED task as a multi-label classification problem. By doing this, we inherently consider multiple one-vs.-rest classificat…

Cited by 0SourceScholar
2022

RIBAC: Towards Robust and Imperceptible Backdoor Attack against Compact DNN

ECCV 2022poster

"Recently backdoor attack has become an emerging threat to the security of deep neural network (DNN) models. To date, most of the existing studies focus on backdoor attack against the uncompressed model; while the vulnerability of compressed DNNs, which are widely used in the practical applications,…

2022

SALSA-Lite: A Fast and Effective Feature for Polyphonic Sound Event Localization and Detection with Microphone Arrays

ICASSP 2022accepted

Polyphonic sound event localization and detection (SELD) has many practical applications in acoustic sensing and monitoring. However, the development of real-time SELD has been limited by the demanding computational requirement of most recent SELD systems. In this work, we introduce SALSA-Lite, a fa…

Cited by 0SourceScholar
2021

A General Network Architecture for Sound Event Localization and Detection Using Transfer Learning and Recurrent Neural Network

ICASSP 2021accepted

Polyphonic sound event detection and localization (SELD) task is challenging because it is difficult to jointly optimize sound event detection (SED) and direction-of-arrival (DOA) estimation in the same network. We propose a general network architecture for SELD in which the SELD network comprises s…

Cited by 0SourceScholar
2021

CHIP: CHannel Independence-based Pruning for Compact Neural Networks

NeurIPS 2021poster

Filter pruning has been widely used for neural network compression because of its enabled practical acceleration. To date, most of the existing filter pruning works explore the importance of filters via using intra-channel information. In this paper, starting from an inter-channel perspective, we pr…

2021

Multi-View Audio And Music Classification

ICASSP 2021accepted

We propose in this work a multi-view learning approach for audio and music classification. Considering four typical low-level representations (i.e. different views) commonly used for audio and music recognition tasks, the proposed multi-view network consists of four subnetworks, each handling one in…

Cited by 19SourceScholar
2021

Self-Attention Generative Adversarial Network for Speech Enhancement

ICASSP 2021accepted

Existing generative adversarial networks (GANs) for speech enhancement solely rely on the convolution operation, which may obscure temporal dependencies across the sequence input. To remedy this issue, we propose a self-attention layer adapted from non-local attention, coupled with the convolutional…

Cited by 0SourceScholar
2019

Forked Recurrent Neural Network for Hand Gesture Classification Using Inertial Measurement Data

ICASSP 2019accepted

For many applications of hand gesture recognition, a delay-free, affordable, and mobile system relying on body signals is mandatory. Therefore, we propose an approach for hand gestures classification given signals of inertial measurement units (IMUs) that works with extremely short windows to avoid…

Cited by 4SourceScholar
2019

Unifying Isolated and Overlapping Audio Event Detection with Multi-label Multi-task Convolutional Recurrent Neural Networks

ICASSP 2019accepted

We propose a multi-label multi-task framework based on a convolutional recurrent neural network to unify detection of isolated and overlapping audio events. The framework leverages the power of convolutional recurrent neural network architectures; convolutional layers learn effective features over w…

Cited by 22SourceScholar
2018

Weighted and Multi-Task Loss for Rare Audio Event Detection

ICASSP 2018accepted

We present in this paper two loss functions tailored for rare audio event detection in audio streams. The weighted loss is designed to tackle the common issue of imbalanced data in background/foreground classification while the multi-task loss enables the networks to simultaneously model the class d…

Cited by 0SourceScholar
2017

CNN-LTE: A class of 1-X pooling convolutional neural networks on label tree embeddings for audio scene classification

ICASSP 2017accepted

We present in this work an approach for audio scene classification. Firstly, given the label set of the scenes, a label tree is automatically constructed where the labels are grouped into meta-classes. This category taxonomy is then used in the feature extraction step in which an audio scene instanc…

Cited by 0SourceScholar
2016

Learning compact structural representations for audio events using regressor banks

ICASSP 2016accepted

We introduce a new learned descriptor for audio signals which is efficient for event representation. The entries of the descriptor are produced by evaluating a set of regressors on the input signal. The regressors are class-specific and trained using the random regression forests framework. Given an…

Cited by 0SourceScholar