← Search

Huaming Wang

9 accepted papers

2024

Natural Language Supervision For General-Purpose Audio Representations

ICASSP 2024accepted

Audio-Language models jointly learn multimodal text and audio representations that enable Zero-Shot inference. Models rely on the encoders to create powerful representations of the input and generalize to multiple tasks ranging from sounds, music, and speech. Although models have achieved remarkable…

Cited by 0SourceScholar
2024

Prompting Audios Using Acoustic Properties for Emotion Representation

ICASSP 2024accepted

Emotions lie on a continuum, but current models treat emotions as a finite valued discrete variable. This representation does not capture the diversity in the expression of emotion. To better represent emotions we propose the use of natural language descriptions (or prompts). In this work, we addres…

Cited by 0SourceScholar
2024

Training Audio Captioning Models without Audio

ICASSP 2024accepted

Automated Audio Captioning (AAC) is the task of generating natural language descriptions given an audio stream. A typical AAC system requires manually curated training data of audio segments and corresponding text caption annotations. The creation of these audio-caption pairs is costly, resulting in…

Cited by 0SourceScholar
2023

CLAP Learning Audio Concepts from Natural Language Supervision

ICASSP 2023accepted

Mainstream machine listening models are trained to learn audio concepts under the paradigm of one class label to many recordings focusing on one task. Learning under such restricted supervision limits the flexibility of models because they require labeled audio for training and can only predict the…

Cited by 0SourceScholar
2023

Pengi: An Audio Language Model for Audio Tasks

NeurIPS 2023poster

In the domain of audio processing, Transfer Learning has facilitated the rise of Self-Supervised Learning and Zero-Shot Learning techniques. These approaches have led to the development of versatile models capable of tackling a wide array of tasks, while delivering state-of-the-art performance. Howe…

2023

Real-Time Audio-Visual End-To-End Speech Enhancement

ICASSP 2023accepted

Audio-visual speech enhancement (AV-SE) methods utilize auxiliary visual cues to enhance speakers’ voices. Therefore, technically they should be able to outperform the audio-only speech enhancement (SE) methods. However, there are few works in the literature on an AV-SE system that can work in real…

Cited by 0SourceScholar
2022

One Model to Enhance Them All: Array Geometry Agnostic Multi-Channel Personalized Speech Enhancement

ICASSP 2022accepted

With the recent surge of video conferencing tools usage, providing high-quality speech signals and accurate captions have become essential to conduct day-to-day business or connect with friends and families. Single-channel personalized speech enhancement (PSE) methods show promising results compared…

Cited by 0SourceScholar
2022

Personalized speech enhancement: new models and Comprehensive evaluation

ICASSP 2022accepted

Personalized speech enhancement (PSE) models utilize additional cues, such as speaker embeddings like d-vectors, to remove background noise and interfering speech in real-time and thus improve the speech quality of online video conferencing systems for various acoustic scenarios. In this work, we pr…

Cited by 0SourceScholar
2018

Attention-Aware Compositional Network for Person Re-Identification

CVPR 2018poster

Person re-identification (ReID) is to identify pedestrians observed from different camera views based on visual appearance. It is a challenging task due to large pose variations, complex background clutters and severe occlusions. Recently, human pose estimation by predicting joint locations was larg…

Cited by 565SourcePDFScholar