← Search

Ho-Hsiang Wu

8 accepted papers

2024

CLAP4Emo: ChatGPT-Assisted Speech Emotion Retrieval with Natural Language Supervision

ICASSP 2024accepted

Speech emotion retrieval is an important technique for large-scale and high-quality data collection. Conventional approach using ensemble of classification models might limit the retrieved emotion diversity and/or underperform in out-of-domain acoustic conditions. Natural language is diverse and agn…

Cited by 6SourceScholar
2024

MOSAIC: Learning Unified Multi-Sensory Object Property Representations for Robot Learning via Interactive Perception

ICRA 2024poster

A holistic understanding of object properties across diverse sensory modalities (e.g., visual, audio, and haptic) is essential for tasks ranging from object categorization to complex manipulation. Drawing inspiration from cognitive science studies that emphasize the significance of multi-sensory int…

Cited by 2SourcecodeScholar
2022

Wav2CLIP: Learning Robust Audio Representations from Clip

ICASSP 2022accepted

We propose Wav2CLIP, a robust audio representation learning method by distilling from Contrastive Language-Image Pre-training (CLIP). We systematically evaluate Wav2CLIP on a variety of audio tasks including classification, retrieval, and generation, and show that Wav2CLIP can outperform several pub…

Cited by 0SourceScholar
2021

Multi-Task Self-Supervised Pre-Training for Music Classification

ICASSP 2021accepted

Deep learning is very data hungry, and supervised learning especially requires massive labeled data to work well. Machine listening research often suffers from limited labeled data problem, as human annotations are costly to acquire, and annotations for audio are time consuming and less intuitive. B…

Cited by 0SourceScholar
2019

Look, Listen, and Learn More: Design Choices for Deep Audio Embeddings

ICASSP 2019accepted

A considerable challenge in applying deep learning to audio classification is the scarcity of labeled data. An increasingly popular solution is to learn deep audio embeddings from large audio collections and use them to train shallow classifiers using small labeled datasets. Look, Listen, and Learn…

Cited by 0SourceScholar