← Search

Yuanbo Hou

8 accepted papers

2026

LEARNING DOMAIN-ROBUST BIOACOUSTIC REPRESENTATIONS FOR MOSQUITO SPECIES CLASSIFICATION WITH CONTRASTIVE LEARNING AND DISTRIBUTION ALIGNMENT

ICASSP 2026poster

Mosquito Species Classification (MSC) is crucial for vector surveillance and disease control. The collection of mosquito bioacoustic data is often limited by mosquito activity seasons and fieldwork. Mosquito recordings across regions, habitats, and laboratories often show non-biological variations f…

Cited by 0SourcePDFScholar
2025

Sound-Based Recognition of Touch Gestures and Emotions for Enhanced Human-Robot Interaction

ICASSP 2025accepted

Emotion recognition and touch gesture decoding are crucial for advancing human-robot interaction (HRI), especially in social environments where emotional cues and tactile perception play important roles. However, many humanoid robots, such as Pepper, Nao, and Furhat, lack full-body tactile skin, lim…

Cited by 0SourceScholar
2024

Boosting Adversarial Transferability across Model Genus by Deformation-Constrained Warping

AAAI 2024technical

Adversarial examples generated by a surrogate model typically exhibit limited transferability to unknown target systems. To address this problem, many transferability enhancement approaches (e.g., input transformation and model augmentation) have been proposed. However, they show poor performances i…

2024

Multi-Level Graph Learning For Audio Event Classification And Human-Perceived Annoyance Rating Prediction

ICASSP 2024accepted

WHO’s report on environmental noise estimates that 22 M people suffer from chronic annoyance related to noise caused by audio events (AEs) from various sources. Annoyance may lead to health issues and adverse effects on metabolic and cognitive systems. In cities, monitoring noise levels does not pro…

Cited by 0SourceScholar
2024

No More Mumbles: Enhancing Robot Intelligibility Through Speech Adaptation

RA-L 2024

Spoken language interaction is at the heart of interpersonal communication, and people flexibly adapt their speech to different individuals and environments. It is surprising that robots, and by extension other digital devices, are not equipped to adapt their speech and instead rely on fixed speech

Cited by 7SourcecodeScholar
2021

Rule-Embedded Network for Audio-Visual Voice Activity Detection in Live Musical Video Streams

ICASSP 2021accepted

Detecting anchor’s voice in live musical streams is an important preprocessing step for music and speech signal processing. Existing approaches to voice activity detection (VAD) primarily rely on audio, however, audio-based VAD is difficult to effectively focus on the target voice in noisy environme…

Cited by 0SourceScholar
2019

Sound Event Detection with Sequentially Labelled Data Based on Connectionist Temporal Classification and Unsupervised Clustering

ICASSP 2019accepted

Sound event detection (SED) methods typically rely on either strongly labelled data or weakly labelled data. As an alternative, sequentially labelled data (SLD) was proposed. In SLD, the events and the order of events in audio clips are known, without knowing the occurrence time of events. This pape…

Cited by 0SourceScholar