← Search

Jaesung Huh

8 accepted papers

2025

Advancing Active Speaker Detection for Egocentric Videos

ICASSP 2025accepted

This paper presents an improved approach to multimodal active speaker detection in egocentric videos, specifically designed to be robust against the rapid movements and motion blur commonly found in such videos. We propose two key techniques to improve the model’s resilience: (i) spatially fixing th…

Cited by 0SourceScholar
2024

TIM: A Time Interval Machine for Audio-Visual Action Recognition

CVPR 2024poster

Diverse actions give rise to rich audio-visual signals in long videos. Recent works showcase that the two modalities of audio and video exhibit different temporal extents of events and distinct labels. We address the interplay between the two modalities in long videos by explicitly modelling the tem…

2023

Epic-Sounds: A Large-Scale Dataset of Actions that Sound

ICASSP 2023accepted

We introduce EPIC-SOUNDS, a large-scale dataset of audio annotations capturing temporal extents and class labels within the audio stream of the egocentric videos from EPIC-KITCHENS-100. We propose an annotation pipeline where annotators temporally label distinguishable audio segments and describe th…

Cited by 0SourceScholar
2023

In Search of Strong Embedding Extractors for Speaker Diarisation

ICASSP 2023accepted

Speaker embedding extractors (EEs), which map input audio to a speaker discriminant latent space, are of paramount importance in speaker diarisation. However, there are several challenges when adopting EEs for diarisation, from which we tackle two key problems. First, the evaluation is not straightf…

Cited by 0SourceScholar
2021

Playing a Part: Speaker Verification at the movies

ICASSP 2021accepted

The goal of this work is to investigate the performance of popular speaker recognition models on speech segments from movies, where often actors intentionally disguise their voice to play a character. We make the following three contributions: (i) We collect a novel, challenging speaker recognition…

Cited by 0SourceScholar
2020

The Sound of My Voice: Speaker Representation Loss for Target Voice Separation

ICASSP 2020accepted

Content and style representations have been widely studied in the field of style transfer. In this paper, we propose a new loss function using speaker content representation for audio source separation, and we call it speaker representation loss. The objective is to extract the target speaker voice…

Cited by 0SourceScholar
2019

Phase-Aware Speech Enhancement with Deep Complex U-Net

ICLR 2019poster

Most deep learning-based models for speech enhancement have mainly focused on estimating the magnitude of spectrogram while reusing the phase from noisy speech for reconstruction. This is due to the difficulty of estimating the phase of clean speech. To improve speech enhancement performance, we tac…

Cited by 476SourceScholar