← Search

Hyung-Min Park

4 accepted papers

2026

Query-Based Asymmetric Modeling with Decoupled Input–Output Rates for Speech Restoration

ICML 2026poster

Speech restoration in real-world conditions is challenging due to compounded distortions and mismatches between input and desired output rates. Most existing systems assume a fixed and shared input–output rate, relying on external resampling that incurs redundant computation and limits generality. W…

Cited by 0SourceScholar
2024

NeXt-TDNN: Modernizing Multi-Scale Temporal Convolution Backbone for Speaker Verification

ICASSP 2024accepted

In speaker verification, ECAPA-TDNN has shown remarkable improvement by utilizing one-dimensional(1D) Res2Net block and squeeze-and-excitation(SE) module, along with multi-layer feature aggregation (MFA). Meanwhile, in vision tasks, ConvNet structures have been modernized by referring to Transformer…

Cited by 0SourceScholar
2024

OLKAVS: An Open Large-Scale Korean Audio-Visual Speech Dataset

ICASSP 2024accepted

Inspired by humans comprehending speech in a multi-modal manner, various audio-visual datasets have been constructed. However, most existing datasets focus on English, developed from pre-existing videos using various prediction models, and have only a small number of multi-view videos. To mitigate t…

Cited by 0SourceScholar
2024

Separate and Reconstruct: Asymmetric Encoder-Decoder for Speech Separation

NeurIPS 2024poster

In speech separation, time-domain approaches have successfully replaced the time-frequency domain with latent sequence feature from a learnable encoder. Conventionally, the feature is separated into speaker-specific ones at the final stage of the network. Instead, we propose a more intuitive strateg…