← Search

Joon Son Chung

59 accepted papers

2026

DIFFUSION-LINK: DIFFUSION PROBABILISTIC MODEL FOR BRIDGING THE AUDIO-TEXT MODALITY GAP

ICASSP 2026poster

Contrastive audio-language pretraining yields powerful joint representations, yet a persistent audio-text modality gap limits the benefits of coupling multimodal encoders with large language models (LLMs). We present Diffusion-Link, a diffusion-based modality-bridging module that generatively maps a…

Cited by 0SourcePDFScholar
2026

How Far Can We Go With Synthetic Data for Audio-Visual Sound Source Localization?

CVPR 2026

We present the first scalable framework for training sound source localization (SSL) models using synthetic data from text-to-X models. Although SSL has made notable progress, existing models remain constrained by limited-scale, uncurated real-world datasets that often suffer from semantic misalignm

Cited by 0SourceScholar
2026

MAGE: A COARSE-TO-FINE SPEECH ENHANCER WITH MASKED GENERATIVE MODEL

ICASSP 2026poster

Speech enhancement remains challenging due to the trade-off between efficiency and perceptual quality. In this paper, we introduce MAGE, a Masked Audio Generative Enhancer that advances generative speech enhancement through a compact and robust design. Unlike prior masked generative models with rand…

Cited by 0SourcePDFScholar
2026

MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence

AAAI 2026technical

Audio comprehension—including speech, non-speech sounds, and music—is essential for achieving human-level intelligence. Consequently, AI agents must demonstrate holistic audio understanding to qualify as generally intelligent. However, evaluating auditory intelligence comprehensively remains challen

Cited by 0SourcePDFScholar
2026

SPADE: STRUCTURED PRUNING AND ADAPTIVE DISTILLATION FOR EFFICIENT LLM-TTS

ICASSP 2026oral

The goal of this paper is to introduce SPADE, a framework for Structured Pruning and Adaptive Distillation for Efficient Large Language Model-based text-to-speech (LLM-TTS). Recent LLM-TTS systems achieve strong controllability and zero-shot generalization, but their large parameter counts and high…

Cited by 0SourcePDFScholar
2026

Seeing Through Touch: Tactile-Driven Visual Localization of Material Regions

CVPR 2026

We address the problem of tactile localization, where the goal is to identify image regions that share the same material properties as a tactile input. Existing visuo-tactile methods rely on global alignment and thus fail to capture the fine-grained local correspondences required for this task. The

Cited by 0SourcecodeScholar
2025

AVCD: Mitigating Hallucinations in Audio-Visual Large Language Models through Contrastive Decoding

NeurIPS 2025poster

Hallucination remains a major challenge in multimodal large language models (MLLMs). To address this, various contrastive decoding (CD) methods have been proposed that contrasts original logits with hallucinated logits generated from perturbed inputs. While CD has shown promise in vision-language mo…

Cited by 0SourcecodeScholar
2025

AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

ICLR 2025poster

Following the success of Large Language Models (LLMs), expanding their boundaries to new modalities represents a significant paradigm shift in multimodal understanding. Human perception is inherently multimodal, relying not only on text but also on auditory and visual cues for a complete understandi…

2025

Accelerating Codec-based Speech Synthesis with Multi-Token Prediction and Speculative Decoding

ICASSP 2025accepted

The goal of this paper is to accelerate codec-based speech synthesis systems with minimum sacrifice to speech quality. We propose an enhanced inference method that allows for flexible trade-offs between speed and quality during inference without requiring additional training. Our core idea is to pre…

Cited by 0SourceScholar
2025

AdaptVC: High Quality Voice Conversion with Adaptive Learning

ICASSP 2025accepted

The goal of voice conversion is to transform the speech of a source speaker to sound like that of a reference speaker while preserving the original content. A key challenge is to extract disentangled linguistic content from the source and voice style from the reference. While existing approaches lev…

Cited by 5SourceScholar
2025

From Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-Speech

CVPR 2025highlight

The objective of this study is to generate high-quality speech from silent talking face videos, a task also known as video-to-speech synthesis. A significant challenge in video-to-speech synthesis lies in the substantial modality gap between silent video and multi-faceted speech. In this paper, we…

2025

High-Quality Joint Image and Video Tokenization with Causal VAE

ICLR 2025poster

Generative modeling has seen significant advancements in image and video synthesis. However, the curse of dimensionality remains a significant obstacle, especially for video generation, given its inherently complex and high-dimensional nature. Many existing works rely on low-dimensional latent space…

Cited by 1SourcePDFScholar
2025

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport

ICASSP 2025accepted

Automated audio captioning is a task that generates textual descriptions for audio content, and recent studies have explored using visual information to enhance captioning quality. However, current methods often fail to effectively fuse audio and visual data, missing important semantic cues from eac…

Cited by 0SourceScholar
2025

Model-Guided Dual-Role Alignment for High-Fidelity Open-Domain Video-to-Audio Generation

NeurIPS 2025poster

We present MGAudio, a novel flow-based framework for open-domain video-to-audio generation, which introduces model-guided dual-role alignment as a central design principle. Unlike prior approaches that rely on classifier-based or classifier-free guidance, MGAudio enables the generative model to guid…

Cited by 0SourcecodeScholar
2025

Seeing Speech and Sound: Distinguishing and Locating Audio Sources in Visual Scenes

CVPR 2025poster

We present a unified model capable of simultaneously grounding both spoken language and non-speech sounds within a visual scene, addressing key limitations in current audio-visual grounding models. Existing approaches are typically limited to handling either speech or non-speech sounds independently…

Cited by 0SourcePDFScholar
2025

V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow

ICASSP 2025accepted

In this paper, we introduce V2SFlow, a novel Video-to-Speech (V2S) framework designed to generate natural and intelligible speech directly from silent talking face videos. While recent V2S systems have shown promising results on constrained datasets with limited speakers and vocabularies, their perf…

Cited by 0SourceScholar
2025

Video Diffusion Models Excel at Tracking Similar-Looking Objects Without Supervision

NeurIPS 2025poster

Distinguishing visually similar objects by their motion remains a critical challenge in computer vision. Although supervised trackers show promise, contemporary self-supervised trackers struggle when visual cues become ambiguous, limiting their scalability and generalization without extensive labele…

Cited by 0SourceScholar
2025

VoiceCraft-Dub: Automated Video Dubbing with Neural Codec Language Models

ICCV 2025poster

We present VoiceCraft-Dub, a novel approach for automated video dubbing that synthesizes high-quality speech from text and facial cues. This task has broad applications in filmmaking, multimedia creation, and assisting voice-impaired individuals. Building on the success of Neural Codec Language Mode…

Cited by 0SourcePDFScholar
2025

VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis

ICASSP 2025accepted

We present VoiceDiT, a multi-modal generative model for producing environment-aware speech and audio from text and visual prompts. While aligning speech with text is crucial for intelligible speech, achieving this alignment in noisy conditions remains a significant and underexplored challenge in the…

Cited by 0SourceScholar
2024

EquiAV: Leveraging Equivariance for Audio-Visual Contrastive Learning

ICML 2024poster

Recent advancements in self-supervised audio-visual representation learning have demonstrated its potential to capture rich and comprehensive representations. However, despite the advantages of data augmentation verified in many learning methods, audio-visual learning has struggled to fully harness…

2024

Faces that Speak: Jointly Synthesising Talking Face and Speech from Text

CVPR 2024poster

The goal of this work is to simultaneously generate natural talking faces and speech outputs from text. We achieve this by integrating Talking Face Generation (TFG) and Text-to-Speech (TTS) systems into a unified framework. We address the main challenges of each task: (1) generating a range of head…

Cited by 11SourcePDFScholar
2024

Fregrad: Lightweight and Fast Frequency-Aware Diffusion Vocoder

ICASSP 2024accepted

The goal of this paper is to generate realistic audio with a lightweight and fast diffusion-based vocoder, named FreGrad. Our framework consists of the following three key components: (1) We employ discrete wavelet transform that decomposes a complicated waveform into sub-band wavelets, which helps…

Cited by 0SourceScholar
2024

From Coarse to Fine: Efficient Training for Audio Spectrogram Transformers

ICASSP 2024accepted

Transformers have become central to recent advances in audio classification. However, training an audio spectrogram transformer, e.g. AST, from scratch can be resource and time-intensive. Furthermore, the complexity of transformers heavily depends on the input audio spectrogram size. In this work, w…

Cited by 0SourceScholar
2024

Let There Be Sound: Reconstructing High Quality Speech from Silent Videos

AAAI 2024technical

The goal of this work is to reconstruct high quality speech from lip motions alone, a task also known as lip-to-speech. A key challenge of lip-to-speech systems is the one-to-many mapping caused by (1) the existence of homophenes and (2) multiple speech variations, resulting in a mispronounced and o…

2024

Rethinking Session Variability: Leveraging Session Embeddings for Session Robustness in Speaker Verification

ICASSP 2024accepted

In the field of speaker verification, session or channel variability poses a significant challenge. While many contemporary methods aim to disentangle session information from speaker embeddings, we introduce a novel approach using an additional embedding to represent the session information. This i…

Cited by 0SourceScholar
2024

Scaling Up Video Summarization Pretraining with Large Language Models

CVPR 2024poster

Long-form video content constitutes a significant portion of internet traffic making automated video summarization an essential research problem. However existing video summarization datasets are notably limited in their size constraining the effectiveness of state-of-the-art methods for generalizat…

Cited by 13SourcePDFScholar
2024

Seeing Through The Conversation: Audio-Visual Speech Separation Based on Diffusion Model

ICASSP 2024accepted

The objective of this work is to extract the target speaker’s voice from a mixture of voices using visual cues. Existing works on audio-visual speech separation have demonstrated their performance with promising intelligibility, but maintaining naturalness remains challenging. To address this issue,…

Cited by 0SourceScholar
2024

Speech Guided Masked Image Modeling for Visually Grounded Speech

ICASSP 2024accepted

The objective of this study is to investigate the learning process of Visually Grounded Speech (VGS) models through joint learning that combines contrastive learning and masked image modeling. Typically, VGS models ahn to establish audio-visual alignment between images and then spoken captions withi…

Cited by 0SourceScholar
2024

TalkNCE: Improving Active Speaker Detection with Talk-Aware Contrastive Learning

ICASSP 2024accepted

The goal of this work is Active Speaker Detection (ASD), a task to determine whether a person is speaking or not in a series of video frames. Previous works have dealt with the task by exploring network architectures while learning effective representations has been less explored. In this work, we p…

Cited by 0SourceScholar
2024

Towards Automated Movie Trailer Generation

CVPR 2024poster

Movie trailers are an essential tool for promoting films and attracting audiences. However the process of creating trailers can be time-consuming and expensive. To streamline this process we propose an automatic trailer generation framework that generates plausible trailers from a full movie by auto…

Cited by 3SourcePDFScholar
2024

VoxMM: Rich Transcription of Conversations in the Wild

ICASSP 2024accepted

This paper presents a multi-modal dataset that contains rich transcriptions of spoken conversations. As diverse multi-modal and multi-task models emerge, there is a growing need for multi-modal training and evaluation datasets accompanied by rich metadata. However, there is no universal dataset that…

Cited by 0SourceScholar
2023

Advancing the Dimensionality Reduction of Speaker Embeddings for Speaker Diarisation: Disentangling Noise and Informing Speech Activity

ICASSP 2023accepted

The objective of this work is to train noise-robust speaker embeddings adapted for speaker diarisation. Speaker embeddings play a crucial role in the performance of diarisation systems, but they often capture spurious information such as noise, adversely affecting performance. Our previous work has…

Cited by 0SourceScholar
2023

Hindi as a Second Language: Improving Visually Grounded Speech with Semantically Similar Samples

ICASSP 2023accepted

The objective of this work is to explore the learning of visually grounded speech models (VGS) from multilingual perspective. Bilingual VGS models are generally trained with an equal number of spoken captions from both languages. However, in reality, there can be an imbalance among the languages for…

Cited by 0SourceScholar
2023

In Search of Strong Embedding Extractors for Speaker Diarisation

ICASSP 2023accepted

Speaker embedding extractors (EEs), which map input audio to a speaker discriminant latent space, are of paramount importance in speaker diarisation. However, there are several challenges when adopting EEs for diarisation, from which we tackle two key problems. First, the evaluation is not straightf…

Cited by 0SourceScholar
2023

Metric Learning for User-Defined Keyword Spotting

ICASSP 2023accepted

The goal of this work is to detect new spoken terms defined by users. While most previous works address Keyword Spotting (KWS) as a closed-set classification problem, this limits their transferability to unseen terms. The ability to define custom keywords has advantages in terms of user experience.I…

Cited by 0SourceScholar
2023

Self-Sufficient Framework for Continuous Sign Language Recognition

ICASSP 2023accepted

The goal of this work is to develop self-sufficient framework for Continuous Sign Language Recognition (CSLR) that addresses key issues of sign language recognition. These include the need for complex multi-scale features such as hands, face, and mouth for understanding, and absence of frame-level a…

Cited by 0SourceScholar
2023

Sound Source Localization is All about Cross-Modal Alignment

ICCV 2023poster

Humans can easily perceive the direction of sound sources in a visual scene, termed sound source localization. Recent studies on learning-based sound source localization have mainly explored the problem from a localization perspective. However, prior arts and existing benchmarks do not account for…

Cited by 18PDFScholar
2022

AASIST: Audio Anti-Spoofing Using Integrated Spectro-Temporal Graph Attention Networks

ICASSP 2022accepted

Artefacts that differentiate spoofed from bona-fide utterances can reside in specific temporal or spectral intervals. Their reliable detection usually depends upon computationally demanding ensemble systems where each subsystem is tuned to some specific artefacts. We seek to develop an efficient, si…

Cited by 0SourceScholar
2022

Multi-Scale Speaker Embedding-Based Graph Attention Networks For Speaker Diarisation

ICASSP 2022accepted

The objective of this work is effective speaker diarisation using multi-scale speaker embeddings. Typically, there is a trade-off between the ability to recognise short speaker segments and the discriminative power of the embedding, according to the segment length used for embedding extraction. To t…

Cited by 0SourceScholar
2021

Playing a Part: Speaker Verification at the movies

ICASSP 2021accepted

The goal of this work is to investigate the performance of popular speaker recognition models on speech segments from movies, where often actors intentionally disguise their voice to play a character. We make the following three contributions: (i) We collect a novel, challenging speaker recognition…

Cited by 0SourceScholar
2021

The ins and outs of speaker recognition: lessons from VoxSRC 2020

ICASSP 2021accepted

The VoxCeleb Speaker Recognition Challenge (VoxSRC) at Interspeech 2020 offers a challenging evaluation for speaker recognition systems, which includes celebrities playing different parts in movies. The goal of this work is robust speaker recognition of utterances recorded in these challenging envir…

Cited by 0SourceScholar
2020

ASR is All You Need: Cross-Modal Distillation for Lip Reading

ICASSP 2020accepted

The goal of this work is to train strong models for visual speech recognition without requiring human annotated ground truth data. We achieve this by distilling from an Automatic Speech Recognition (ASR) model that has been trained on a large-scale audio-only corpus. We use a cross-modal distillatio…

Cited by 0SourceScholar
2020

BSL-1K: Scaling up co-articulated sign language recognition using mouthing cues

ECCV 2020poster

Recent progress in fine-grained gesture and action classification, and machine translation, point to the possibility of automated sign language recognition becoming a reality. A key stumbling block in making progress towards this goal is a lack of appropriate training data, stemming from the high co…

Cited by 221SourcePDFScholar
2020

Disentangled Speech Embeddings Using Cross-Modal Self-Supervision

ICASSP 2020accepted

The objective of this paper is to learn representations of speaker identity without access to manually annotated data. To do so, we develop a self-supervised learning objective that exploits the natural cross-modal synchrony between faces and audio in video. The key idea behind our approach is to te…

Cited by 0SourceScholar
2020

Self-Supervised Learning of Audio-Visual Objects from Video

ECCV 2020poster

Our objective is to transform a video into a set of discrete audio-visual objects using self-supervised learning. To this end we introduce a model that uses attention to localize and group sound sources, and optical flow to aggregate information over time. We demonstrate the effectiveness of the aud…

Cited by 313SourcePDFScholar
2020

The Sound of My Voice: Speaker Representation Loss for Target Voice Separation

ICASSP 2020accepted

Content and style representations have been widely studied in the field of style transfer. In this paper, we propose a new loss function using speaker content representation for audio source separation, and we call it speaker representation loss. The objective is to extract the target speaker voice…

Cited by 0SourceScholar
2019

Perfect Match: Improved Cross-modal Embeddings for Audio-visual Synchronisation

ICASSP 2019accepted

This paper proposes a new strategy for learning powerful cross-modal embeddings for audio-to-video synchronisation. Here, we set up the problem as one of cross-modal retrieval, where the objective is to find the most relevant audio segment given a short video clip. The method builds on the recent ad…

Cited by 0SourceScholar
2019

Utterance-level Aggregation for Speaker Recognition in the Wild

ICASSP 2019accepted

The objective of this paper is speaker recognition `in the wild' - where utterances may be of variable length and also contain irrelevant signals. Crucial elements in the design of deep networks for this task are the type of trunk (frame level) network, and the method of temporal aggregation. We pro…

Cited by 0SourceScholar