← Search

Chaeyoung Jung

7 accepted papers

2025

AVCD: Mitigating Hallucinations in Audio-Visual Large Language Models through Contrastive Decoding

NeurIPS 2025poster

Hallucination remains a major challenge in multimodal large language models (MLLMs). To address this, various contrastive decoding (CD) methods have been proposed that contrasts original logits with hallucinated logits generated from perturbed inputs. While CD has shown promise in vision-language mo…

Cited by 0SourcecodeScholar
2025

From Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-Speech

CVPR 2025highlight

The objective of this study is to generate high-quality speech from silent talking face videos, a task also known as video-to-speech synthesis. A significant challenge in video-to-speech synthesis lies in the substantial modality gap between silent video and multi-faceted speech. In this paper, we…

2025

VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis

ICASSP 2025accepted

We present VoiceDiT, a multi-modal generative model for producing environment-aware speech and audio from text and visual prompts. While aligning speech with text is crucial for intelligible speech, achieving this alignment in noisy conditions remains a significant and underexplored challenge in the…

Cited by 0SourceScholar
2024

Seeing Through The Conversation: Audio-Visual Speech Separation Based on Diffusion Model

ICASSP 2024accepted

The objective of this work is to extract the target speaker’s voice from a mixture of voices using visual cues. Existing works on audio-visual speech separation have demonstrated their performance with promising intelligibility, but maintaining naturalness remains challenging. To address this issue,…

Cited by 0SourceScholar
2024

TalkNCE: Improving Active Speaker Detection with Talk-Aware Contrastive Learning

ICASSP 2024accepted

The goal of this work is Active Speaker Detection (ASD), a task to determine whether a person is speaking or not in a series of video frames. Previous works have dealt with the task by exploring network architectures while learning effective representations has been less explored. In this work, we p…

Cited by 0SourceScholar