← Search

Youngjoon Jang

13 accepted papers

2026

Improving Semantic Proximity in English-Centric Information Retrieval through Cross-Lingual Alignment

ICLR 2026poster

With the increasing accessibility and utilization of multilingual documents, Cross-Lingual Information Retrieval (CLIR) has emerged as an important research area. Conventionally, CLIR tasks have been conducted under settings where the language of documents differs from that of queries, and typically…

Cited by 0SourceScholar
2025

AVCD: Mitigating Hallucinations in Audio-Visual Large Language Models through Contrastive Decoding

NeurIPS 2025poster

Hallucination remains a major challenge in multimodal large language models (MLLMs). To address this, various contrastive decoding (CD) methods have been proposed that contrasts original logits with hallucinated logits generated from perturbed inputs. While CD has shown promise in vision-language mo…

Cited by 0SourcecodeScholar
2025

Lost in Translation, Found in Context: Sign Language Translation with Contextual Cues

CVPR 2025poster

Our objective is to translate continuous sign language into spoken language text. Inspired by the way human interpreters rely on context for accurate translation, we incorporate additional contextual cues together with the signing video, into a new translation framework. Specifically, besides visual…

Cited by 3SourcePDFScholar
2025

VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis

ICASSP 2025accepted

We present VoiceDiT, a multi-modal generative model for producing environment-aware speech and audio from text and visual prompts. While aligning speech with text is crucial for intelligible speech, achieving this alignment in noisy conditions remains a significant and underexplored challenge in the…

Cited by 0SourceScholar
2024

Faces that Speak: Jointly Synthesising Talking Face and Speech from Text

CVPR 2024poster

The goal of this work is to simultaneously generate natural talking faces and speech outputs from text. We achieve this by integrating Talking Face Generation (TFG) and Text-to-Speech (TTS) systems into a unified framework. We address the main challenges of each task: (1) generating a range of head…

Cited by 11SourcePDFScholar
2024

Fregrad: Lightweight and Fast Frequency-Aware Diffusion Vocoder

ICASSP 2024accepted

The goal of this paper is to generate realistic audio with a lightweight and fast diffusion-based vocoder, named FreGrad. Our framework consists of the following three key components: (1) We employ discrete wavelet transform that decomposes a complicated waveform into sub-band wavelets, which helps…

Cited by 0SourceScholar
2024

Seeing Through The Conversation: Audio-Visual Speech Separation Based on Diffusion Model

ICASSP 2024accepted

The objective of this work is to extract the target speaker’s voice from a mixture of voices using visual cues. Existing works on audio-visual speech separation have demonstrated their performance with promising intelligibility, but maintaining naturalness remains challenging. To address this issue,…

Cited by 0SourceScholar
2024

TalkNCE: Improving Active Speaker Detection with Talk-Aware Contrastive Learning

ICASSP 2024accepted

The goal of this work is Active Speaker Detection (ASD), a task to determine whether a person is speaking or not in a series of video frames. Previous works have dealt with the task by exploring network architectures while learning effective representations has been less explored. In this work, we p…

Cited by 0SourceScholar
2024

VoxMM: Rich Transcription of Conversations in the Wild

ICASSP 2024accepted

This paper presents a multi-modal dataset that contains rich transcriptions of spoken conversations. As diverse multi-modal and multi-task models emerge, there is a growing need for multi-modal training and evaluation datasets accompanied by rich metadata. However, there is no universal dataset that…

Cited by 0SourceScholar
2024

Where am I? Large Language Models Wandering between Semantics and Structures in Long Contexts

EMNLP 2024main

As the utilization of Large Language Models (LLMs) becomes more widespread, there is a growing demand for their ability to handle more complex and longer external knowledge across various use cases. Most existing evaluations of the open-ended question answering (ODQA) task, which necessitates the us…

2023

Metric Learning for User-Defined Keyword Spotting

ICASSP 2023accepted

The goal of this work is to detect new spoken terms defined by users. While most previous works address Keyword Spotting (KWS) as a closed-set classification problem, this limits their transferability to unseen terms. The ability to define custom keywords has advantages in terms of user experience.I…

Cited by 0SourceScholar
2023

Self-Sufficient Framework for Continuous Sign Language Recognition

ICASSP 2023accepted

The goal of this work is to develop self-sufficient framework for Continuous Sign Language Recognition (CSLR) that addresses key issues of sign language recognition. These include the need for complex multi-scale features such as hands, face, and mouth for understanding, and absence of frame-level a…

Cited by 0SourceScholar