← Search

Se Jin Park

9 accepted papers

2025

Long-Form Speech Generation with Spoken Language Models

ICML 2025oral

We consider the generative modeling of speech over multiple minutes, a requirement for long-form multimedia generation and audio-native voice assistants. However, textless spoken language models struggle to generate plausible speech past tens of seconds, due to high temporal resolution of speech tok…

2025

MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens

ACL 2025finding

Audio-Visual Speech Recognition (AVSR) achieves robust speech recognition in noisy environments by combining auditory and visual information. However, recent Large Language Model (LLM) based AVSR systems incur high computational costs due to the high temporal resolution of audio-visual speech proces…

2024

AV2AV: Direct Audio-Visual Speech to Audio-Visual Speech Translation with Unified Audio-Visual Speech Representation

CVPR 2024highlight

This paper proposes a novel direct Audio-Visual Speech to Audio-Visual Speech Translation (AV2AV) framework where the input and output of the system are multimodal (i.e. audio and visual speech). With the proposed AV2AV two key advantages can be brought: 1) We can perform real-like conversations wit…

2024

Exploring Phonetic Context-Aware Lip-Sync for Talking Face Generation

ICASSP 2024accepted

Talking face generation is the challenging task of synthesizing a natural and realistic face that requires accurate synchronization with a given audio. Due to co-articulation, where an isolated phone is influenced by the preceding or following phones, the articulation of a phone varies upon the phon…

Cited by 0SourceScholar
2024

Persona Extraction Through Semantic Similarity for Emotional Support Conversation Generation

ICASSP 2024accepted

Providing emotional support through dialogue systems is becoming increasingly important in today’s world, as it can support both mental health and social interactions in many conversation scenarios. Previous works have shown that using persona is effective for generating empathetic and supportive re…

Cited by 0SourceScholar
2024

Text-Driven Talking Face Synthesis by Reprogramming Audio-Driven Models

ICASSP 2024accepted

In this paper, we present a method for reprogramming pre-trained audio-driven talking face synthesis models to operate in a text-driven manner. Consequently, we can easily generate face videos that articulate the provided textual sentences, eliminating the necessity of recording speech for each infe…

Cited by 0SourceScholar
2023

Intuitive Multilingual Audio-Visual Speech Recognition with a Single-Trained Model

EMNLP 2023short findings

We present a novel approach to multilingual audio-visual speech recognition tasks by introducing a single model on a multilingual dataset. Motivated by a human cognitive system where humans can intuitively distinguish different languages without any conscious effort or guidance, we propose a model t…

Cited by 0SourceScholar
2022

SyncTalkFace: Talking Face Generation with Precise Lip-Syncing via Audio-Lip Memory

AAAI 2022technical

The challenge of talking face generation from speech lies in aligning two different modal information, audio and video, such that the mouth region corresponds to input audio. Previous methods either exploit audio-visual representation learning or leverage intermediate structural information such as…

Cited by 90SourcePDFScholar
2021

Multi-Modality Associative Bridging Through Memory: Speech Sound Recollected From Face Video

ICCV 2021poster

In this paper, we introduce a novel audio-visual multi-modal bridging framework that can utilize both audio and visual information, even with uni-modal inputs. We exploit a memory network that stores source (i.e., visual) and target (i.e., audio) modal representations, where source modal representat…

Cited by 52PDFScholar