← Search

Joanna Hong

9 accepted papers

2024

Let’s Go Real Talk: Spoken Dialogue Model for Face-to-Face Conversation

ACL 2024long

In this paper, we introduce a novel Face-to-Face spoken dialogue model. It processes audio-visual speech from user input and generates audio-visual speech as the response, marking the initial step towards creating an avatar chatbot system without relying on intermediate text. To this end, we newly i…

2023

DiffV2S: Diffusion-Based Video-to-Speech Synthesis with Vision-Guided Speaker Embedding

ICCV 2023poster

Recent research has demonstrated impressive results in video-to-speech synthesis which involves reconstructing speech solely from visual input. However, previous works have struggled to accurately synthesize speech due to a lack of sufficient guidance for the model to infer the correct content with…

Cited by 19PDFcodeScholar
2023

Intuitive Multilingual Audio-Visual Speech Recognition with a Single-Trained Model

EMNLP 2023short findings

We present a novel approach to multilingual audio-visual speech recognition tasks by introducing a single model on a multilingual dataset. Motivated by a human cognitive system where humans can intuitively distinguish different languages without any conscious effort or guidance, we propose a model t…

Cited by 0SourceScholar
2023

Watch or Listen: Robust Audio-Visual Speech Recognition With Visual Corruption Modeling and Reliability Scoring

CVPR 2023poster

This paper deals with Audio-Visual Speech Recognition (AVSR) under multimodal input corruption situation where audio inputs and visual inputs are both corrupted, which is not well addressed in previous research directions. Previous studies have focused on how to complement the corrupted audio inputs…

2022

SyncTalkFace: Talking Face Generation with Precise Lip-Syncing via Audio-Lip Memory

AAAI 2022technical

The challenge of talking face generation from speech lies in aligning two different modal information, audio and video, such that the mouth region corresponds to input audio. Previous methods either exploit audio-visual representation learning or leverage intermediate structural information such as…

Cited by 90SourcePDFScholar
2022

VisageSynTalk: Unseen Speaker Video-to-Speech Synthesis via Speech-Visage Feature Selection

ECCV 2022poster

"The goal of this work is to reconstruct speech from a silent talking face video. Recent studies have shown impressive performance on synthesizing speech from silent talking face videos. However, they have not explicitly considered on varying identity characteristics of different speakers, which pla…

Cited by 7SourcePDFScholar
2021

Multi-Modality Associative Bridging Through Memory: Speech Sound Recollected From Face Video

ICCV 2021poster

In this paper, we introduce a novel audio-visual multi-modal bridging framework that can utilize both audio and visual information, even with uni-modal inputs. We exploit a memory network that stores source (i.e., visual) and target (i.e., audio) modal representations, where source modal representat…

Cited by 52PDFScholar