← Search

Chae Won Kim

8 accepted papers

2025

Personalized Lip Reading: Adapting to Your Unique Lip Movements with Vision and Language

AAAI 2025technical

Lip reading aims to predict spoken language by analyzing lip movements. Despite advancements in lip reading technologies, performance degrades when models are applied to unseen speakers due to their sensitivity to variations in visual information such as lip appearances. To address this challenge, s…

2025

Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations

ICCV 2025poster

We explore a novel zero-shot Audio-Visual Speech Recognition (AVSR) framework, dubbed Zero-AVSR, which enables speech recognition in target languages without requiring any audio-visual speech data in those languages. Specifically, we introduce the Audio-Visual Speech Romanizer (AV-Romanizer), which…

2024

CoLLaVO: Crayon Large Language and Vision mOdel

ACL 2024findings

The remarkable success of Large Language Models (LLMs) and instruction tuning drives the evolution of Vision Language Models (VLMs) towards a versatile general-purpose model. Yet, it remains unexplored whether current VLMs genuinely possess quality object-level image understanding capabilities deter…

2024

Meteor: Mamba-based Traversal of Rationale for Large Language and Vision Models

NeurIPS 2024poster

The rapid development of large language and vision models (LLVMs) has been driven by advances in visual instruction tuning. Recently, open-source LLVMs have curated high-quality visual instruction tuning datasets and utilized additional vision encoders or multiple computer vision models in order to…

2024

MoAI: Mixture of All Intelligence for Large Language and Vision Models

ECCV 2024poster

"The rise of large language models (LLMs) and instruction tuning has led to the current trend of instruction-tuned large language and vision models (LLVMs). This trend involves either meticulously curating numerous instruction tuning datasets tailored to specific objectives or enlarging LLVMs to man…

2024

Persona Extraction Through Semantic Similarity for Emotional Support Conversation Generation

ICASSP 2024accepted

Providing emotional support through dialogue systems is becoming increasingly important in today’s world, as it can support both mental health and social interactions in many conversation scenarios. Previous works have shown that using persona is effective for generating empathetic and supportive re…

Cited by 0SourceScholar
2024

TroL: Traversal of Layers for Large Language and Vision Models

EMNLP 2024main

Large language and vision models (LLVMs) have been driven by the generalization power of large language models (LLMs) and the advent of visual instruction tuning. Along with scaling them up directly, these models enable LLVMs to showcase powerful vision language (VL) performances by covering diverse…

2023

Deep Visual Forced Alignment: Learning to Align Transcription with Talking Face Video

AAAI 2023technical

Forced alignment refers to a technology that time-aligns a given transcription with a corresponding speech. However, as the forced alignment technologies have developed using speech audio, they might fail in alignment when the input speech audio is noise-corrupted or is not accessible. We focus on t…

Cited by 3SourcePDFScholar