← Search

Hyeongseop Rha

4 accepted papers

2026

Emotion-Coherent Reasoning for Multimodal LLMs via Emotional Rationale Verifier

AAAI 2026technical

The recent advancement of Multimodal Large Language Models (MLLMs) is transforming human-computer interaction (HCI) from surface-level exchanges into more nuanced and emotionally intelligent communication. To realize this shift, emotion understanding becomes essential allowing systems to capture sub

Cited by 0SourcePDFScholar
2025

MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens

ACL 2025finding

Audio-Visual Speech Recognition (AVSR) achieves robust speech recognition in noisy environments by combining auditory and visual information. However, recent Large Language Model (LLM) based AVSR systems incur high computational costs due to the high temporal resolution of audio-visual speech proces…

2025

Personalized Lip Reading: Adapting to Your Unique Lip Movements with Vision and Language

AAAI 2025technical

Lip reading aims to predict spoken language by analyzing lip movements. Despite advancements in lip reading technologies, performance degrades when models are applied to unseen speakers due to their sensitivity to variations in visual information such as lip appearances. To address this challenge, s…

2024

Let’s Go Real Talk: Spoken Dialogue Model for Face-to-Face Conversation

ACL 2024long

In this paper, we introduce a novel Face-to-Face spoken dialogue model. It processes audio-visual speech from user input and generates audio-visual speech as the response, marking the initial step towards creating an avatar chatbot system without relying on intermediate text. To this end, we newly i…