← Search

Yong Man Ro

60 accepted papers

2026

Emotion-Coherent Reasoning for Multimodal LLMs via Emotional Rationale Verifier

AAAI 2026technical

The recent advancement of Multimodal Large Language Models (MLLMs) is transforming human-computer interaction (HCI) from surface-level exchanges into more nuanced and emotionally intelligent communication. To realize this shift, emotion understanding becomes essential allowing systems to capture sub

Cited by 0SourcePDFScholar
2026

MAD: Modality-Adaptive Decoding for Mitigating Cross-Modal Hallucinations in Multimodal Large Language Models

CVPR 2026

Multimodal Large Language Models (MLLMs) suffer from cross-modal hallucinations, where one modality inappropriately influences generation about another, leading to fabricated output. This exposes a more fundamental deficiency in modality-interaction control. To address this, we propose Modality-Adap

Cited by 0SourcecodeScholar
2025

Long-Form Speech Generation with Spoken Language Models

ICML 2025oral

We consider the generative modeling of speech over multiple minutes, a requirement for long-form multimedia generation and audio-native voice assistants. However, textless spoken language models struggle to generate plausible speech past tens of seconds, due to high temporal resolution of speech tok…

2025

MMS-LLaMA: Efficient LLM-based Audio-Visual Speech Recognition with Minimal Multimodal Speech Tokens

ACL 2025finding

Audio-Visual Speech Recognition (AVSR) achieves robust speech recognition in noisy environments by combining auditory and visual information. However, recent Large Language Model (LLM) based AVSR systems incur high computational costs due to the high temporal resolution of audio-visual speech proces…

2025

Personalized Lip Reading: Adapting to Your Unique Lip Movements with Vision and Language

AAAI 2025technical

Lip reading aims to predict spoken language by analyzing lip movements. Despite advancements in lip reading technologies, performance degrades when models are applied to unseen speakers due to their sensitivity to variations in visual information such as lip appearances. To address this challenge, s…

2025

SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis

CVPR 2025poster

Despite advances in Large Multi-modal Models, applying them to long and untrimmed video content remains challenging due to limitations in context length and substantial memory overhead. These constraints often lead to significant information loss and reduced relevance in the model responses. With th…

Cited by 3SourcePDFScholar
2025

Unified Reinforcement and Imitation Learning for Vision-Language Models

NeurIPS 2025poster

Vision-Language Models (VLMs) have achieved remarkable progress, yet their large scale often renders them impractical for resource-constrained environments. This paper introduces Unified Reinforcement and Imitation Learning (RIL), a novel and efficient training algorithm designed to create powerful,…

Cited by 0SourceScholar
2025

VLsI: Verbalized Layers-to-Interactions from Large to Small Vision Language Models

CVPR 2025poster

The recent surge in high-quality visual instruction tuning samples from closed-source vision-language models (VLMs) such as GPT-4V has accelerated the release of open-source VLMs across various model sizes. However, scaling VLMs to improve performance using larger models brings significant computati…

Cited by 0SourcePDFScholar
2025

Zero-AVSR: Zero-Shot Audio-Visual Speech Recognition with LLMs by Learning Language-Agnostic Speech Representations

ICCV 2025poster

We explore a novel zero-shot Audio-Visual Speech Recognition (AVSR) framework, dubbed Zero-AVSR, which enables speech recognition in target languages without requiring any audio-visual speech data in those languages. Specifically, we introduce the Audio-Visual Speech Romanizer (AV-Romanizer), which…

2024

AV2AV: Direct Audio-Visual Speech to Audio-Visual Speech Translation with Unified Audio-Visual Speech Representation

CVPR 2024highlight

This paper proposes a novel direct Audio-Visual Speech to Audio-Visual Speech Translation (AV2AV) framework where the input and output of the system are multimodal (i.e. audio and visual speech). With the proposed AV2AV two key advantages can be brought: 1) We can perform real-like conversations wit…

2024

CODE: Contrasting Self-generated Description to Combat Hallucination in Large Multi-modal Models

NeurIPS 2024poster

Large Multi-modal Models (LMMs) have recently demonstrated remarkable abilities in visual context understanding and coherent response generation. However, alongside these advancements, the issue of hallucinations has emerged as a significant challenge, producing erroneous responses that are unrelate…

Cited by 52SourcePDFScholar
2024

Causal Mode Multiplexer: A Novel Framework for Unbiased Multispectral Pedestrian Detection

CVPR 2024poster

RGBT multispectral pedestrian detection has emerged as a promising solution for safety-critical applications that require day/night operations. However the modality bias problem remains unsolved as multispectral pedestrian detectors learn the statistical bias in datasets. Specifically datasets in mu…

2024

CoLLaVO: Crayon Large Language and Vision mOdel

ACL 2024findings

The remarkable success of Large Language Models (LLMs) and instruction tuning drives the evolution of Vision Language Models (VLMs) towards a versatile general-purpose model. Yet, it remains unexplored whether current VLMs genuinely possess quality object-level image understanding capabilities deter…

2024

Exploring Phonetic Context-Aware Lip-Sync for Talking Face Generation

ICASSP 2024accepted

Talking face generation is the challenging task of synthesizing a natural and realistic face that requires accurate synchronization with a given audio. Due to co-articulation, where an isolated phone is influenced by the preceding or following phones, the articulation of a phone varies upon the phon…

Cited by 0SourceScholar
2024

Improving Open Set Recognition via Visual Prompts Distilled from Common-Sense Knowledge

AAAI 2024technical

Open Set Recognition (OSR) poses significant challenges in distinguishing known from unknown classes. In OSR, the overconfidence problem has become a persistent obstacle, where visual recognition models often misclassify unknown objects as known objects with high confidence. This issue stems from th…

Cited by 9SourcePDFScholar
2024

Meteor: Mamba-based Traversal of Rationale for Large Language and Vision Models

NeurIPS 2024poster

The rapid development of large language and vision models (LLVMs) has been driven by advances in visual instruction tuning. Recently, open-source LLVMs have curated high-quality visual instruction tuning datasets and utilized additional vision encoders or multiple computer vision models in order to…

2024

Persona Extraction Through Semantic Similarity for Emotional Support Conversation Generation

ICASSP 2024accepted

Providing emotional support through dialogue systems is becoming increasingly important in today’s world, as it can support both mental health and social interactions in many conversation scenarios. Previous works have shown that using persona is effective for generating empathetic and supportive re…

Cited by 0SourceScholar
2024

Text-Driven Talking Face Synthesis by Reprogramming Audio-Driven Models

ICASSP 2024accepted

In this paper, we present a method for reprogramming pre-trained audio-driven talking face synthesis models to operate in a text-driven manner. Consequently, we can easily generate face videos that articulate the provided textual sentences, eliminating the necessity of recording speech for each infe…

Cited by 0SourceScholar
2024

Towards Practical and Efficient Image-to-Speech Captioning with Vision-Language Pre-Training and Multi-Modal Tokens

ICASSP 2024accepted

In this paper, we propose methods to build a powerful and efficient Image-to-Speech captioning (Im2Sp) model. To this end, we start with importing the rich knowledge related to image comprehension and language modeling from a large-scale pre-trained vision-language model into Im2Sp. We set the outpu…

Cited by 0SourceScholar
2024

TroL: Traversal of Layers for Large Language and Vision Models

EMNLP 2024main

Large language and vision models (LLVMs) have been driven by the generalization power of large language models (LLMs) and the advent of visual instruction tuning. Along with scaling them up directly, these models enable LLVMs to showcase powerful vision language (VL) performances by covering diverse…

2024

Visual Speech Recognition for Languages with Limited Labeled Data Using Automatic Labels from Whisper

ICASSP 2024accepted

This paper proposes a powerful Visual Speech Recognition (VSR) method for multiple languages, especially for low-resource languages that have a limited number of labeled data. Different from previous methods that tried to improve the VSR performance for the target language by using knowledge learned…

Cited by 0SourceScholar
2024

What if...?: Thinking Counterfactual Keywords Helps to Mitigate Hallucination in Large Multi-modal Models

EMNLP 2024finding

This paper presents a way of enhancing the reliability of Large Multi-modal Models (LMMs) in addressing hallucination, where the models generate cross-modal inconsistent responses. Without additional training, we propose Counterfactual Inception, a novel method that implants counterfactual thinking…

2024

Where Visual Speech Meets Language: VSP-LLM Framework for Efficient and Context-Aware Visual Speech Processing

EMNLP 2024finding

In visual speech processing, context modeling capability is one of the most important requirements due to the ambiguous nature of lip movements. For example, homophenes, words that share identical lip movements but produce different sounds, can be distinguished by considering the context. In this pa…

2023

Deep Visual Forced Alignment: Learning to Align Transcription with Talking Face Video

AAAI 2023technical

Forced alignment refers to a technology that time-aligns a given transcription with a corresponding speech. However, as the forced alignment technologies have developed using speech audio, they might fail in alignment when the input speech audio is noise-corrupted or is not accessible. We focus on t…

Cited by 3SourcePDFScholar
2023

Demystifying Causal Features on Adversarial Examples and Causal Inoculation for Robust Network by Adversarial Instrumental Variable Regression

CVPR 2023poster

The origin of adversarial examples is still inexplicable in research fields, and it arouses arguments from various viewpoints, albeit comprehensive investigations. In this paper, we propose a way of delving into the unexpected vulnerability in adversarially trained networks from a causal perspective…

2023

DiffV2S: Diffusion-Based Video-to-Speech Synthesis with Vision-Guided Speaker Embedding

ICCV 2023poster

Recent research has demonstrated impressive results in video-to-speech synthesis which involves reconstructing speech solely from visual input. However, previous works have struggled to accurately synthesize speech due to a lack of sufficient guidance for the model to infer the correct content with…

Cited by 19PDFcodeScholar
2023

Intuitive Multilingual Audio-Visual Speech Recognition with a Single-Trained Model

EMNLP 2023short findings

We present a novel approach to multilingual audio-visual speech recognition tasks by introducing a single model on a multilingual dataset. Motivated by a human cognitive system where humans can intuitively distinguish different languages without any conscious effort or guidance, we propose a model t…

Cited by 0SourceScholar
2023

Lip Reading for Low-resource Languages by Learning and Combining General Speech Knowledge and Language-specific Knowledge

ICCV 2023poster

This paper proposes a novel lip reading framework, especially for low-resource languages, which has not been well addressed in the previous literature. Since low-resource languages do not have enough video-text paired data to train the model to have sufficient power to model lip movements and langua…

Cited by 18PDFScholar
2023

Mitigating Adversarial Vulnerability through Causal Parameter Estimation by Adversarial Double Machine Learning

ICCV 2023poster

Adversarial examples derived from deliberately crafted perturbations on visual inputs can easily harm decision process of deep neural networks. To prevent potential threats, various adversarial training-based defense methods have grown rapidly and become a de facto standard approach for robustness.…

Cited by 11PDFcodeScholar
2023

Multispectral Invisible Coating: Laminated Visible-Thermal Physical Attack against Multispectral Object Detectors Using Transparent Low-E Films

AAAI 2023technical

Multispectral object detection plays a vital role in safety-critical vision systems that require an around-the-clock operation and encounter dynamic real-world situations(e.g., self-driving cars and autonomous surveillance systems). Despite its crucial competence in safety-related applications, its…

Cited by 9SourcePDFScholar
2023

Similarity Relation Preserving Cross-Modal Learning for Multispectral Pedestrian Detection Against Adversarial Attacks

ICASSP 2023accepted

Although multispectral pedestrian detection studies have shown remarkable detection performances, they are still vulnerable to adversarial attacks. We see the similarity relations between object candidates were not maintained because of the adversarial attacks, resulting in performance degradation.…

Cited by 0SourceScholar
2023

Watch or Listen: Robust Audio-Visual Speech Recognition With Visual Corruption Modeling and Reliability Scoring

CVPR 2023poster

This paper deals with Audio-Visual Speech Recognition (AVSR) under multimodal input corruption situation where audio inputs and visual inputs are both corrupted, which is not well addressed in previous research directions. Previous studies have focused on how to complement the corrupted audio inputs…

2022

Audio-Visual Mismatch-Aware Video Retrieval via Association and Adjustment

ECCV 2022poster

"Retrieving desired videos using natural language queries has attracted increasing attention in research and industry fields as a huge number of videos appear on the internet. Some existing methods attempted to address this video retrieval problem by exploiting multi-modal information, especially au…

Cited by 8SourcePDFScholar
2022

Distinguishing Homophenes Using Multi-Head Visual-Audio Memory for Lip Reading

AAAI 2022technical

Recognizing speech from silent lip movement, which is called lip reading, is a challenging task due to 1) the inherent information insufficiency of lip movement to fully represent the speech, and 2) the existence of homophenes that have similar lip movement with different pronunciations. In this pap…

Cited by 66SourcePDFScholar
2022

Masking Adversarial Damage: Finding Adversarial Saliency for Robust and Sparse Network

CVPR 2022poster

Adversarial examples provoke weak reliability and potential security issues in deep neural networks. Although adversarial training has been widely studied to improve adversarial robustness, it works in an over-parameterized regime and requires high computations and large memory budgets. To bridge ad…

Cited by 18PDFcodeScholar
2022

Robust Thermal Infrared Pedestrian Detection By Associating Visible Pedestrian Knowledge

ICASSP 2022accepted

Recently, pedestrian detection on thermal infrared images has shown the robust pedestrian detection performance. In this paper, we propose a novel thermal infrared pedestrian detection framework which can associate and utilize the complementary pedestrian knowledge from visible images. Motivated by…

Cited by 0SourceScholar
2022

SyncTalkFace: Talking Face Generation with Precise Lip-Syncing via Audio-Lip Memory

AAAI 2022technical

The challenge of talking face generation from speech lies in aligning two different modal information, audio and video, such that the mouth region corresponds to input audio. Previous methods either exploit audio-visual representation learning or leverage intermediate structural information such as…

Cited by 90SourcePDFScholar
2022

Towards Versatile Pedestrian Detector with Multisensory-Matching and Multispectral Recalling Memory

AAAI 2022technical

Recently, automated surveillance cameras can change a visible sensor and a thermal sensor for all-day operation. However, existing single-modal pedestrian detectors mainly focus on detecting pedestrians in only one specific modality (i.e., visible or thermal), so they cannot cope with other modal in…

Cited by 30SourcePDFScholar
2022

VisageSynTalk: Unseen Speaker Video-to-Speech Synthesis via Speech-Visage Feature Selection

ECCV 2022poster

"The goal of this work is to reconstruct speech from a silent talking face video. Recent studies have shown impressive performance on synthesizing speech from silent talking face videos. However, they have not explicitly considered on varying identity characteristics of different speakers, which pla…

Cited by 7SourcePDFScholar
2022

Weakly Paired Associative Learning for Sound and Image Representations via Bimodal Associative Memory

CVPR 2022poster

Data representation learning without labels has attracted increasing attention due to its nature that does not require human annotation. Recently, representation learning has been extended to bimodal data, especially sound and image which are closely related to basic human senses. Existing sound and…

Cited by 7PDFScholar
2021

Distilling Robust and Non-Robust Features in Adversarial Examples by Information Bottleneck

NeurIPS 2021poster

Adversarial examples, generated by carefully crafted perturbation, have attracted considerable attention in research fields. Recent works have argued that the existence of the robust and non-robust features is a primary cause of the adversarial examples, and investigated their internal interactions…

2021

Multi-Modality Associative Bridging Through Memory: Speech Sound Recollected From Face Video

ICCV 2021poster

In this paper, we introduce a novel audio-visual multi-modal bridging framework that can utilize both audio and visual information, even with uni-modal inputs. We exploit a memory network that stores source (i.e., visual) and target (i.e., audio) modal representations, where source modal representat…

Cited by 52PDFScholar
2021

Towards Robust Training of Multi-Sensor Data Fusion Network Against Adversarial Examples in Semantic Segmentation

ICASSP 2021accepted

The success of multi-sensor data fusions in deep learning appears to be attributed to the use of complementary information among multiple sensor datasets. Compared to their predictive performance, relatively less attention has been devoted to the adversarial robustness of multi-sensor data fusion mo…

Cited by 0SourceScholar
2021

Towards a Better Understanding of VR Sickness: Physical Symptom Prediction for VR Contents

AAAI 2021technical

We address the black-box issue of VR sickness assessment (VRSA) by evaluating the level of physical symptoms of VR sickness. For the VR contents inducing the similar VR sickness level, the physical symptoms can vary depending on the characteristics of the contents. Most of existing VRSA methods focu…

Cited by 12SourcePDFScholar
2021

Video Prediction Recalling Long-Term Motion Context via Memory Alignment Learning

CVPR 2021poster

Our work addresses long-term motion context issues for predicting future frames. To predict the future precisely, it is required to capture which long-term motion context (e.g., walking or running) the input motion (e.g., leg movement) belongs to. The bottlenecks arising when dealing with the long-t…

Cited by 146PDFcodeScholar
2021

Visual Comfort Aware-Reinforcement Learning for Depth Adjustment of Stereoscopic 3D Images

AAAI 2021technical

Depth adjustment aims to enhance the visual experience of stereoscopic 3D (S3D) images, which accompanied with improving visual comfort and depth perception. For a human expert, the depth adjustment procedure is a sequence of iterative decision making. The human expert iteratively adjusted the depth…

Cited by 9SourcePDFScholar
2020

SACA Net: Cybersickness Assessment of Individual Viewers for VR Content via Graph-based Symptom Relation Embedding

ECCV 2020poster

Recently, cybersickness assessment for VR content is required to deal with viewing safety issues. Assessing physical symptoms of individual viewers is challenging but important to provide detailed and personalized guides for viewing safety. In this paper, we propose a novel symptom-aware cybersickne…

Cited by 8SourcePDFScholar
2020

Structure Boundary Preserving Segmentation for Medical Image With Ambiguous Boundary

CVPR 2020poster

In this paper, we propose a novel image segmentation method to tackle two critical problems of medical image, which are (i) ambiguity of structure boundary in the medical image domain and (ii) uncertainty of the segmented region without specialized domain knowledge. To solve those two problems in au…

Cited by 169PDFScholar
2020

Towards High-Performance Object Detection: Task-Specific Design Considering Classification and Localization Separation

ICASSP 2020accepted

Object detection performs two tasks (classification and localization) simultaneously. Two tasks share a similarity: they need robust features that effectively represent the visual appearance of the objects. However, two tasks also have different properties. First, classification mainly requires feat…

Cited by 0SourceScholar
2018

Facial Dynamics Interpreter Network: What are the Important Relations between Local Dynamics for Facial Trait Estimation?

ECCV 2018poster

Human face analysis is an important task in computer vision. According to cognitive-psychological studies, facial dynamics could provide crucial cues for face analysis. The motion of a facial local region in facial expression is related to the motion of other facial local regions. In this paper, a n…

Cited by 7SourcePDFScholar
2017

Color channel-wise recurrent learning for facial expression recognition

ICASSP 2017accepted

Facial expression recognition is increasingly gaining importance in emerging affective computing applications. In practice, achieving accurate facial expression recognition is still challenging due to environmental variations. In this paper, we propose a color channel-wise recurrent facial feature l…

Cited by 0SourceScholar
2016

Latent feature representation with 3-D multi-view deep convolutional neural network for bilateral analysis in digital breast tomosynthesis

ICASSP 2016accepted

In clinical studies of breast cancer, masses appear as asymmetric densities between the left and the right breasts, which show different breast tissue structures. For classifying breast masses, most researchers have developed hand-crafted bilateral features by extracting the asymmetric information i…

Cited by 0SourceScholar