← Search

Sungwoo Cho

4 accepted papers

2025

Flex-Judge: Text-Only Reasoning Unleashes Zero-Shot Multimodal Evaluators

NeurIPS 2025poster

Human-generated reward signals are critical for aligning generative models with human preferences, guiding both training and inference-time evaluations. While large language models (LLMs) employed as proxy evaluators, i.e., LLM-as-a-Judge, significantly reduce the costs associated with manual annota…

Cited by 0SourceScholar
2025

MAVFlow: Preserving Paralinguistic Elements with Conditional Flow Matching for Zero-Shot AV2AV Multilingual Translation

ICCV 2025poster

Despite recent advances in text-to-speech (TTS) models, audio-visual-to-audio-visual (AV2AV) translation still faces a critical challenge: maintaining speaker consistency between the original and translated vocal and facial features. To address this issue, we propose a conditional flow matching (CFM…

2025

MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition

ICML 2025poster

Audio-visual speech recognition (AVSR) has become critical for enhancing speech recognition in noisy environments by integrating both auditory and visual modalities. However, existing AVSR systems struggle to scale up without compromising computational efficiency. In this study, we introduce MoHAVE…

Cited by 2SourcePDFScholar
2025

Multi-Task Corrupted Prediction for Learning Robust Audio-Visual Speech Representation

ICLR 2025poster

Audio-visual speech recognition (AVSR) incorporates auditory and visual modalities to improve recognition accuracy, particularly in noisy environments where audio-only speech systems are insufficient. While previous research has largely addressed audio disruptions, few studies have dealt with visual…