← Search

Chanho Park

8 accepted papers

2026

Token Warping Helps MLLMs Look from Nearby Viewpoints

CVPR 2026

Can warping tokens, rather than pixels, help multimodal large language models (MLLMs) understand how a scene appears from a nearby viewpoint? While MLLMs perform well on visual reasoning, they remain fragile to viewpoint changes, as pixel-wise warping is highly sensitive to small depth errors and of

Cited by 0SourcecodeScholar
2025

Fast Word Error Rate Estimation Using Self-Supervised Representations for Speech and Text

ICASSP 2025accepted

Word error rate (WER) estimation aims to evaluate the quality of an automatic speech recognition (ASR) system’s output without requiring ground-truth labels. This task has gained increasing attention as advanced ASR systems are trained on large amounts of data. In this context, the computational eff…

Cited by 0SourceScholar
2025

Perspective-Aware Reasoning in Vision-Language Models via Mental Imagery Simulation

ICCV 2025poster

We present a framework for perspective-aware reasoning in vision-language models (VLMs) through mental imagery simulation. Perspective-taking, the ability to perceive an environment or situation from an alternative viewpoint, is a key benchmark for human-level visual understanding, essential for env…

Cited by 0SourcePDFScholar
2025

SHARE: Shared Memory-Aware Open-Domain Long-Term Dialogue Dataset Constructed from Movie Script

ACL 2025long

Shared memories between two individuals strengthen their bond and are crucial for facilitating their ongoing conversations. This study aims to make long-term dialogue more engaging by leveraging these shared memories. To this end, we introduce a new long-term dialogue dataset named SHARE, constructe…

2024

Automatic Speech Recognition System-Independent Word Error Rate Estimation

COLING 2024main

Word error rate (WER) is a metric used to evaluate the quality of transcriptions produced by Automatic Speech Recognition (ASR) systems. In many applications, it is of interest to estimate WER given a pair of a speech utterance and a transcript. Previous work on WER estimation focused on building mo…

2024

SignSGD with Federated Defense: Harnessing Adversarial Attacks through Gradient Sign Decoding

ICML 2024poster

Distributed learning is an effective approach to accelerate model training by using parallel computing power of multiple workers. However, substantial communication delays arise between workers and a parameter server due to the massive costs associated with communicating gradients. SignSGD with majo…

Cited by 2SourcePDFScholar
2022

Unsupervised Data Selection for Speech Recognition with Contrastive Loss Ratios

ICASSP 2022accepted

This paper proposes an unsupervised data selection method by using a submodular function based on contrastive loss ratios of target and training data sets. A model using a contrastive loss function is trained on both sets. Then the ratio of frame-level losses for each model is used by a submodular f…

Cited by 0SourceScholar