← Search

Tae-Ho Kim

2 accepted papers

2026

ERGO: Efficient High-Resolution Visual Understanding for Vision-Language Models

ICLR 2026poster

Efficient processing of high-resolution images is crucial for real-world vision–language applications. However, existing Large Vision-Language Models (LVLMs) incur substantial computational overhead due to the large number of vision tokens. With the advent of "thinking with images" models, reasoning…

Cited by 0SourcecodeScholar
2020

Emotional Voice Conversion Using Multitask Learning with Text-To-Speech

ICASSP 2020accepted

Voice conversion (VC) is a task that alters the voice of a person to suit different styles while conserving the linguistic content. Previous state-of-the-art technology used in VC was based on the sequence-to-sequence (seq2seq) model, which could lose linguistic information. There was an attempt to…

Cited by 0SourceScholar