← Search

Taehwan Kim

11 accepted papers

2026

Cross-Modal Emotion Transfer for Emotion Editing in Talking Face Video

CVPR 2026

Talking face generation has gained significant attention as a core application of generative models.To enhance the expressiveness and realism of synthesized videos, emotion editing in talking face video plays a crucial role.However, existing approaches often limit expressive flexibility and struggle

Cited by 0SourcecodeScholar
2026

MM-SeR: Multimodal Self-Refinement for Lightweight Image Captioning

CVPR 2026

Systems such as video chatbots and navigation robots often depend on streaming image captioning to interpret visual inputs. Existing approaches typically employ large multimodal language models (MLLMs) for this purpose, but their substantial computational cost hinders practical application.This limi

Cited by 0SourcecodeScholar
2025

RingFormer: Rethinking Recurrent Transformer with Adaptive Level Signals

EMNLP 2025

Transformers have achieved great success in effectively processing sequential data such as text. Their architecture consisting of several attention and feedforward blocks can model relations between elements of a sequence in parallel manner, which makes them very efficient to train and effective in

Cited by 0SourcePDFScholar
2025

VEHME: A Vision-Language Model For Evaluating Handwritten Mathematics Expressions

EMNLP 2025

Automatically assessing handwritten mathematical solutions is an important problem in educational technology with practical applications, but remains a significant challenge due to the diverse formats, unstructured layouts, and symbolic complexity of student work. To address this challenge, we intro

Cited by 0SourcePDFScholar
2023

Sound of Story: Multi-modal Storytelling with Audio

EMNLP 2023long findings

Storytelling is multi-modal in the real world. When one tells a story, one may use all of the visualizations and sounds along with the story itself. However, prior studies on storytelling datasets and tasks have paid little attention to sound even though sound also conveys meaningful semantics of th…

Cited by 0SourcecodeScholar
2018

Optimal Spectral Estimation and System Trade-Off in Long-Distance Frequency-Modulated Continuous-Wave Lidar

ICASSP 2018accepted

Frequency-modulated continuous-wave (FMCW) LIDAR is a promising technology for next-generation integrated 3D imaging systems. However, it has been considered difficult to apply FMCW LIDAR for long-distance (> 100m) targets, such as those in automotive and airborne applications. Maintaining coherence…

Cited by 0SourceScholar
2016

Signer-independent fingerspelling recognition with deep neural network adaptation

ICASSP 2016accepted

We study the problem of recognition of fingerspelled letter sequences in American Sign Language in a signer-independent setting. Fingerspelled sequences are both challenging and important to recognize, as they are used for many content words such as proper nouns and technical terms. Previous work ha…

Cited by 0SourceScholar