← Search

Taesoo Kim

7 accepted papers

2026

Cross-Modal Emotion Transfer for Emotion Editing in Talking Face Video

CVPR 2026

Talking face generation has gained significant attention as a core application of generative models.To enhance the expressiveness and realism of synthesized videos, emotion editing in talking face video plays a crucial role.However, existing approaches often limit expressive flexibility and struggle

Cited by 0SourcecodeScholar
2025

Do Not Mimic My Voice : Speaker Identity Unlearning for Zero-Shot Text-to-Speech

ICML 2025poster

The rapid advancement of Zero-Shot Text-to-Speech (ZS-TTS) technology has enabled high-fidelity voice synthesis from minimal audio cues, raising significant privacy and ethical concerns. Despite the threats to voice privacy, research to selectively remove the knowledge to replicate unwanted individu…

2025

Heterogeneous Graph Neural Network on Semantic Tree

AAAI 2025technical

The recent past has seen an increasing interest in Heterogeneous Graph Neural Networks (HGNNs), since many real-world graphs are heterogeneous in nature, from citation graphs to email graphs. However, existing methods ignore a tree hierarchy among metapaths, naturally constituted by different node t…

2024

Selective Generation for Controllable Language Models

NeurIPS 2024spotlight

Trustworthiness of generative language models (GLMs) is crucial in their deployment to critical decision making systems. Hence, certified risk control methods such as selective prediction and conformal prediction have been applied to mitigating the hallucination problem in various supervised downstr…

2023

Regression to Classification: Waveform Encoding for Neural Field-Based Audio Signal Representation

ICASSP 2023accepted

Neural fields, also known as coordinate-based representations, are an emerging signal representation framework. This approach has also been used to represent audio signals, but the generated audio often contains noise. To reduce noise and improve representation quality, we propose using waveform enc…

Cited by 0SourceScholar
2022

ADA-VAD: Unpaired Adversarial Domain Adaptation for Noise-Robust Voice Activity Detection

ICASSP 2022accepted

Voice Activity Detection (VAD) is becoming an essential front-end component in various speech processing systems. As those systems are commonly deployed in environments with diverse noise types and low signal-to-noise ratios (SNRs), an effective VAD method should perform robust detection of speech r…

Cited by 0SourceScholar