← Search

Taehoon Kim

10 accepted papers

2026

CATALYST: Cognitive-To-Autonomy-Inspired Two-Stage Training Data Generation with Local-System-Aware Selection Technique

ICRA 2026poster

In conventional learning-based robotic dynamics modeling, physical information is mostly incorporated into the model or loss function, while the design of training data often relies on random sampling or uniform coverage, which can limit performance. To address this gap, this paper proposes the CATA…

Cited by 0Scholar
2026

Test-Time Alignment of Text-to-Image Diffusion Models via Null-Text Embedding Optimisation

CVPR 2026

Test-time alignment (TTA) aims to adapt models to specific rewards during inference. However, existing methods tend to either under-optimise or over-optimise (reward hack) the target reward function. We propose Null-Text Test-Time Alignment (Null-TTA), which aligns diffusion models by optimising the

Cited by 0SourceScholar
2025

Beyond Spatial Frequency: Pixel-wise Temporal Frequency-based Deepfake Video Detection

ICCV 2025poster

We introduce a deepfake video detection approach that exploits pixel-wise temporal inconsistencies, which traditional spatial frequency-based detectors often overlook. The traditional detectors represent temporal information merely by stacking spatial frequency spectra across frames, resulting in th…

2024

D3T: Distinctive Dual-Domain Teacher Zigzagging Across RGB-Thermal Gap for Domain-Adaptive Object Detection

CVPR 2024poster

Domain adaptation for object detection typically entails transferring knowledge from one visible domain to another visible domain. However there are limited studies on adapting from the visible to the thermal domain because the domain gap between the visible and thermal domains is much larger than e…

2024

Exploiting Style Latent Flows for Generalizing Deepfake Video Detection

CVPR 2024poster

This paper presents a new approach for the detection of fake videos based on the analysis of style latent vectors and their abnormal behavior in temporal changes in the generated videos. We discovered that the generated facial videos suffer from the temporal distinctiveness in the temporal changes o…

Cited by 34SourcePDFScholar
2022

L-Verse: Bidirectional Generation Between Image and Text

CVPR 2022oral

Far beyond learning long-range interactions of natural language, transformers are becoming the de-facto standard for many vision tasks with their power and scalability. Especially with cross-modal tasks between image and text, vector quantized variational autoencoders (VQ-VAEs) are widely used to ma…

Cited by 33PDFcodeScholar
2021

Prosodic Clustering for Phoneme-Level Prosody Control in End-to-End Speech Synthesis

ICASSP 2021accepted

This paper presents a method for controlling the prosody at the phoneme level in an autoregressive attention-based text-to-speech system. Instead of learning latent prosodic features with a variational framework as is commonly done, we directly extract phoneme-level F0 and duration features from the…

Cited by 0SourceScholar
2019

Quantifying Generalization in Reinforcement Learning

ICML 2019oral

In this paper, we investigate the problem of overfitting in deep reinforcement learning. Among the most common benchmarks in RL, it is customary to use the same environments for both training and testing. This practice offers relatively little insight into an agent’s ability to generalize. We addres…

2019

Transfer Learning via Unsupervised Task Discovery for Visual Question Answering

CVPR 2019poster

We study how to leverage off-the-shelf visual and linguistic data to cope with out-of-vocabulary answers in visual question answering task. Existing large-scale visual datasets with annotations such as image class labels, bounding boxes and region descriptions are good sources for learning rich and…

Cited by 22PDFcodeScholar