← Search

Seong-Whan Lee

40 accepted papers

2026

TF-CADE: Foreground-Concentrated Text-Video Alignment for Zero-Shot Temporal Action Detection

CVPR 2026

Zero-Shot Temporal Action Detection (ZSTAD) aims to localize and recognize action instances from unseen action categories in untrimmed videos. Although existing methods have shown effectiveness by advancing architectural text-video alignment, they still struggle with capturing semantic distinctions

Cited by 0SourceScholar
2026

Toward Complex-Valued Neural Networks for Waveform Generation

ICLR 2026poster

Neural vocoders have recently advanced waveform generation, yielding natural and expressive audio. Among these approaches, iSTFT-based vocoders have recently gained attention. They predict a complex-valued spectrogram and then synthesize the waveform via iSTFT, thereby avoiding learned upsampling st…

Cited by 0SourcecodeScholar
2025

Comprehensive Information Bottleneck for Unveiling Universal Attribution to Interpret Vision Transformers

CVPR 2025highlight

The feature attribution method reveals the contribution of input variables to the decision-making process to provide an attribution map for explanation. Existing methods grounded on the information bottleneck principle compute information in a specific layer to obtain attributions, compressing the f…

Cited by 0SourcePDFScholar
2025

DiGIT: Multi-Dilated Gated Encoder and Central-Adjacent Region Integrated Decoder for Temporal Action Detection Transformer

CVPR 2025poster

In this paper, we examine a key limitation in query-based detectors for temporal action detection (TAD), which arises from their direct adaptation of originally designed architectures for object detection. Despite the effectiveness of the existing models, they struggle to fully address the unique ch…

2025

FLowHigh: Towards Efficient and High-Quality Audio Super-Resolution with Single-Step Flow Matching

ICASSP 2025accepted

Audio super-resolution is challenging owing to its ill-posed nature. Recently, the application of diffusion models in audio super-resolution has shown promising results in alleviating this challenge. However, diffusion-based models have limitations, primarily the necessity for numerous sampling step…

Cited by 0SourceScholar
2025

FillerSpeech: Towards Human-Like Text-to-Speech Synthesis with Filler Insertion and Filler Style Control

EMNLP 2025

Recent advancements in speech synthesis have significantly improved the audio quality and pronunciation of synthesized speech. To further advance toward human-like conversational speech synthesis, this paper presents FillerSpeech, a novel speech synthesis framework that enables natural filler insert

2025

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis

ICASSP 2025accepted

Recently, there has been a growing demand for conversational speech synthesis (CSS) that generates more natural speech by considering the conversational context. To address this, we introduce JELLY, a novel CSS framework that integrates emotion recognition and context reasoning for generating approp…

Cited by 0SourceScholar
2025

MIRe: Enhancing Multimodal Queries Representation via Fusion-Free Modality Interaction for Multimodal Retrieval

ACL 2025finding

Recent multimodal retrieval methods have endowed text-based retrievers with multimodal capabilities by utilizing pre-training strategies for visual-text alignment. They often directly fuse the two modalities for cross-reference during the alignment to understand multimodal queries. However, existing…

2025

Multi-Context Temporal Consistent Modeling for Referring Video Object Segmentation

ICASSP 2025accepted

Referring video object segmentation aims to segment objects within a video corresponding to a given text description. Existing transformer-based temporal modeling approaches face challenges related to query inconsistency and the limited consideration of context. Query in-consistency produces unstabl…

Cited by 0SourceScholar
2025

PeriodWave: Multi-Period Flow Matching for High-Fidelity Waveform Generation

ICLR 2025poster

Recently, universal waveform generation tasks have been investigated conditioned on various out-of-distribution scenarios. Although one-step GAN-based methods have shown their strength in fast waveform generation, they are vulnerable to train-inference mismatch scenarios such as two-stage text-to-sp…

2025

PoseAnchor: Robust Root Position Estimation for 3D Human Pose Estimation

ICCV 2025poster

Standard 3D human pose estimation (HPE) benchmarks employ root-centering, which normalizes poses relative to the pelvis but discards absolute root position information. While effective for evaluation, this approach limits real-world applications such as motion tracking, AR/VR, and human-computer int…

2025

ProPose: Probabilistic 3D Human Pose Estimation with Instance-Level Distribution and Normalizing Flow

AAAI 2025technical

3D Human Pose Estimation (HPE) is a one-to-many problem by nature, making it challenging to estimate an accurate 3D pose from a single 2D pose. Some prior works have attempted to tackle this problem by using a conditional generative network. They generate 3D poses from a given 2D pose with noises fr…

2025

Towards Dynamic Neural Communication and Speech Neuroprosthesis Based on Viseme Decoding

ICASSP 2025accepted

Decoding text, speech, or images from human neural signals holds promising potential both as neuroprosthesis for patients and as innovative communication tools for general users. Although neural signals contain various information on speech intentions, movements, and phonetic details, generating inf…

Cited by 0SourceScholar
2025

Towards Fine-Grained Interpretability: Counterfactual Explanations for Misclassification with Saliency Partition

CVPR 2025poster

Attribution-based explanation techniques capture key patterns to enhance visual interpretability. However, these patterns often lack the granularity needed for insight in fine-grained tasks, particularly in cases of model misclassification, where explanations may be insufficiently detailed. To addre…

2025

Towards Generalizable 3D Human Pose Estimation via Ensembles on Flat Loss Landscapes

NeurIPS 2025poster

3D Human Pose Estimation (HPE) is a fundamental task in the computer vision. Generalization in 3D HPE task is crucial due to the need for robustness across diverse environments and datasets. Existing methods often focus on learning relationships between joints to enhance the generalization capabilit…

Cited by 0SourceScholar
2025

XLQA: A Benchmark for Locale-Aware Multilingual Open-Domain Question Answering

EMNLP 2025

Large Language Models (LLMs) have shown significant progress in Open-domain question answering (ODQA), yet most evaluations focus on English and assume locale-invariant answers across languages. This assumption neglects the cultural and regional variations that affect question understanding and answ

2024

DDDM-VC: Decoupled Denoising Diffusion Models with Disentangled Representation and Prior Mixup for Verified Robust Voice Conversion

AAAI 2024technical

Diffusion-based generative models have recently exhibited powerful generative performance. However, as many attributes exist in the data distribution and owing to several limitations of sharing the model parameters across all levels of the generation process, it remains challenging to control specif…

2024

Midi-Voice: Expressive Zero-Shot Singing Voice Synthesis via Midi-Driven Priors

ICASSP 2024accepted

Recently, singing voice synthesis (SVS) models have shown significant progress with generative models. However, previous SVS models inaccurately predict prior and fundamental frequency (F0) for unseen speakers, resulting in a low-quality generated singing voice. To address these issues, in this pape…

Cited by 0SourceScholar
2024

TE-TAD: Towards Full End-to-End Temporal Action Detection via Time-Aligned Coordinate Expression

CVPR 2024poster

In this paper we investigate that the normalized coordinate expression is a key factor as reliance on hand-crafted components in query-based detectors for temporal action detection (TAD). Despite significant advancements towards an end-to-end framework in object detection query-based detectors have…

2024

Text-Infused Attention and Foreground-Aware Modeling for Zero-Shot Temporal Action Detection

NeurIPS 2024poster

Zero-Shot Temporal Action Detection (ZSTAD) aims to classify and localize action segments in untrimmed videos for unseen action categories. Most existing ZSTAD methods utilize a foreground-based approach, limiting the integration of text and visual features due to their reliance on pre-extracted pro…

2024

TranSentence: speech-to-speech Translation via Language-Agnostic Sentence-Level Speech Encoding without Language-Parallel Data

ICASSP 2024accepted

Although there has been significant advancement in the field of speech-to-speech translation, conventional models still require language-parallel speech data between the source and target languages for training. In this paper, we introduce TranSentence, a novel speech-to-speech translation without l…

Cited by 0SourceScholar
2024

Unknown-Aware Graph Regularization for Robust Semi-supervised Learning from Uncurated Data

AAAI 2024technical

Recent advances in semi-supervised learning (SSL) have relied on the optimistic assumption that labeled and unlabeled data share the same class distribution. However, this assumption is often violated in real-world scenarios, where unlabeled data may contain out-of-class samples. SSL with such uncur…

2023

Compensatory Debiasing For Gender Imbalances In Language Models

ICASSP 2023accepted

Pre-trained language models (PLMs) learn gender bias from imbalances in human-written corpora. This bias leads to critical social issues when deploying PLMs in real-world scenarios. However, minimizing bias is limited by the trade-off due to the degradation of language modeling performance. It is pa…

Cited by 0SourceScholar
2023

Towards Better Visualizing the Decision Basis of Networks via Unfold and Conquer Attribution Guidance

AAAI 2023technical

Revealing the transparency of Deep Neural Networks (DNNs) has been widely studied to describe the decision mechanisms of network inner structures. In this paper, we propose a novel post-hoc framework, Unfold and Conquer Attribution Guidance (UCAG), which enhances the explainability of the network de…

2023

Towards Voice Reconstruction from EEG during Imagined Speech

AAAI 2023technical

Translating imagined speech from human brain activity into voice is a challenging and absorbing research issue that can provide new means of human communication via brain signals. Efforts to reconstruct speech from brain activity have shown their potential using invasive measures of spoken speech da…

2022

EMOQ-TTS: Emotion Intensity Quantization for Fine-Grained Controllable Emotional Text-to-Speech

ICASSP 2022accepted

Although recent advances in text-to-speech (TTS) have shown significant improvement, it is still limited to emotional speech synthesis. To produce emotional speech, most works utilize emotion information extracted from emotion labels or reference audio. However, they result in monotonous emotional e…

Cited by 0SourceScholar
2022

Emergence of Hierarchical Layers in a Single Sheet of Self-Organizing Spiking Neurons

NeurIPS 2022accept

Traditionally convolutional neural network architectures have been designed by stacking layers on top of each other to form deeper hierarchical networks. The cortex in the brain however does not just stack layers as done in standard convolution neural networks, instead different regions are organize…

Cited by 2SourcePDFScholar
2022

FRE-GAN 2: Fast and Efficient Frequency-Consistent Audio Synthesis

ICASSP 2022accepted

Although recent advances in neural vocoder have shown significant improvement, most of these models have a trade-off between audio quality and computational complexity. Since the large model has a limitation on the low-resource devices, a more efficient neural vocoder should synthesize high-quality…

Cited by 0SourceScholar
2022

HierSpeech: Bridging the Gap between Text and Speech by Hierarchical Variational Inference using Self-supervised Representations for Speech Synthesis

NeurIPS 2022accept

This paper presents HierSpeech, a high-quality end-to-end text-to-speech (TTS) system based on a hierarchical conditional variational autoencoder (VAE) utilizing self-supervised speech representations. Recently, single-stage TTS systems, which directly generate raw speech waveform from text, have be…

Cited by 57SourcePDFScholar
2022

PVAE-TTS: Adaptive Text-to-Speech via Progressive Style Adaptation

ICASSP 2022accepted

Adaptive text-to-speech (TTS) has attracted increasing interests for the purpose of training TTS systems without tons of high quality data. Nevertheless, existing adaptive TTS systems still show low adaptation quality for novel speakers, since it is hard to learn an extensive speaking style with lim…

Cited by 0SourceScholar
2021

Interpreting Deep Neural Networks with Relative Sectional Propagation by Analyzing Comparative Gradients and Hostile Activations

AAAI 2021technical

The clear transparency of Deep Neural Networks (DNNs) is hampered by complex internal structures and nonlinear transformations along deep hierarchies. In this paper, we propose a new attribution method, Relative Sectional Propagation (RSP), for fully decomposing the output predictions with the chara…

Cited by 18SourcePDFScholar
2021

Multi-SpectroGAN: High-Diversity and High-Fidelity Spectrogram Generation with Adversarial Style Combination for Speech Synthesis

AAAI 2021technical

While generative adversarial networks (GANs) based neural text-to-speech (TTS) systems have shown significant improvement in neural speech synthesis, there is no TTS system to learn to synthesize speech from text sequences with only adversarial feedback. Because adversarial feedback alone is not suf…

Cited by 65SourcePDFScholar
2020

Classification of High-Dimensional Motor Imagery Tasks Based on An End-To-End Role Assigned Convolutional Neural Network

ICASSP 2020accepted

A brain-computer interface (BCI) provides a direct communication pathway between user and external devices. EEG-based motor imagery paradigm is widely used in non-invasive BCI to obtain encoded signals contained user intention of movement execution. However, EEG has intricate and non-stationary prop…

Cited by 0SourceScholar
2020

Decoding Movement Imagination and Execution From Eeg Signals Using Bci-Transfer Learning Method Based on Relation Network

ICASSP 2020accepted

A brain-computer interface (BCI) is used to control external devices for healthy people as well as to rehabilitate motor functions for motor-disabled patients. Decoding movement intention is one of the most significant aspects for performing arm movement tasks using brain signals. Decoding movement…

Cited by 0SourceScholar
2018

Deep Reinforcement Learning in Continuous Action Spaces: a Case Study in the Game of Simulated Curling

ICML 2018oral

Many real-world applications of reinforcement learning require an agent to select optimal actions from continuous spaces. Recently, deep neural networks have successfully been applied to games with discrete actions spaces. However, deep neural networks for discrete actions are not suitable for devis…