← Search

Junhyeok Lee

7 accepted papers

2026

MASKVCT: MASKED VOICE CODEC TRANSFORMER FOR ZERO-SHOT VOICE CONVERSION WITH INCREASED CONTROLLABILITY VIA MULTIPLE GUIDANCES

ICASSP 2026poster

We introduce MaskVCT, a zero-shot voice conversion (VC) model that offers multi-factor controllability through multiple classifier-free guidances (CFGs). While previous VC models rely on a fixed conditioning scheme, MaskVCT integrates diverse conditions in a single model. To further enhance robustne…

Cited by 0SourcePDFScholar
2025

REVECA: Adaptive Planning and Trajectory-Based Validation in Cooperative Language Agents Using Information Relevance and Relative Proximity

AAAI 2025technical

We address the challenge of multi-agent cooperation, where agents achieve a common goal by cooperating with decentralized agents under complex partial observations. Existing cooperative agent systems often struggle with efficiently processing continuously accumulating information, managing globally…

Cited by 0SourcePDFScholar
2023

Direct Preference-based Policy Optimization without Reward Modeling

NeurIPS 2023poster

Preference-based reinforcement learning (PbRL) is an approach that enables RL agents to learn from preference, which is particularly useful when formulating a reward function is challenging. Existing PbRL methods generally involve a two-step procedure: they first learn a reward model based on given…

2023

PhaseAug: A Differentiable Augmentation for Speech Synthesis to Simulate One-to-Many Mapping

ICASSP 2023accepted

Previous generative adversarial network (GAN)-based neural vocoders are trained to reconstruct the exact ground truth wave-form from the paired mel-spectrogram and do not consider the one-to-many relationship of speech synthesis. This conventional training causes overfitting for both the discriminat…

Cited by 0SourceScholar
2022

ASSEM-VC: Realistic Voice Conversion by Assembling Modern Speech Synthesis Techniques

ICASSP 2022accepted

Recent works on voice conversion (VC) focus on preserving the rhythm and the intonation as well as the linguistic content. To preserve these features from the source, we decompose current non-parallel VC systems into two encoders and one decoder. We analyze each module with several experiments and r…

Cited by 0SourceScholar
2022

Query-Efficient and Scalable Black-Box Adversarial Attacks on Discrete Sequential Data via Bayesian Optimization

ICML 2022spotlight

We focus on the problem of adversarial attacks against models on discrete sequential data in the black-box setting where the attacker aims to craft adversarial examples with limited query access to the victim model. Existing black-box attacks, mostly based on greedy algorithms, find adversarial exam…