← Search

Jihwan Lee

6 accepted papers

2026

AMPED: Adaptive Multi-objective Projection for balancing Exploration and skill Diversification

ICLR 2026poster

Skill-based reinforcement learning (SBRL) enables rapid adaptation in environments with sparse rewards by pretraining a skill-conditioned policy. Effective skill learning requires jointly maximizing both exploration and skill diversity. However, existing methods often face challenges in simultaneous…

Cited by 0SourcecodeScholar
2026

Follow the Saliency: Supervised Saliency for Retrieval-augmented Dense Video Captioning

CVPR 2026

Existing retrieval-augmented approaches for Dense Video Captioning (DVC) often fail to achieve accurate temporal segmentation aligned with true event boundaries, as they rely on heuristic strategies that overlook ground truth event boundaries.The proposed framework, STaRC, overcomes this limitation

Cited by 0SourcecodeScholar
2026

TRACED: Transition-aware Regret Approximation with Co-learnability for Environment Design

ICLR 2026poster

Generalizing deep reinforcement learning agents to unseen environments remains a significant challenge. One promising solution is Unsupervised Environment Design (UED), a co‑evolutionary framework in which a teacher adaptively generates tasks with high learning potential, while a student learns a ro…

Cited by 0SourcecodeScholar
2025

Enhancing Listened Speech Decoding from EEG via Parallel Phoneme Sequence Prediction

ICASSP 2025accepted

Brain-computer interfaces (BCI) offer numerous human-centered application possibilities, particularly affecting people with neurological disorders. Text or speech decoding from brain activities is a relevant domain that could augment the quality of life for people with impaired speech perception. We…

Cited by 0SourceScholar
2024

Mels-Tts : Multi-Emotion Multi-Lingual Multi-Speaker Text-To-Speech System Via Disentangled Style Tokens

ICASSP 2024accepted

This paper proposes a multi-emotion, multi-lingual, and multi-speaker text-to-speech (MELS-TTS) system, employing disentangled style tokens for effective emotion transfer. In speech encompassing various attributes, such as emotional state, speaker identity, and linguistic style, disentangling these…

Cited by 0SourceScholar