← Search

Jihyun Lee

26 accepted papers

2026

Large-scale Codec Avatars: The Unreasonable Effectiveness of Large-scale Avatar Pretraining

CVPR 2026

High-quality 3D avatar modeling faces a critical trade-off between fidelity and generalization. On the one hand, multi-view studio data enables high-fidelity modeling of humans with precise control over expressions and poses, but it struggles to generalize to real-world data due to limited scale and

Cited by 0SourcecodeScholar
2026

PhysHanDI: Physics-Based Reconstruction of Hand-Deformable Object Interactions

ICML 2026poster

While existing methods for reconstructing hand–object interactions have made impressive progress, they either focus on rigid or part-wise rigid objects—limiting their ability to model real-world objects (e.g., cloth, stuffed animals) that exhibit highly non-rigid deformations—or model deformable obj…

Cited by 0SourceScholar
2025

Generative Modeling of Shape-Dependent Self-Contact Human Poses

ICCV 2025poster

One can hardly model self-contact of human poses without considering underlying body shapes. For example, the pose of rubbing a belly for a person with a low BMI leads to penetration of the hand into the belly for a person with a high BMI. Despite its relevance, existing self-contact datasets lack t…

Cited by 0SourcePDFScholar
2025

MIRROR: Multimodal Cognitive Reframing Therapy for Rolling with Resistance

EMNLP 2025

Recent studies have explored the use of large language models (LLMs) in psychotherapy; however, text-based cognitive behavioral therapy (CBT) models often struggle with client resistance, which can weaken therapeutic alliance. To address this, we propose a multimodal approach that incorporates nonve

2025

MPMAvatar: Learning 3D Gaussian Avatars with Accurate and Robust Physics-Based Dynamics

NeurIPS 2025poster

While there has been significant progress in the field of 3D avatar creation from visual observations, modeling physically plausible dynamics of humans with loose garments remains a challenging problem. Although a few existing works address this problem by leveraging physical simulation, they suffer…

Cited by 0SourceScholar
2025

ORIGEN: Zero-Shot 3D Orientation Grounding in Text-to-Image Generation

NeurIPS 2025poster

We introduce ORIGEN, the first zero-shot method for 3D orientation grounding in text-to-image generation across multiple objects and diverse categories. While previous work on spatial grounding in image generation has mainly focused on 2D positioning, it lacks control over 3D orientation. To address…

Cited by 0SourceScholar
2025

PanicToCalm: A Proactive Counseling Agent for Panic Attacks

EMNLP 2025

Panic attacks are acute episodes of fear and distress, in which timely, appropriate intervention can significantly help individuals regain stability. However, suitable datasets for training such models remain scarce due to ethical and logistical issues. To address this, we introduce Pace, which is a

Cited by 0SourcePDFScholar
2025

PicPersona-TOD : A Dataset for Personalizing Utterance Style in Task-Oriented Dialogue with Image Persona

NAACL 2025long

Task-Oriented Dialogue (TOD) systems are designed to fulfill user requests through natural language interactions, yet existing systems often produce generic, monotonic responses that lack individuality and fail to adapt to users’ personal attributes. To address this, we introduce PicPersona-TOD, a n…

2025

Prediction-Feedback DETR for Temporal Action Detection

AAAI 2025technical

Temporal Action Detection (TAD) is fundamental yet challenging for real-world video applications. Leveraging the unique benefits of transformers, various DETR-based approaches have been adopted in TAD. However, it has recently been identified that the attention collapse in self-attention causes the…

Cited by 1SourcePDFScholar
2025

Progressive Facial Granularity Aggregation with Bilateral Attribute-based Enhancement for Face-to-Speech Synthesis

EMNLP 2025

For individuals who have experienced traumatic events such as strokes, speech may no longer be a viable means of communication. While text-to-speech (TTS) can be used as a communication aid since it generates synthetic speech, it fails to preserve the user’s own voice. As such, face-to-voice (FTV) s

2025

Prompt-Guided Selective Masking Loss for Context-Aware Emotive Text-to-Speech

NAACL 2025findings

Emotional dialogue speech synthesis (EDSS) aims to generate expressive speech by leveraging the dialogue context between interlocutors. This is typically done by concatenating global representations of previous utterances as conditions for text-to-speech (TTS) systems. However, such approaches overl…

Cited by 4SourcePDFScholar
2025

REWIND: Real-Time Egocentric Whole-Body Motion Diffusion with Exemplar-Based Identity Conditioning

CVPR 2025poster

We present REWIND (Real-Time Egocentric Whole-Body Motion Diffusion), a one-step diffusion model for real-time, high-fidelity human motion estimation from egocentric image inputs. While an existing method for egocentric whole-body (i.e., body and hands) motion estimation is non-real-time and acausal…

Cited by 0SourcePDFScholar
2024

Dense Hand-Object(HO) GraspNet with Full Grasping Taxonomy and Dynamics

ECCV 2024poster

"Existing datasets for 3D hand-object interaction are limited either in the data cardinality, data variations in interaction scenarios, or the quality of annotations. In this work, we present a comprehensive new training dataset for hand-object interaction called HOGraspNet. It is the only real data…

2024

InterHandGen: Two-Hand Interaction Generation via Cascaded Reverse Diffusion

CVPR 2024poster

We present InterHandGen a novel framework that learns the generative prior of two-hand interaction. Sampling from our model yields plausible and diverse two-hand shapes in close interaction with or without an object. Our prior can be incorporated into any optimization or learning methods to reduce a…

Cited by 13SourcePDFScholar
2023

End-to-End Neural Audio Coding in the MDCT Domain

ICASSP 2023accepted

Modern deep neural network (DNN)-based audio coding approaches utilize complicated non-linear functions (e.g., convolutional neural networks and non-linear activations), which leads to high complexity and memory usage. However, their decoded audio quality is still not much higher than that of signal…

Cited by 0SourceScholar
2023

FourierHandFlow: Neural 4D Hand Representation Using Fourier Query Flow

NeurIPS 2023poster

Recent 4D shape representations model continuous temporal evolution of implicit shapes by (1) learning query flows without leveraging shape and articulation priors or (2) decoding shape occupancies separately for each time value. Thus, they do not effectively capture implicit correspondences between…

Cited by 5SourcePDFScholar
2023

Im2Hands: Learning Attentive Implicit Representation of Interacting Two-Hand Shapes

CVPR 2023poster

We present Implicit Two Hands (Im2Hands), the first neural implicit representation of two interacting hands. Unlike existing methods on two-hand reconstruction that rely on a parametric hand model and/or low-resolution meshes, Im2Hands can produce fine-grained geometry of two hands with high hand-to…

2023

Progressive Multi-Stage Neural Audio Codec with Psychoacoustic Loss and Discriminator

ICASSP 2023accepted

In this paper, we improve the efficiency of the progressive multi-stage neural audio codec (PR-Codec) by utilizing perceptually motivated training criteria. Although our baseline PR-Codec successfully reconstructs full-band signals by progressively decoding the pre-defined subband signals, transpare…

Cited by 6SourceScholar
2022

Adversarial Audio Synthesis Using a Harmonic-Percussive Discriminator

ICASSP 2022accepted

In this paper, we propose a discriminator design scheme for generative adversarial network-based audio signal generation. Unlike conventional discriminators that take an entire signal as input, our discriminator separates the audio signal into harmonic and percussive components and analyzes each com…

Cited by 0SourceScholar
2022

Pop-Out Motion: 3D-Aware Image Deformation via Learning the Shape Laplacian

CVPR 2022poster

We propose a framework that can deform an object in a 2D image as it exists in 3D space. Most existing methods for 3D-aware image manipulation are limited to (1) only changing the global scene information or depth, or (2) manipulating an object of specific categories. In this paper, we present a 3D-…

Cited by 3PDFScholar
2022

Progressive Multi-Stage Neural Audio Coding with Guided References

ICASSP 2022accepted

In this paper, we propose an effective multi-stage neural audio coding algorithm that encodes full-band audio signals (up to 20 kHz) using an end-to-end training criterion. By predefining several dyadic subband signals as training targets, we progressively encode input audio signals in each stage su…

Cited by 0SourceScholar
2021

Adaptable Multi-Domain Language Model for Transformer ASR

ICASSP 2021accepted

We propose an adapter based multi-domain Transformer based language model (LM) for Transformer ASR. The model consists of a big size common LM and small size adapters. The model can perform multi-domain adaptation with only the small size adapters and its related layers. The proposed model can reuse…

Cited by 0SourceScholar
2021

Partially Overlapped Inference for Long-Form Speech Recognition

ICASSP 2021accepted

While the end-to-end speech recognition models show impressive performance on many domains, they have difficulties in decoding long-form utterances. The overlapped inference algorithm with tie-breaking between two parallel hypotheses has been proposed for long-form speech recognition and shows drama…

Cited by 0SourceScholar
2019

Knowledge Distillation Using Output Errors for Self-attention End-to-end Models

ICASSP 2019accepted

Most automatic speech recognition (ASR) neural network models are not suitable for mobile devices due to their large model sizes. Therefore, it is required to reduce the model size to meet the limited hardware resources. In this study, we investigate sequence-level knowledge distillation techniques…

Cited by 0SourceScholar