← Search

Jisoo Kim

14 accepted papers

2026

GEM: A Scale-Aware and Distribution-Sensitive Sparse Fine-Tuning Framework for Effective Downstream Adaptation

AAAI 2026technical

Parameter-efficient fine-tuning (PEFT) has become a popular way to adapt large pre-trained models to new tasks. Most PEFT methods update only a small subset of parameters while freezing the rest, avoiding redundant computation. As they maximize the absolute size of the updates without regard to the

Cited by 0SourcePDFScholar
2026

MFCC Inspired Spectral Feature Extraction for Robust Touch Interaction in Social_Robots

ICRA 2026poster

Touch is a fundamental modality for conveying emotions and intentions in Human–Robot Interaction. However, conventional approaches to touch pattern recognition often lack robustness to inter-user variability, whereas alternative solutions are frequently bulky or costly. This study proposes a novel f…

Cited by 0Scholar
2026

NEAF: Natural Image Editing with Attention Fusion for Generalizable Test-time Optimization in Text-Guided Image Editing

CVPR 2026

Diffusion-based text-to-image (T2I) models have enabled remarkable generative capabilities, yet precise text-based image editing that preserves the original's structural and perceptual fidelity remains non-trivial. Existing approaches either rely on retraining with large bespoke datasets, incurring

Cited by 0SourceScholar
2025

DEEPTalk: Dynamic Emotion Embedding for Probabilistic Speech-Driven 3D Face Animation

AAAI 2025technical

Speech-driven 3D facial animation has garnered lots of attention thanks to its broad range of applications. Despite recent advancements in achieving realistic lip motion, current methods fail to capture the nuanced emotional undertones conveyed through speech and produce monotonous facial motion. Th…

2025

DisCoRD: Discrete Tokens to Continuous Motion via Rectified Flow Decoding

ICCV 2025poster

Human motion is inherently continuous and dynamic, posing significant challenges for generative models. While discrete generation methods are widely used, they suffer from limited expressiveness and frame-wise noise artifacts. In contrast, continuous approaches produce smoother, more natural motion…

Cited by 0SourcePDFScholar
2025

EgoSpeak: Learning When to Speak for Egocentric Conversational Agents in the Wild

NAACL 2025findings

Predicting when to initiate speech in real-world environments remains a fundamental challenge for conversational agents. We introduce , a novel framework for real-time speech initiation prediction in egocentric streaming video. By modeling the conversation from the speaker’s first-person viewpoint,…

Cited by 0SourcePDFScholar
2025

Layer-wise Update Aggregation with Recycling for Communication-Efficient Federated Learning

NeurIPS 2025poster

Expensive communication cost is a common performance bottleneck in Federated Learning (FL), which makes it less appealing in real-world applications. Many communication-efficient FL methods focus on discarding a part of model updates mostly based on gradient magnitude. In this study, we find that re…

Cited by 0SourceScholar
2025

ParaHome: Parameterizing Everyday Home Activities Towards 3D Generative Modeling of Human-Object Interactions

CVPR 2025poster

To enable machines to understand the way humans interact with the physical world in daily life, 3D interaction signals should be captured in natural settings, allowing people to engage with multiple objects in a range of sequential and casual manipulations. To achieve this goal, we introduce our Par…

Cited by 17SourcePDFScholar
2025

Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues

ACL 2025long

Nonverbal communication is integral to human interaction, with gestures, facial expressions, and body language conveying critical aspects of intent and emotion. However, existing large language models (LLMs) fail to effectively incorporate these nonverbal elements, limiting their capacity to create…

2025

UDC-VIT: A Real-World Video Dataset for Under-Display Cameras

ICCV 2025poster

Even though an Under-Display Camera (UDC) is an advanced imaging system, the display panel significantly degrades captured images or videos, introducing low transmittance, blur, noise, and flare issues. Tackling such issues is challenging because of the complex degradation of UDCs, including diverse…

2025

V.I.P. : Iterative Online Preference Distillation for Efficient Video Diffusion Models

ICCV 2025poster

With growing interest in deploying text-to-video (T2V) models in resource-constrained environments, reducing their high computational cost has become crucial, leading to extensive research on pruning and knowledge distillation methods while maintaining performance. However, existing distillation met…

2023

Abusive Activity Detection with Multi-Modality Based on Convolutional Neural Network

ICASSP 2023accepted

Abusive activity frequently occurs in various places, including nursing homes for the elderly. However, it is difficult to detect because it has various forms and is not easy to define. Therefore, in this study, we try to detect using the Convolutional Neural Network (CNN). Data based on multimodali…

Cited by 0SourceScholar
2023

TextManiA: Enriching Visual Feature by Text-driven Manifold Augmentation

ICCV 2023poster

We propose TextManiA, a text-driven manifold augmentation method that semantically enriches visual feature spaces, regardless of class distribution. TextManiA augments visual data with intra-class semantic perturbation by exploiting easy-to-understand visually mimetic words, i.e., attributes. This w…

Cited by 10PDFScholar
2020

Multi-Temporal Recurrent Neural Networks For Progressive Non-Uniform Single Image Deblurring With Incremental Temporal Training

ECCV 2020poster

Blind non-uniform image deblurring for severe blurs induced by large motions is still challenging. Multi-scale (MS) approach has been widely used for deblurring that sequentially recovers the downsampled original image in low spatial scale first and then further restores in high spatial scale using…