← Search

Yunji Kim

10 accepted papers

2026

Co-occurring Associated REtained concepts in Diffusion Unlearning

ICLR 2026poster

Unlearning has emerged as a key technique to mitigate harmful content generation in diffusion models. However, existing methods often remove not only the target concept, but also benign co-occurring concepts. Unlearning nudity can unintentionally suppress the concept of person, preventing a model fr…

Cited by 0SourceScholar
2025

Read, Watch and Scream! Sound Generation from Text and Video

AAAI 2025technical

Despite the impressive progress of multimodal generative models, video-to-audio generation still suffers from limited performance and limits the flexibility to prioritize sound synthesis for specific objects within the scene. Conversely, text-to-audio generation methods generate high-quality audio b…

2024

STELLA: Continual Audio-Video Pre-training with SpatioTemporal Localized Alignment

ICML 2024poster

Continuously learning a variety of audio-video semantics over time is crucial for audio-related reasoning tasks in our ever-evolving world. However, this is a nontrivial problem and poses two critical challenges: sparse spatio-temporal correlation between audio-video pairs and multimodal correlation…

Cited by 4SourcePDFScholar
2023

Dense Text-to-Image Generation with Attention Modulation

ICCV 2023poster

Existing text-to-image diffusion models struggle to synthesize realistic images given dense captions, where each text prompt provides a detailed description for a specific image region. To address this, we propose DenseDiffusion, a training-free method that adapts a pre-trained text-to-image model t…

Cited by 125PDFcodeScholar
2023

Self-Supervised Set Representation Learning for Unsupervised Meta-Learning

ICLR 2023poster

Unsupervised meta-learning (UML) essentially shares the spirit of self-supervised learning (SSL) in that their goal aims at learning models without any human supervision so that the models can be adapted to downstream tasks. Further, the learning objective of self-supervised learning, which pulls po…

Cited by 11SourcePDFScholar
2023

Text-Conditioned Sampling Framework for Text-to-Image Generation with Masked Generative Models

ICCV 2023poster

Token-based masked generative models are gaining popularity for their fast inference time with parallel decoding. While recent token-based approaches achieve competitive performance to diffusion-based models, their generation performance is still suboptimal as they sample multiple tokens simultaneou…

Cited by 5PDFScholar
2022

Mutual Information Divergence: A Unified Metric for Multimodal Generative Models

NeurIPS 2022accept

Text-to-image generation and image captioning are recently emerged as a new experimental paradigm to assess machine intelligence. They predict continuous quantity accompanied by their sampling techniques in the generation, making evaluation complicated and intractable to get marginal distributions.…

2019

Unsupervised Keypoint Learning for Guiding Class-Conditional Video Prediction

NeurIPS 2019poster

We propose a deep video prediction model conditioned on a single image and an action class. To generate future frames, we first detect keypoints of a moving object and predict future motion as a sequence of keypoints. The input image is then translated following the predicted keypoints sequence to c…

Cited by 58SourcePDFScholar
2018

Text-Adaptive Generative Adversarial Networks: Manipulating Images with Natural Language

NeurIPS 2018spotlight

This paper addresses the problem of manipulating images using natural language description. Our task aims to semantically modify visual attributes of an object in an image according to the text describing the new visual appearance. Although existing methods synthesize images having new attributes, t…

Cited by 263SourcePDFScholar