← Search

Sungrack Yun

21 accepted papers

2026

Ar2Can: An Architect and an Artist Leveraging a Canvas for Multi-Human Generation

CVPR 2026

Despite recent advances in personalized image generation, existing models consistently fail to produce reliable multi-human scenes, often merging or losing facial identity. We present Ar2Can, a novel two-stage framework that disentangles spatial planning from identity rendering for multi-human gener

Cited by 0SourcecodeScholar
2026

Memory-Efficient Fine-Tuning Diffusion Transformers via Dynamic Patch Sampling and Block Skipping

CVPR 2026

Diffusion Transformers (DiTs) have significantly enhanced text-to-image (T2I) generation quality, enabling high-quality personalized content creation. However, fine-tuning these models requires substantial computational complexity and memory, limiting practical deployment under resource constraints.

Cited by 0SourceScholar
2025

ConsNoTrainLoRA: Data-driven Weight Initialization of Low-rank Adapters using Constraints

ICCV 2025poster

Foundation models are pre-trained on large-scale datasets and subsequently fine-tuned on small-scale datasets using parameter-efficient fine-tuning (PEFT) techniques like low-rank adapters (LoRA). In most previous works, LoRA weight matrices are randomly initialized with a fixed rank across all atta…

Cited by 0SourcePDFScholar
2025

MultiHuman-Testbench: Benchmarking Image Generation for Multiple Humans

NeurIPS 2025poster

Generation of images containing multiple humans, performing complex actions, while preserving their facial identities, is a significant challenge. A major factor contributing to this is the lack of a a dedicated benchmark. To address this, we introduce MultiHuman-Testbench, a novel benchmark for rig…

Cited by 0SourceScholar
2025

Steering Guidance for Personalized Text-to-Image Diffusion Models

ICCV 2025poster

Personalizing text-to-image diffusion models is crucial for adapting the pre-trained models to specific target concepts, enabling diverse image generation. However, fine-tuning with few images introduces an inherent trade-off between aligning with the target distribution (e.g., subject fidelity) and…

Cited by 0SourcePDFScholar
2024

Balanced Learning for Multi-Domain Long-Tailed Speaker Recognition

ICASSP 2024accepted

This paper considers two types of imbalance problems commonly inherent in large-scale datasets: multiple domain and class imbalance. Class imbalance causes the algorithm to be biased toward the majority classes, and multiple-domain data results in significant performance disparities for different do…

Cited by 0SourceScholar
2024

Feature Diversification and Adaptation for Federated Domain Generalization

ECCV 2024poster

"Federated learning, a distributed learning paradigm, utilizes multiple clients to build a robust global model. In real-world applications, local clients often operate within their limited domains, leading to a ‘domain shift’ across clients. Privacy concerns limit each client’s learning to its own d…

Cited by 1SourcePDFScholar
2024

Hollowed Net for On-Device Personalization of Text-to-Image Diffusion Models

NeurIPS 2024poster

Recent advancements in text-to-image diffusion models have enabled the personalization of these models to generate custom images from textual prompts. This paper presents an efficient LoRA-based personalization approach for on-device subject-driven generation, where pre-trained diffusion models are…

Cited by 0SourcePDFScholar
2023

Few-Shot Common Action Localization via Cross-Attentional Fusion of Context and Temporal Dynamics

ICCV 2023poster

The goal of this paper is to localize action instances in a long untrimmed query video using just meager trimmed support videos representing a common action whose class information is not given. In this task, it is crucial to mine reliable temporal cues representing a common action from handful supp…

Cited by 7PDFScholar
2023

Label Shift Adapter for Test-Time Adaptation under Covariate and Label Shifts

ICCV 2023poster

Test-time adaptation (TTA) aims to adapt a pre-trained model to the target domain in a batch-by-batch manner during inference. While label distributions often exhibit imbalances in real-world scenarios, most previous TTA approaches typically assume that both source and target domain datasets have ba…

Cited by 22PDFScholar
2023

Progressive Random Convolutions for Single Domain Generalization

CVPR 2023poster

Single domain generalization aims to train a generalizable model with only one source domain to perform well on arbitrary unseen target domains. Image augmentation based on Random Convolutions (RandConv), consisting of one convolution layer randomly initialized for each mini-batch, enables the model…

2022

Improving Test-Time Adaptation via Shift-Agnostic Weight Regularization and Nearest Source Prototypes

ECCV 2022poster

"This paper proposes a novel test-time adaptation strategy that adjusts the model pre-trained on the source domain using only unlabeled online data from the target domain to alleviate the performance degradation due to the distribution shift between the source and target domains. Adapting the entire…

Cited by 78SourcePDFScholar
2022

Multi-Head Modularization to Leverage Generalization Capability in Multi-Modal Networks

AAAI 2022technical

It has been crucial to leverage the rich information of multiple modalities in many tasks. Existing works have tried to design multi-modal networks with descent multi-modal fusion modules. Instead, we focus on improving generalization capability of multi-modal networks, especially the fusion module.…

Cited by 1SourcePDFScholar
2021

Cross-Attentional Audio-Visual Fusion for Weakly-Supervised Action Localization

ICLR 2021poster

Temporally localizing actions in videos is one of the key components for video understanding. Learning from weakly-labeled data is seen as a potential solution towards avoiding expensive frame-level annotations. Different from other works which only depend on visual-modality, we propose to learn ric…

Cited by 77SourcePDFScholar
2021

Efficient Action Recognition via Dynamic Knowledge Propagation

ICCV 2021poster

Efficient action recognition has become crucial to extend the success of action recognition to many real-world applications. Contrary to most existing methods, which mainly focus on selecting salient frames to reduce the computation cost, we focus more on making the most of the selected frames. To t…

Cited by 30PDFScholar
2021

Federated Learning of User Verification Models Without Sharing Embeddings

ICML 2021spotlight

We consider the problem of training User Verification (UV) models in federated setup, where each user has access to the data of only one class and user embeddings cannot be shared with the server or other users. To address this problem, we propose Federated User Verification (FedUV), a framework in…

Cited by 31SourcePDFScholar
2021

Subspectral Normalization for Neural Audio Data Processing

ICASSP 2021accepted

Convolutional Neural Networks are widely used in various machine learning domains. In image processing, the features can be obtained by applying 2D convolution to all spatial dimensions of the input. However, in the audio case, frequency domain input like Mel-Spectrogram has different and unique cha…

Cited by 0SourceScholar