← Search

Hyoungwoo Park

6 accepted papers

2026

Memory-Efficient Fine-Tuning Diffusion Transformers via Dynamic Patch Sampling and Block Skipping

CVPR 2026

Diffusion Transformers (DiTs) have significantly enhanced text-to-image (T2I) generation quality, enabling high-quality personalized content creation. However, fine-tuning these models requires substantial computational complexity and memory, limiting practical deployment under resource constraints.

Cited by 0SourceScholar
2025

ConsNoTrainLoRA: Data-driven Weight Initialization of Low-rank Adapters using Constraints

ICCV 2025poster

Foundation models are pre-trained on large-scale datasets and subsequently fine-tuned on small-scale datasets using parameter-efficient fine-tuning (PEFT) techniques like low-rank adapters (LoRA). In most previous works, LoRA weight matrices are randomly initialized with a fixed rank across all atta…

Cited by 0SourcePDFScholar
2025

Steering Guidance for Personalized Text-to-Image Diffusion Models

ICCV 2025poster

Personalizing text-to-image diffusion models is crucial for adapting the pre-trained models to specific target concepts, enabling diverse image generation. However, fine-tuning with few images introduces an inherent trade-off between aligning with the target distribution (e.g., subject fidelity) and…

Cited by 0SourcePDFScholar
2024

Balanced Learning for Multi-Domain Long-Tailed Speaker Recognition

ICASSP 2024accepted

This paper considers two types of imbalance problems commonly inherent in large-scale datasets: multiple domain and class imbalance. Class imbalance causes the algorithm to be biased toward the majority classes, and multiple-domain data results in significant performance disparities for different do…

Cited by 0SourceScholar
2021

Cross-Attentional Audio-Visual Fusion for Weakly-Supervised Action Localization

ICLR 2021poster

Temporally localizing actions in videos is one of the key components for video understanding. Learning from weakly-labeled data is seen as a potential solution towards avoiding expensive frame-level annotations. Different from other works which only depend on visual-modality, we propose to learn ric…

Cited by 77SourcePDFScholar
2021

Subspectral Normalization for Neural Audio Data Processing

ICASSP 2021accepted

Convolutional Neural Networks are widely used in various machine learning domains. In image processing, the features can be obtained by applying 2D convolution to all spatial dimensions of the input. However, in the audio case, frequency domain input like Mel-Spectrogram has different and unique cha…

Cited by 0SourceScholar